Skip to content
SayGMDocs
Dashboard

Routing

Most models on SayGM are available from several providers, each with its own price and speed. SayGM picks one for every request. Two optional headers let you steer that choice:

Header Values Use it to
X-GM-Routing default, cheapest, fastest Prefer the lowest price or the quickest response.
X-GM-Providers Comma-separated provider names Choose which providers may receive your request.

Add them to any Chat Completions, Responses, Anthropic Messages, Gemini, or Images request. They apply to that request only, so you can mix modes freely, or set them once as default headers on your client.

X-GM-Routing: cheapest
  • default uses SayGM’s standard routing for the model. Leaving the header out does the same.
  • cheapest sends the request to the provider with the lowest expected charge, favoring providers with a strong recent success rate.
  • fastest sends the request to the provider with the quickest recent time to first token and output speed.

Both choices are based on recent prices and performance. Your actual charge is the settled cost returned with each response. If SayGM retries your request on another provider, it keeps the same mode.

In a multi-turn conversation, SayGM keeps sending turns to the same provider while it stays available, so prompt caching keeps working and later turns stay fast and cheap. Your routing mode picks the provider at the start of a conversation, after five minutes of inactivity, or when the previous provider becomes unavailable.

Routing modes work with single-model requests. Fusion and multi-model cascade requests, and some multimodal requests, use default routing; setting cheapest or fastest on them returns HTTP 400.

X-GM-Providers: bedrock,anthropic

Your request goes only to the providers you list. Use this when a contract, data-residency requirement, or your own review means only certain companies may process your prompts.

Provider names identify the service that runs the model. They include cloud platforms such as bedrock, azure, and foundry, model makers such as anthropic, openai, and google, and inference hosts such as kubetee, chutes, near, and deepinfra.

List up to 16 provider names, separated by commas.

Your list applies with every routing mode and on every retry. If none of your chosen providers can serve the request, SayGM returns an error, so your prompt only ever reaches providers you picked.

Send a request with a provider name you know is not listed, and the error message names the providers currently offering that model:

{
"error": {
"message": "no_allowed_provider: No route for gpt-5.5 is available on requested providers [bedrock]. Current providers offering this model: [azure, openai].",
"type": "invalid_request_error",
"code": "no_allowed_provider",
"param": null
}
}

Providers join and leave over time, so check the list again when you change models or see this error.

Status Code Meaning What to do
400 invalid_provider_restriction The provider list could not be read. Check the provider names and send the header once.
400 no_allowed_provider None of your providers offers this model. Add a provider from the error message, or choose another model.
400 — The request is a fusion or multi-model cascade request. Remove the header.
400 — A Responses follow-up (previous_response_id or encrypted reasoning) belongs to a provider outside your list. Add that provider, or start a new conversation.
503 — Your providers offer the model but are busy right now. Retry with backoff, or add a provider.

cURL:

Terminal window
curl "https://api.saygm.com/v1/chat/completions" \
-H "Authorization: Bearer $GM_API_KEY" \
-H "Content-Type: application/json" \
-H "X-GM-Routing: fastest" \
-H "X-GM-Providers: openai,azure" \
-d '{
"model": "gpt-5.5",
"messages": [{"role": "user", "content": "Say hello."}]
}'

OpenAI Python SDK, for every request from a client:

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.saygm.com/v1",
api_key=os.environ["GM_API_KEY"],
default_headers={"X-GM-Routing": "cheapest"},
)

Or for a single request:

response = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "Say hello."}],
extra_headers={"X-GM-Providers": "openai"},
)

The Anthropic SDKs accept the same default_headers and extra_headers options. In Node.js, both SDKs use defaultHeaders on the client and a headers option per request.