Skip to content

OpenAI-compatible providers

Ten providers expose an OpenAI-compatible chat-completions API and require no provider-specific gateway configuration beyond the BYOK key. They share the same request body, Authorization: Bearer <KEY> authentication, and response shape. The differences between them are the base URL, the expected model naming convention, and the upstream model catalogue.

Providers covered by their own page are excluded from this list. See Mistral, Cohere, HuggingFace, Cloudflare Workers AI, Azure OpenAI, Google Vertex AI for those.

Provider list

Provider id Drop-down label Model naming convention Notes
groq Groq bare names, e.g. llama-3.3-70b-versatile Custom-hardware inference.
together Together AI org-prefixed, e.g. meta-llama/Llama-3.3-70B-Instruct-Turbo or zai-org/GLM-5.2 Large open-model catalogue, incl. Zhipu GLM (zai-org/).
fireworks Fireworks full path accounts/fireworks/models/<name> Requires the full account-prefixed model path.
cerebras Cerebras bare names, e.g. llama3.1-70b Very fast inference.
deepseek DeepSeek deepseek-chat, deepseek-reasoner DeepSeek frontier models.
openrouter OpenRouter provider-prefixed, e.g. anthropic/claude-3.5-sonnet Aggregator with 300+ models from many vendors.
perplexity Perplexity sonar, sonar-pro, etc. Web-grounded generation.
sambanova SambaNova Meta-Llama-3.1-405B-Instruct, etc. Enterprise inference.
xai xAI grok-3, grok-3-mini Grok models.
nvidia NVIDIA NIM nvidia/llama-3.1-nemotron-70b-instruct, etc. nvidia/-prefixed model names.

💡 Note: Myra also uses an OpenAI-compatible wire format, but it is a Myra-hosted provider that uses an internal token rather than a BYOK key. Myra is therefore not part of this BYOK list. See Providers Overview.

BYOK setup

The procedure for storing a BYOK key is the same for every provider. See Provider keys (BYOK) for the steps. Select the matching provider id value (column above) in the Provider drop-down list and paste the API key in the API Key field.

curl -X POST "https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/keys" \
  -H "Content-Type: application/json" \
  -d '{"provider": "groq", "alias": "default", "key": "gsk_..."}'

Replace "groq" with any value from the Provider id column.

Inference endpoints

Each provider responds at its native gateway path:

POST /v1/<TENANT_SLUG>/<GATEWAY_SLUG>/<PROVIDER>/chat/completions

For example, POST /v1/myapp/production/groq/chat/completions.

Compat endpoint

The unified compat endpoint resolves the provider from the model name. When the model name carries an unambiguous prefix, the compat endpoint routes to the correct provider automatically:

curl -X POST "https://<your-gateway-host>/v1/myapp/production/compat/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <TOKEN>" \
  -d '{
    "model": "deepseek-chat",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

When the compat endpoint cannot match a model name to a known provider, it falls back to OpenRouter — but only when the gateway has an OpenRouter key configured. Without one, no fallback is attempted and the request is rejected with a gateway-side 400 model_not_found. To use OpenRouter explicitly, address it through its native path (/v1/{tenant}/{gateway}/openrouter/chat/completions).

💡 Note: The OpenRouter fallback requires a stored OpenRouter BYOK key. Without a stored key the compat endpoint does not attempt the fallback at all — it returns 400 model_not_found from the gateway, and no request is sent to OpenRouter.

Request example

curl -X POST "https://<your-gateway-host>/v1/myapp/production/groq/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <TOKEN>" \
  -d '{
    "model": "llama-3.3-70b-versatile",
    "messages": [
      {"role": "system", "content": "Be brief."},
      {"role": "user",   "content": "What is LoRA fine-tuning?"}
    ],
    "temperature": 0.5,
    "max_tokens": 256
  }'

See also