OpenAI-compatible providers
Ten providers expose an OpenAI-compatible chat-completions API and require no provider-specific gateway configuration beyond the BYOK key. They share the same request body, Authorization: Bearer <KEY> authentication, and response shape. The differences between them are the base URL, the expected model naming convention, and the upstream model catalogue.
Providers covered by their own page are excluded from this list. See Mistral, Cohere, HuggingFace, Cloudflare Workers AI, Azure OpenAI, Google Vertex AI for those.
Provider list
| Provider id | Drop-down label | Model naming convention | Notes |
|---|---|---|---|
groq |
Groq | bare names, e.g. llama-3.3-70b-versatile |
Custom-hardware inference. |
together |
Together AI | org-prefixed, e.g. meta-llama/Llama-3.3-70B-Instruct-Turbo or zai-org/GLM-5.2 |
Large open-model catalogue, incl. Zhipu GLM (zai-org/). |
fireworks |
Fireworks | full path accounts/fireworks/models/<name> |
Requires the full account-prefixed model path. |
cerebras |
Cerebras | bare names, e.g. llama3.1-70b |
Very fast inference. |
deepseek |
DeepSeek | deepseek-chat, deepseek-reasoner |
DeepSeek frontier models. |
openrouter |
OpenRouter | provider-prefixed, e.g. anthropic/claude-3.5-sonnet |
Aggregator with 300+ models from many vendors. |
perplexity |
Perplexity | sonar, sonar-pro, etc. |
Web-grounded generation. |
sambanova |
SambaNova | Meta-Llama-3.1-405B-Instruct, etc. |
Enterprise inference. |
xai |
xAI | grok-3, grok-3-mini |
Grok models. |
nvidia |
NVIDIA NIM | nvidia/llama-3.1-nemotron-70b-instruct, etc. |
nvidia/-prefixed model names. |
💡 Note: Myra also uses an OpenAI-compatible wire format, but it is a Myra-hosted provider that uses an internal token rather than a BYOK key. Myra is therefore not part of this BYOK list. See Providers Overview.
BYOK setup
The procedure for storing a BYOK key is the same for every provider. See Provider keys (BYOK) for the steps. Select the matching provider id value (column above) in the Provider drop-down list and paste the API key in the API Key field.
curl -X POST "https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/keys" \
-H "Content-Type: application/json" \
-d '{"provider": "groq", "alias": "default", "key": "gsk_..."}'
Replace "groq" with any value from the Provider id column.
Inference endpoints
Each provider responds at its native gateway path:
For example, POST /v1/myapp/production/groq/chat/completions.
Compat endpoint
The unified compat endpoint resolves the provider from the model name. When the model name carries an unambiguous prefix, the compat endpoint routes to the correct provider automatically:
curl -X POST "https://<your-gateway-host>/v1/myapp/production/compat/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <TOKEN>" \
-d '{
"model": "deepseek-chat",
"messages": [{"role": "user", "content": "Hello"}]
}'
When the compat endpoint cannot match a model name to a known provider, it falls back to OpenRouter — but only when the gateway has an OpenRouter key configured. Without one, no fallback is attempted and the request is rejected with a gateway-side 400 model_not_found. To use OpenRouter explicitly, address it through its native path (/v1/{tenant}/{gateway}/openrouter/chat/completions).
💡 Note: The OpenRouter fallback requires a stored OpenRouter BYOK key. Without a stored key the compat endpoint does not attempt the fallback at all — it returns
400 model_not_foundfrom the gateway, and no request is sent to OpenRouter.
Request example
curl -X POST "https://<your-gateway-host>/v1/myapp/production/groq/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <TOKEN>" \
-d '{
"model": "llama-3.3-70b-versatile",
"messages": [
{"role": "system", "content": "Be brief."},
{"role": "user", "content": "What is LoRA fine-tuning?"}
],
"temperature": 0.5,
"max_tokens": 256
}'