Providers overview
Myra AI Workspace supports 21 inference providers. Each provider is accessible via its own native endpoint and via the unified OpenAI-compatible (compat) endpoint.
Supported providers
| Provider | Wire format | Auth | Notes |
|---|---|---|---|
| OpenAI | Native OpenAI | Bearer | Direct pass-through. |
| Azure OpenAI | OpenAI | api-key header |
Native path reads azure_resource, azure_deployment, azure_api_version from gateway config (each has a default). See Azure OpenAI. |
| Anthropic | Messages API | x-api-key |
System prompt extraction, extended thinking, prompt caching. |
| Google Gemini | GenerateContent | x-goog-api-key |
AI Studio API key. |
| Vertex AI | GenerateContent | Authorization: Bearer |
Service-account JSON. The gateway mints a short-lived OAuth2 token (RFC 7523 JWT-bearer). Requires vertex_project (or a deployment-wide default project set by your operator); vertex_region is optional, defaulting to a deployment-wide default region and then us-central1. |
| AWS Bedrock | Bedrock InvokeModel (per-family) | SigV4 | HMAC-SHA256 signing; requires bedrock_region in gateway config. Streaming not implemented. |
| Mistral AI | OpenAI-compatible | Bearer | |
| Groq | OpenAI-compatible | Bearer | |
| Together AI | OpenAI-compatible | Bearer | meta-llama/, deepseek-ai/, Qwen/, zai-org/ (Zhipu GLM) model prefixes. |
| Fireworks | OpenAI-compatible | Bearer | accounts/fireworks/models/ model prefix. |
| Cerebras | OpenAI-compatible | Bearer | Fast inference. |
| DeepSeek | OpenAI-compatible | Bearer | deepseek- model prefix. |
| OpenRouter | OpenAI-compatible | Bearer | 300+ models; universal compat fallback. |
| Perplexity | OpenAI-compatible | Bearer | sonar- model prefix. |
| SambaNova | OpenAI-compatible | Bearer | |
| xAI | OpenAI-compatible | Bearer | grok- model prefix. |
| NVIDIA NIM | OpenAI-compatible | Bearer | nvidia/ model prefix. |
| Cloudflare Workers AI | OpenAI-compatible | Bearer | Requires cf_account_id in gateway config. |
| Cohere | Cohere Chat V2 | Bearer | Native request/response translation. |
| HuggingFace | OpenAI-compatible | Bearer | Serverless or dedicated endpoint via hf_endpoint. |
| Myra | OpenAI-compatible | Internal | Myra-hosted models (Qwen3, Gemma, Mistral) served on Myra's EU-hosted inference fleet. Uses an internal token, not a BYOK key. Web search is supported via the gateway's streaming tool loop. |
Endpoint patterns
Native (provider-specific) endpoint
Each provider has a dedicated path that preserves the exact wire format of the provider:
⭐ Example: The following table shows native endpoint paths per provider.
| Provider | Native endpoint |
|---|---|
| OpenAI | POST /v1/{tenant}/{gateway}/openai/chat/completions |
| Anthropic | POST /v1/{tenant}/{gateway}/anthropic/chat/completions |
| Google Gemini | POST /v1/{tenant}/{gateway}/gemini/chat/completions |
| AWS Bedrock | POST /v1/{tenant}/{gateway}/bedrock/chat/completions |
| Azure OpenAI | POST /v1/{tenant}/{gateway}/azure/chat/completions |
All providers accept the OpenAI chat completions request body. The gateway translates to the native wire format of each provider automatically.
Compat (unified) endpoint
The compat endpoint routes to the correct provider by inspecting the model field in the request body. It always returns an OpenAI-shaped response regardless of which provider handled the request.
curl -s -X POST "https://gateway.example.com/v1/myapp/prod/compat/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <token>" \
-d '{
"model": "claude-opus-4-6",
"messages": [{"role": "user", "content": "Hello"}]
}'
Compat model resolution
When a request arrives at the compat endpoint, the gateway resolves the provider in three tiers:
See Unified compat endpoint for the resolution flow diagram.
| Tier | Mechanism | Example |
|---|---|---|
| 1 | Exact model-name match | A known model name maps directly to its provider (e.g. claude-opus-4-6 → Anthropic) |
| 2 | Model-name prefix match | claude- → Anthropic; grok- → xAI |
| 3 | OpenRouter fallback | A model name that matches no exact entry and no prefix — only when OpenRouter is configured on the gateway; otherwise 400 model_not_found |
To pin a specific provider instead of resolving by model name, address the provider directly in the URL path (/v1/{tenant}/{gateway}/{provider}/chat/completions) rather than the compat pseudo-provider.
Tier 2 prefix examples:
| Model prefix | Resolved provider |
|---|---|
gpt-, o1-, o3-, o4- |
OpenAI |
claude- |
Anthropic |
gemini- |
Google Gemini |
mistral-, codestral- |
Mistral AI |
meta-llama/, deepseek-ai/, Qwen/, zai-org/ |
Together AI |
accounts/fireworks/models/ |
Fireworks |
deepseek- |
DeepSeek |
sonar- |
Perplexity |
grok- |
xAI |
nvidia/ |
NVIDIA NIM |
@cf/ |
Cloudflare Workers AI |
command-, embed- |
Cohere |
myra/ |
Myra (self-hosted EU fleet) |
💡 Note: If a model name matches no exact entry and no prefix, the request is forwarded to OpenRouter (which supports 300+ models) only when the gateway has an
openrouterprovider configured (a stored OpenRouter BYOK key). Without OpenRouter configured, an unclassifiable model id returns400 model_not_found.
Model id normalization
Provider inference (above) and model id normalization are two different mechanisms — do not conflate them.
After the provider is resolved, the gateway removes a leading <prefix>/ from the model id only when that exact prefix is registered as a routing namespace for the resolved provider. The normalized id is what the gateway uses everywhere downstream: the model price catalog lookup, the recorded cost, the model column of the request log and analytics, the token-limit and capability catalog, and the X-AIG-Model response header. Sending myra/qwen3.8-27b and sending qwen3.8-27b therefore produce the same routing decision, the same cost and one single analytics row.
The registered routing namespaces are:
| Provider | Stripped prefix |
|---|---|
| Google Gemini | gemini/ |
| Vertex AI | vertex_ai/ |
| Azure | azure_ai/, azure/ |
| Groq | groq/ |
| Mistral AI | mistral/, text-completion-codestral/ |
| Together AI | together_ai/ |
| Fireworks | fireworks_ai/ |
| NVIDIA NIM | nvidia_nim/ |
| SambaNova | sambanova/ |
| DeepSeek | deepseek/ |
| xAI | xai/ |
| Perplexity | perplexity/ |
| Cerebras | cerebras/ |
| Cohere | cohere/ |
| Bedrock | bedrock/ |
| OpenRouter | openrouter/ |
| Myra | myra/ |
Every other prefix is left untouched, deliberately. For many providers the slash is part of the model's real catalog name rather than a routing prefix — openai/gpt-4o-mini and meta-llama/llama-3-70b-instruct on OpenRouter, meta-llama/… and Qwen/… on Together AI, accounts/fireworks/models/… on Fireworks, nvidia/… and meta/… on NVIDIA NIM, @cf/… on Cloudflare Workers AI, and the repository ids on Hugging Face. Those names are stored and billed exactly as the provider publishes them; stripping them would break both routing and pricing.
Two things worth knowing:
- A prefix is only removed when it matches the provider the request actually resolved to.
myra/qwen3.8-27bsent to an OpenRouter-resolved route keeps its prefix. - Only one leading segment is ever removed.
💡 Note: Azure, Bedrock and OpenRouter are a known exception to the rule above: parts of their catalog use
azure/…,bedrock/…andopenrouter/…as genuine model names, so for those three providers sending the prefixed form and the bare form is not equivalent. Address them with the exact model id the catalog lists (GET /admin/v1/models).
Provider base URL override
Every provider has a hardcoded default base URL (for example, https://api.openai.com). Override this per gateway using the provider_base_urls config field. The value is a base (scheme, host, optional port and path prefix): the gateway appends the provider's own endpoint path to it, and a query string or fragment on the override is discarded before that (so an override cannot change which endpoint a request is read as).
Each key must be a supported provider slug — one the gateway can route to (in the gateway settings UI the provider is a dropdown of these, not free text). A key that is not a supported provider is rejected on save (400), because an override is looked up at request time by the exact provider name, so an unknown key would never route. A legacy alias such as vllm is rejected with a message naming its canonical id (myra); use the canonical id. (This differs from provider_allowlist, which accepts the vllm alias and folds it to myra — that field is consumed by canonical lookup, whereas a base-URL key is matched verbatim.)
curl -X PATCH "https://gateway.example.com/admin/v1/gateways/{id}" \
-H "Cookie: aig_admin=<SESSION>" \
-H "Content-Type: application/json" \
-d '{
"config": {
"provider_base_urls": {
"openai": "https://proxy.example.com/openai"
}
}
}'
⚠️ The override is SSRF-guarded. The gateway resolves the override URL once and rejects any that resolves to a private, loopback, or otherwise internal (RFC1918) address — the request fails with
CONFIGURATION_ERROR("URL resolves to a private/internal address"), and the pinned IP prevents a later DNS-rebind. An override must resolve to a public address, so a proxy or staging endpoint has to be publicly reachable (a purely internal*.internalhost is blocked).
Use cases:
- Routing through a publicly reachable OpenAI-compatible proxy
- Testing against a public staging endpoint
Provider header pass-through
A request header prefixed with x-aig-provider- is stripped of that prefix and forwarded to the upstream provider only when the stripped name is on the allow-list. The forwardable names are exactly anthropic-beta, openai-organization, and openai-project — the documented client-controlled provider headers. Every other stripped name is dropped and logged at warn; nothing else reaches the upstream request. This is an allow-list, not a blocklist: a credential (authorization, x-api-key), a request-framing / hop-by-hop header (host, content-length, transfer-encoding, connection, expect, te, upgrade, trailer, proxy-authorization), and any unrecognised name are all refused — so x-aig-provider-authorization cannot inject an upstream credential and x-aig-provider-content-length cannot tamper with request framing on the shared upstream connection pool.
At most 16 overrides and 8 KB of override data are forwarded per request; anything beyond that is dropped with a warn.
Client sends:
Gateway forwards to provider:This lets you pass provider-specific beta flags, versioning headers, or experimental features without waiting for gateway-level support.
⚠️ Caution: An allow-listed
x-aig-provider-*header is forwarded regardless of which provider handles the request. Sendingx-aig-provider-openai-organizationto an Anthropic request is harmless (Anthropic ignores the unknown header), but the header is not filtered by provider. Only the three allow-listed names are ever forwarded.