Skip to content

Providers overview

Myra AI Workspace supports 21 inference providers. Each provider is accessible via its own native endpoint and via the unified OpenAI-compatible (compat) endpoint.

Supported providers

Provider Wire format Auth Notes
OpenAI Native OpenAI Bearer Direct pass-through.
Azure OpenAI OpenAI api-key header Native path reads azure_resource, azure_deployment, azure_api_version from gateway config (each has a default). See Azure OpenAI.
Anthropic Messages API x-api-key System prompt extraction, extended thinking, prompt caching.
Google Gemini GenerateContent x-goog-api-key AI Studio API key.
Vertex AI GenerateContent Authorization: Bearer Service-account JSON. The gateway mints a short-lived OAuth2 token (RFC 7523 JWT-bearer). Requires vertex_project (or a deployment-wide default project set by your operator); vertex_region is optional, defaulting to a deployment-wide default region and then us-central1.
AWS Bedrock Bedrock InvokeModel (per-family) SigV4 HMAC-SHA256 signing; requires bedrock_region in gateway config. Streaming not implemented.
Mistral AI OpenAI-compatible Bearer
Groq OpenAI-compatible Bearer
Together AI OpenAI-compatible Bearer meta-llama/, deepseek-ai/, Qwen/, zai-org/ (Zhipu GLM) model prefixes.
Fireworks OpenAI-compatible Bearer accounts/fireworks/models/ model prefix.
Cerebras OpenAI-compatible Bearer Fast inference.
DeepSeek OpenAI-compatible Bearer deepseek- model prefix.
OpenRouter OpenAI-compatible Bearer 300+ models; universal compat fallback.
Perplexity OpenAI-compatible Bearer sonar- model prefix.
SambaNova OpenAI-compatible Bearer
xAI OpenAI-compatible Bearer grok- model prefix.
NVIDIA NIM OpenAI-compatible Bearer nvidia/ model prefix.
Cloudflare Workers AI OpenAI-compatible Bearer Requires cf_account_id in gateway config.
Cohere Cohere Chat V2 Bearer Native request/response translation.
HuggingFace OpenAI-compatible Bearer Serverless or dedicated endpoint via hf_endpoint.
Myra OpenAI-compatible Internal Myra-hosted models (Qwen3, Gemma, Mistral) served on Myra's EU-hosted inference fleet. Uses an internal token, not a BYOK key. Web search is supported via the gateway's streaming tool loop.

Endpoint patterns

Native (provider-specific) endpoint

Each provider has a dedicated path that preserves the exact wire format of the provider:

POST /v1/{tenant}/{gateway}/{provider}/chat/completions

⭐ Example: The following table shows native endpoint paths per provider.

Provider Native endpoint
OpenAI POST /v1/{tenant}/{gateway}/openai/chat/completions
Anthropic POST /v1/{tenant}/{gateway}/anthropic/chat/completions
Google Gemini POST /v1/{tenant}/{gateway}/gemini/chat/completions
AWS Bedrock POST /v1/{tenant}/{gateway}/bedrock/chat/completions
Azure OpenAI POST /v1/{tenant}/{gateway}/azure/chat/completions

All providers accept the OpenAI chat completions request body. The gateway translates to the native wire format of each provider automatically.

Compat (unified) endpoint

POST /v1/{tenant}/{gateway}/compat/chat/completions

The compat endpoint routes to the correct provider by inspecting the model field in the request body. It always returns an OpenAI-shaped response regardless of which provider handled the request.

curl -s -X POST "https://gateway.example.com/v1/myapp/prod/compat/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <token>" \
  -d '{
    "model": "claude-opus-4-6",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Compat model resolution

When a request arrives at the compat endpoint, the gateway resolves the provider in three tiers:

See Unified compat endpoint for the resolution flow diagram.

Tier Mechanism Example
1 Exact model-name match A known model name maps directly to its provider (e.g. claude-opus-4-6 → Anthropic)
2 Model-name prefix match claude- → Anthropic; grok- → xAI
3 OpenRouter fallback A model name that matches no exact entry and no prefix — only when OpenRouter is configured on the gateway; otherwise 400 model_not_found

To pin a specific provider instead of resolving by model name, address the provider directly in the URL path (/v1/{tenant}/{gateway}/{provider}/chat/completions) rather than the compat pseudo-provider.

Tier 2 prefix examples:

Model prefix Resolved provider
gpt-, o1-, o3-, o4- OpenAI
claude- Anthropic
gemini- Google Gemini
mistral-, codestral- Mistral AI
meta-llama/, deepseek-ai/, Qwen/, zai-org/ Together AI
accounts/fireworks/models/ Fireworks
deepseek- DeepSeek
sonar- Perplexity
grok- xAI
nvidia/ NVIDIA NIM
@cf/ Cloudflare Workers AI
command-, embed- Cohere
myra/ Myra (self-hosted EU fleet)

💡 Note: If a model name matches no exact entry and no prefix, the request is forwarded to OpenRouter (which supports 300+ models) only when the gateway has an openrouter provider configured (a stored OpenRouter BYOK key). Without OpenRouter configured, an unclassifiable model id returns 400 model_not_found.

Model id normalization

Provider inference (above) and model id normalization are two different mechanisms — do not conflate them.

After the provider is resolved, the gateway removes a leading <prefix>/ from the model id only when that exact prefix is registered as a routing namespace for the resolved provider. The normalized id is what the gateway uses everywhere downstream: the model price catalog lookup, the recorded cost, the model column of the request log and analytics, the token-limit and capability catalog, and the X-AIG-Model response header. Sending myra/qwen3.8-27b and sending qwen3.8-27b therefore produce the same routing decision, the same cost and one single analytics row.

The registered routing namespaces are:

Provider Stripped prefix
Google Gemini gemini/
Vertex AI vertex_ai/
Azure azure_ai/, azure/
Groq groq/
Mistral AI mistral/, text-completion-codestral/
Together AI together_ai/
Fireworks fireworks_ai/
NVIDIA NIM nvidia_nim/
SambaNova sambanova/
DeepSeek deepseek/
xAI xai/
Perplexity perplexity/
Cerebras cerebras/
Cohere cohere/
Bedrock bedrock/
OpenRouter openrouter/
Myra myra/

Every other prefix is left untouched, deliberately. For many providers the slash is part of the model's real catalog name rather than a routing prefix — openai/gpt-4o-mini and meta-llama/llama-3-70b-instruct on OpenRouter, meta-llama/… and Qwen/… on Together AI, accounts/fireworks/models/… on Fireworks, nvidia/… and meta/… on NVIDIA NIM, @cf/… on Cloudflare Workers AI, and the repository ids on Hugging Face. Those names are stored and billed exactly as the provider publishes them; stripping them would break both routing and pricing.

Two things worth knowing:

  • A prefix is only removed when it matches the provider the request actually resolved to. myra/qwen3.8-27b sent to an OpenRouter-resolved route keeps its prefix.
  • Only one leading segment is ever removed.

💡 Note: Azure, Bedrock and OpenRouter are a known exception to the rule above: parts of their catalog use azure/…, bedrock/… and openrouter/… as genuine model names, so for those three providers sending the prefixed form and the bare form is not equivalent. Address them with the exact model id the catalog lists (GET /admin/v1/models).

Provider base URL override

Every provider has a hardcoded default base URL (for example, https://api.openai.com). Override this per gateway using the provider_base_urls config field. The value is a base (scheme, host, optional port and path prefix): the gateway appends the provider's own endpoint path to it, and a query string or fragment on the override is discarded before that (so an override cannot change which endpoint a request is read as).

Each key must be a supported provider slug — one the gateway can route to (in the gateway settings UI the provider is a dropdown of these, not free text). A key that is not a supported provider is rejected on save (400), because an override is looked up at request time by the exact provider name, so an unknown key would never route. A legacy alias such as vllm is rejected with a message naming its canonical id (myra); use the canonical id. (This differs from provider_allowlist, which accepts the vllm alias and folds it to myra — that field is consumed by canonical lookup, whereas a base-URL key is matched verbatim.)

curl -X PATCH "https://gateway.example.com/admin/v1/gateways/{id}" \
  -H "Cookie: aig_admin=<SESSION>" \
  -H "Content-Type: application/json" \
  -d '{
    "config": {
      "provider_base_urls": {
        "openai": "https://proxy.example.com/openai"
      }
    }
  }'

⚠️ The override is SSRF-guarded. The gateway resolves the override URL once and rejects any that resolves to a private, loopback, or otherwise internal (RFC1918) address — the request fails with CONFIGURATION_ERROR ("URL resolves to a private/internal address"), and the pinned IP prevents a later DNS-rebind. An override must resolve to a public address, so a proxy or staging endpoint has to be publicly reachable (a purely internal *.internal host is blocked).

Use cases:

  • Routing through a publicly reachable OpenAI-compatible proxy
  • Testing against a public staging endpoint

Provider header pass-through

A request header prefixed with x-aig-provider- is stripped of that prefix and forwarded to the upstream provider only when the stripped name is on the allow-list. The forwardable names are exactly anthropic-beta, openai-organization, and openai-project — the documented client-controlled provider headers. Every other stripped name is dropped and logged at warn; nothing else reaches the upstream request. This is an allow-list, not a blocklist: a credential (authorization, x-api-key), a request-framing / hop-by-hop header (host, content-length, transfer-encoding, connection, expect, te, upgrade, trailer, proxy-authorization), and any unrecognised name are all refused — so x-aig-provider-authorization cannot inject an upstream credential and x-aig-provider-content-length cannot tamper with request framing on the shared upstream connection pool.

At most 16 overrides and 8 KB of override data are forwarded per request; anything beyond that is dropped with a warn.

Client sends:

x-aig-provider-anthropic-beta: interleaved-thinking-2025-05-14
Gateway forwards to provider:
anthropic-beta: interleaved-thinking-2025-05-14

This lets you pass provider-specific beta flags, versioning headers, or experimental features without waiting for gateway-level support.

⚠️ Caution: An allow-listed x-aig-provider-* header is forwarded regardless of which provider handles the request. Sending x-aig-provider-openai-organization to an Anthropic request is harmless (Anthropic ignores the unknown header), but the header is not filtered by provider. Only the three allow-listed names are ever forwarded.

See also