Error codes
All errors returned by the gateway (admin API and inference endpoints) follow a consistent JSON structure.
Error response format
{
"error": {
"code": "invalid_request",
"message": "Human-readable detail about what went wrong"
}
}
The code field is a stable machine-readable string. The message is informational and may change between releases. Always branch logic on code, not message.
Error codes
💡 Note: Inference endpoint errors (
/v1/...) use the structured format below withcodeandmessagefields. Admin API errors (/admin/v1/...) use a simpler flat format,{"error": "message string"}— though a few admin endpoints add a sibling top-levelcodefor outcomes a client must branch on (e.g. the last-administrator delete guard returns{"error": …, "code": "last_admin_confirm_required" | "last_admin_blocked"}— see Deleting a user; voucher creation returns acodefor every validation rejection — see Vouchers; a tenant create, or a gateway create withcreate_only, whose name already exists returnsslug_taken(409) — see Tenants & gateways; deleting a custom role that a user outside your reach holds returnsrole_held_by_unreachable_user(409, with acount) — see Deleting a custom role; app-feedback submission returnsinvalid_type,description_required(400 — a missing/blank Details field) andsupport_rate_limited— see Feedback; the Contact us enquiryPOST /admin/v1/contactreturnsinvalid_contact_email(400 — a present-but-malformed email; an absent one falls back to the account email),invalid_contact_field(400 — a non-stringname/message/url),support_rate_limited(429 — the shared support-egress window),contact_send_failed(502 — the relay rejected the message, logged and auditeddelivered=false, retryable) andcontact_unavailable(502 — the configured recipient is invalid, nothing sent); the one-time-code sign-in routes returninvalid_email(400) for a malformed address and the three429capsotp_request_cap(too many codes requested for the address),otp_attempt_cap(this code's five guesses are spent — request a new code) andotp_failure_cap(twenty rejected attempts in an hour, wrong guesses andotp_attempt_capanswers alike — wait) — see Authentication; the public cancellation button carries its machine code inerroritself —email_required/invalid_email/invalid_code/malformed_body(400),code_invalid(401: wrong, expired or no live code),code_request_cap(429: five codes an hour for the address — a mailed code still works),code_attempt_cap(429: request a new code) andcode_failure_cap(429: wait an hour) — see Billing; a project or tenant knowledge write whose file name already exists returnsknowledge_name_conflict(409) — see Projects and Tenant Knowledge; the route-resolution refusals ofPOST/PATCH /admin/v1/conversationsreturn the machine string in BOTHerrorand a siblingcode—no_runnable_route(no gateway that routes can serve the request: the tenant has nopurpose: "production"gateway at all, or none of them serves a model that survives the plan / residency / allowlist filters) andno_local_route(409), andmodel_override_gateway_not_visible,project_gateway_not_visibleandunknown_task(400) — see Conversations; on those refusals the two fields are equal and that value is frozen, while the endpoint's OTHER failures keep the plain flat shape). Read such acodefrom the response body's top level, not from the message string: theerrorprose is the operator/log text, it is English only, it names internal fields, and it may change between releases. A client that wants to show the user something must key its own localised message off thecode.A third shape exists for the typed policy refusals in this table (
share_restricted_project,project_release_forbidden,share_source_project_forbidden,project_tier_unresolved,project_access_unresolved,agent_area_unresolved, …): those put the machine code inerrorand the prose in a siblingmessage—{"error": "share_restricted_project", "message": "This conversation is in a project…"}. The rule is unchanged and is why the shape exists: branch on the machine string, render your own localised text, never parse the prose.
| Code | HTTP status | Description |
|---|---|---|
unauthorized |
401 | No auth token, or the token is unknown for the gateway named in the path. Tokens are gateway-scoped: a valid token used on the wrong gateway also returns 401. The message distinguishes the cases — missing token, a token scoped to a different gateway of the same tenant, or an unknown/other-tenant token. An expired token has its own code, token_expired; an explicitly revoked one has token_revoked. A token that disappeared together with its agent, user or gateway (a cascade, not an explicit revocation) answers this generic code. See Token security model. |
token_revoked |
401 | The token was valid and was revoked — deleted under Gateways → Tokens, Profile → My tokens, by a chat-bridge teardown, or because the token's owner was disabled, deleted, or SCIM-deprovisioned (ending a user's access revokes their /v1 tokens in the same step). Distinct from unauthorized because nothing about the URL or the gateway is wrong, and distinct from token_expired because re-minting is not something the caller can do alone: an administrator (or the token's owner) has to mint a new token and update the caller. The message names the revocation date (UTC) when it is known. The decision comes from a revocation tombstone written in the same transaction as the delete, read for the gateway named in the path only — a token revoked on another gateway or in another tenant is reported as the generic unauthorized, never as revoked. A client must not retry on this code. |
token_expired |
401 | The token was valid and has expired. Distinct from unauthorized because it is a different situation with a different remedy: the credential was real, nothing is misconfigured, and the call succeeds again once a fresh token is used. The message names the expiry and points to Gateways → Tokens, so a scheduled agent or client stuck on a stale token gets an actionable next step. A client may safely re-mint and retry on this code, which it must not do for unauthorized. The chat app does exactly that: when a turn's request is refused with this code before any byte streamed (the short-lived play token lapsed between the pre-turn check and the request), it re-mints once and replays the same turn — safe because authentication runs before the turn is claimed or anything is persisted. When it cannot re-mint (the browser session itself expired) it shows a typed "session expired" state in the conversation with a sign-in action, instead of leaving the page mid-turn; a refusal that is not an expiry (unauthorized, or a refused mint) is shown as a typed refusal and is never retried. The accepted 401 body shape is { "error": { "code": string, "message": string } }; a body whose code is absent or not a string is treated as unauthorized (no retry). |
forbidden |
403 | The token is valid but does not have permission. Causes: viewer role on inference, or IP not in the ip_allowlist of the gateway. Admin message writes (POST/PATCH/DELETE /admin/v1/conversations/{id}/messages…) return the same code, in the admin { "error": <prose>, "code": "forbidden" } shape, when the caller can read the conversation (a project member it is shared with) but may not write to it right now — off-session only the owner writes. A refusal the app shows as a notice, not an error. See Messages. |
impersonation_forbidden |
403 | An administrator who is viewing the app as a user (see Admin impersonation) tried to mint a durable credential or create a persistent out-of-band effect — the short-lived chat token, a personal or gateway token, a share link, a webhook trigger. Refused so nothing survives Stop. Shape: { "error": "impersonation_forbidden", "code": "impersonation_forbidden" } — the machine code is BOTH the error value and the sibling code (added the sibling); branch on it, never render it. The chat app localises typed codes (code, or error.code), so the banner carries "Chat is unavailable while you are viewing the app as this user…" (the copy already ships; the "Could not create gateway token:" prefix stays until the banner's own follow-up). Stop viewing as the user to continue. |
viewer_read_only |
403 | The caller holds the viewer role and tried a chat write that spends inference credit or reshapes a conversation's content/context on the admin session plane — creating a conversation (POST /admin/v1/conversations), forking one (.../fork), a message POST/PATCH/DELETE, an attachment upload, a chat-file upload (POST /admin/v1/chat/files), summarize, summaries, context-drop, or a knowledge/<KID> PUT. Emitted in the admin { "error": <prose>, "code": "viewer_read_only" } shape and decided before any conversation lookup, so it is not an existence oracle. Distinct from forbidden (which the app reads as "you can read this conversation but not write to it right now"). It is not a blanket write ban: a viewer may still read and do owner-housekeeping (rename, delete, rate, delete an attachment), plus share links, comments and read-aloud. The /v1 inference twin is forbidden (see above). See Viewers cannot spend credit or reshape a conversation. |
processing_restricted |
403 | The user is under a GDPR Art. 18 processing restriction (a reversible admin-set legal hold). No model is called. An admin lifts it via PATCH /admin/v1/users/{id} with {"processing_restricted": false}. |
data_residency_blocked |
403 | The gateway requires EU model hosting (eu_region_routing: true) but the selected provider/region is not EU-hosted. The request is refused before any upstream call — it never reaches a US endpoint. Route the model via a native EU provider (Myra, Mistral) or an EU adapter (Bedrock eu-*, Vertex europe-*, or an EU-region Azure resource). See Data residency. |
provider_not_allowed |
403 | The gateway enforces an explicit provider allowlist (provider_allowlist_enforced: true, or a deployment-wide enforcement setting configured by your operator) and the selected model resolves to a provider that is not on provider_allowlist. The request is refused before any upstream call — no attempt, primary or fallback, can silently reach a non-approved provider. Pick a model served by an approved provider, or add the provider to the allowlist. Fails closed: under enforcement a missing/empty/malformed allowlist denies every provider. See Provider allowlist. |
role_model_not_allowed |
403 | The gateway defines a per-role model mask (role_policies.<role>.models) and the request's resolved model is not in the user role's allowance. Enforced on the primary (user-chosen) model at the guardrail stage, and re-asserted on the primary turn's upstream retry/fallback swap targets before any network call — so a system-initiated swap can never carry the user's turn onto a model the role excludes. A deterministic policy refusal (not a provider outage; no retry, no model-error triage row). Pick a model in your role's allowance. See Per-role policy. |
model_disabled_on_gateway |
403 | A tenant admin disabled this Myra-hosted local model on the gateway (Settings › Gateways › Add Model › Use Platform Key › Myra models). A hard block: the model is neither offered in any model picker nor routable on that gateway — enforced per dispatch attempt (primary and every fallback/failover re-target), at agent/project/scheduled-task save time for new pins (those save-time rejections return 400 with this same code), and excluded from the attachment auto-upgrade and the 413 switch-model suggestion. A deterministic policy refusal (not a provider outage; no retry, no model-error triage row). A pre-existing pin to a since-disabled model returns this code until re-pinned. Re-enable the model in the same dialog, or pick a different one. |
provider_objected |
403 | The selected model resolves to a provider whose sub-processor this tenant has objected to (Art. 28 Abs. 6 DSGVO — a per-tenant deactivation recorded by Myra Ops via the subprocessor-objection API). A hard block scoped to that one tenant: the provider is neither offered in any model picker nor routable for that tenant — enforced per dispatch attempt (primary and every fallback/failover re-target, plus the observed LiteLLM-overflow leg), at agent/project/scheduled-task save time for new pins, and excluded from the model-offer set. A deterministic policy refusal (not a provider outage; no retry, no model-error triage row). No other tenant is affected. Pick a model served by a different provider, or ask Myra to withdraw the objection. |
provider_not_configured |
400 | Returned by a save-time validation (pinning a project default model, pinning an agent's model, or the /compact conversation-summary call) when the pinned/summary (provider, model) resolves to a keyless provider this gateway never opted into — no provider_config row and no per-gateway base-URL override for it. Refused at save so the pin cannot be accepted and then block every conversation ("offer-then-block"). It is the config-time twin of the runtime invalid_request case below, decided by the same opt-in predicate, and — like provider_not_allowed — is a caller-fixable routing error, not a provider outage (no retry, no operations alert). Pick a provider this gateway serves, or opt the provider in on the gateway. (A provider that requires an API key returns provider_key_missing (424) instead.) |
live_not_driver |
403 | The conversation has a live shared session and the caller is not the current driver — only the driver may write. Request the turn via POST /admin/v1/conversations/{id}/live/driver. On the /v1 turn path it is also returned (fail-closed) when the session state cannot be resolved. The three admin message writes (POST/PATCH content/DELETE /admin/v1/conversations/{id}/messages…) return it in the admin { "error": <prose>, "code": "live_not_driver" } shape — there a session-read fault is a retryable 500, not this code, so a database blip never reads as "someone else is driving". See Live shared session and Messages. |
model_not_found |
400 | The requested model cannot be served, from one of two origins. (a) Pre-routing (a client error): a raw OpenAI-compatible (/compat) request names a model id the gateway cannot place on any provider it serves — the id is not a catalogued model, matches no known provider prefix, and the gateway does not configure openrouter (which passes an arbitrary id through). Refused at the routing boundary before any upstream call — not a provider outage, no retry, no operations alert. Check the model id, or pick a model this gateway serves. Note: a gateway that does configure openrouter still passes an uncatalogued id through to openrouter (its purpose as an aggregator), so this origin is returned only where there is genuinely no provider to route to; an unknown id is never silently attributed to openrouter when openrouter is not configured. (b) Post-upstream (a serving-side registration gap): the gateway's own model proxy accepted the request but reports the routed model group as an unknown alias (the group is not registered on the proxy). Rather than dumping the raw upstream NotFoundError string, the gateway surfaces the same typed model_not_found and, for a chat turn, the "Continue with <default>?" recovery. Unlike origin (a), this happens after an upstream attempt and does leave a model-error triage row (currently categorized user_fault, still notified to ops). Retrying the same model will not help until the group is registered. (c) Unpriced self-hosted chat model (a config gap): a Myra self-hosted chat model with no usable internal price is refused at dispatch with this same model_not_found (deliberately indistinguishable on the wire from a deprecated/absent model) so it is never served at an implicit $0. Deterministic (no retry) and not paged (no triage row) — a platform admin setting a real price makes it usable. Non-chat fleet models (embeddings/rerank/image/speech) are priced at zero on purpose and are not affected. |
model_not_allowed_for_project |
403 | The conversation is in a project restricted to local models only, and the requested model does not resolve to a local model. The request is refused before any upstream call. Pick a local model to continue. Also returned by the conversation-route PATCH (see Route policy validation) when a pin would move a local_only conversation onto a non-local model — refused at save so the pin cannot brick the conversation. And by the conversation-summary (compaction) call, which is generated from the conversation's own text and so is refused for a local_only project on a non-local model before anything is sent. |
project_tier_unresolved |
503 | The conversation's project binding could not be confirmed for this request, so its data-residency tier is unknown and the turn is refused (fail-closed) rather than risk egressing a restricted conversation. Retryable — this is a transient state (the binding resolves on retry), NOT the permanent model_not_allowed_for_project; the message avoids the false "only local models" claim for a project that may not be local at all. See How the egress tier is resolved. Also returned by create-from-share (POST /admin/v1/conversations with source_share_token) when the source conversation's binding cannot be read transiently — the continuation is refused rather than created unbound, and retryably, because the source's project may not be restricted at all. See Continuing a shared conversation. |
share_source_project_forbidden |
403 | Continuing a shared conversation (POST /admin/v1/conversations with source_share_token) whose source lives in a project the caller is not authorized on and whose access tier is restrictive (Local only or PII mandatory) — or whose binding cannot be confirmed at all (the project belongs to another workspace, was deleted, or its tier is unrecognized). Copying a restricted transcript into an unbound conversation would strip the residency guarantee that project exists to provide, so the continuation is refused instead. Permanent — unlike project_tier_unresolved there is nothing to retry; ask the project's owner for access. "Not authorized" and "could not confirm" deliberately share this one code, and the message names no project, workspace, tier or owner, so the response is not an oracle over a stranger's project. See Continuing a shared conversation. |
guardrail_unavailable |
503 | A request-phase safety guardrail is DOWN and the gateway is configured to refuse rather than skip it, so the request was rejected instead of being sent unscanned. Retryable — nothing about the content was wrong, which is why it is deliberately not guardrail_blocked (400, permanent). On a chat turn this same condition is delivered as a "temporarily unavailable" policy bubble rather than a typed error; the typed code is what non-chat callers (currently the conversation-summary call) receive. |
guardrail_doc_too_large |
503 | A PII scan could not finish inside its wall-clock budget because the document was too large for the sidecar's current throughput, so the request was refused (fail-closed) rather than sent unscanned. Retryable — and it genuinely converges: the chunks already scanned are cached, so a re-upload measures faster and fits. Deliberately distinct from guardrail_unavailable (a generic transient outage) and provider_quota_exhausted (a provider billing 503): the honest guidance is "too large right now, try again", not "temporarily unavailable". On a chat turn it is delivered as the guardrail_doc_too_large policy bubble with a Try again action; the typed code is what non-chat callers (currently the conversation-summary/compact route) receive. See Large messages: chunked analysis and the scan budget. |
project_access_unresolved |
503 | The authoritative check of your access to a conversation's project failed transiently while continuing a shared conversation. Retryable, and deliberately distinct from a denial: a database fault is not "you are not a member", and refusing on it would hand out a permanent-sounding refusal for a project you may well belong to. Also distinct from project_tier_unresolved, which is the same class of fault on the project's tier read — keeping them apart says which of the two authorities is degraded. See Continuing a shared conversation. |
share_restricted_project |
403 | Minting a public share link (POST /admin/v1/conversations/{id}/share) for a conversation whose project restricts data residency (Local only or PII mandatory), or whose project binding cannot be confirmed. A share link is a persistent, unauthenticated read credential, so one click would publish the transcript of a project that exists to keep its data inside the estate. Permanent — the remedy is to move the conversation out of the project first (which requires owner rank on it and is audited), or to ask the project's owner. Also returned by GET .../share, so an owner is never handed a live-looking URL for a link the public read refuses. Addressed to the owner, who has a remedy — unlike the cause-neutral share_source_project_forbidden a public recipient receives. |
project_release_forbidden |
403 | Releasing a conversation from a project that restricts data residency — detaching it (PATCH …/conversations/{id} with project_id: null) or moving it to another project — without owner rank on that project. Attaching is membership-gated because the conversation would gain the project's knowledge; releasing is the direction that drops a residency guarantee, so it is gated on the rank that could lower the project's tier or delete the project anyway. Also returned, with no owner exception (platform admins aside), when the current project cannot be confirmed at all — nobody holds owner rank on a project that no longer resolves. Permanent; the committed binding is unchanged. |
agent_area_unresolved |
403 | A saved agent is bound to a knowledge area that no longer exists (it was deleted after the agent was configured). The gateway cannot tell what data-residency policy that area carried, so the run is forced to the most restrictive setting on both axes — local models only, PII masking on — and a cloud-pinned agent therefore cannot dispatch. Permanent, not retried: unlike project_tier_unresolved the area is gone for good, so a scheduled agent fails every run until it is fixed; and unlike model_not_allowed_for_project it does not claim a project that no longer exists or offer a remedy ("pick a local model") the owner of an unattended run cannot perform. Open the agent and save its knowledge areas. See Knowledge areas. |
conversation_project_unresolved |
403 | The conversation belongs to a project that no longer exists, so its data-residency tier cannot be read and the turn is restricted to local models. Unlike the retryable project_tier_unresolved (a transient read fault), the binding is gone for good: pick a local model to continue here, or start a new conversation. |
gateway_not_allowed_for_project |
403 | Returned by the conversation-route PATCH (see Route policy validation) when a pin's resulting gateway cannot satisfy the conversation's PII policy — a pii_mandatory project, or a tenant with mandatory PII masking, pinned to a gateway that does not scrub request-phase PII — or when the gateway_id is unknown (any caller), or belongs to another tenant, or names a gateway that does not route (purpose other than production); the last two apply to non-admin callers only (platform admins may route cross-tenant and may pin a non-production gateway deliberately for triage). The routability refusal fires only on a MOVE — a request naming a DIFFERENT gateway than the one the conversation already sits on — so an unrelated edit, or a model switch that re-sends the committed gateway, is never refused and an existing conversation cannot be bricked. Refused at save so a picker click cannot brick the conversation for every subsequent message. Platform admins may route through another tenant's gateway. Also returned by create-from-share (POST /admin/v1/conversations with source_share_token), which validates the supplied gateway_id through the same fence before writing anything. |
plan_model_not_allowed |
403 | The workspace is on a self-serve plan and the requested model is not included in that plan's model list. The request is refused before any upstream call. Pick a model included in your subscription. Fails closed: a misconfigured (empty) plan model list denies every model. On save the same code is returned with 400 when an agent, a project's gateway + model pin, a scheduled task or a gateway's agentic_fetch.model names a model outside the plan (judged on the retired-id successor); a model saved earlier that the plan no longer permits runs on the gateway's Auto choice instead (stored pins only — see Tenant entitlements). Also returned (fail-closed) in the over-cap degraded state when the degrade target is unavailable or a routing override would escape the efficient set. See Billing. |
eu_gov_model_not_allowed |
403 | The workspace is on a paid self-serve plan and requested a Myra-hosted (non-commercial, provider-class myra) model without the EU-Gov add-on entitlement (myra_hosted_models). Commercial providers (Claude, OpenAI, …) are always allowed. The request is refused before any upstream call — pick a commercial model, or have the EU-Gov add-on granted. The 7-day free trial (self_serve_trial) is exempt — its two Myra EU models are entitled by plan design, so a trial never sees this code for them. Fails closed: an unresolvable/absent grant, and an unresolved provider, both deny; the over-cap degrade target is checked the same way (never a silent Myra swap). Grant source: a tenant entitlement row for the eu_gov add-on. On save: 400 with the same code for a stored pin or agentic_fetch.model naming a Myra-hosted model without the grant. Platform infrastructure services (OCR, vision, embeddings, guards, speech, image generation, the default fetch inner model) are exempt. See Billing. |
subscription_inactive |
402 | The workspace is on a self-serve plan whose subscription is inactive (payment overdue past its grace period, cancelled, or under dispute). Inference is blocked; the account and billing pages stay reachable so the subscription can be reactivated. During the grace period after a failed payment, inference still works. See Billing. |
trial_expired |
402 | A trial workspace whose trial phase has ended (trial_ends_at is in the past). Inference is blocked; the account/admin pages stay reachable to upgrade (self-serve) or for a platform admin to extend the trial (admin-provisioned demo). For a self-serve free trial, a companion trial_budget_exhausted (402) covers the other end condition; for an admin-provisioned demo the spend cap surfaces as quota_exceeded (429). See Billing / Tenants. |
trial_budget_exhausted |
402 | A self-serve free-trial workspace that has spent its whole trial credit before the 7 days elapse. Distinct from trial_expired (time) so the SPA shows a "credit used up" upgrade screen; inference is blocked while the account/billing pages stay reachable to upgrade. Not the generic quota_exceeded (429) — a trial has no subscription to renew against. See Billing. |
workflows_disabled |
403 | The workflows feature is not enabled for the workspace. The workflow builder and workflow runs are unavailable until an admin enables the feature for the tenant. See Workflows. |
agents_disabled |
403 | The Agents feature is not enabled for the workspace. A human agent invoke (interactive /agents/{slug}/invoke, the run-as-self preview, the in-chat agent bridge) is refused until an admin enables Agents for the tenant (agents_enabled). The admin /agents routes return the plain feature_disabled form of this block. See Agents. |
feature_disabled |
403 | An admin route for a per-tenant feature (Workflows, Agents, Playground, or Scheduled tasks) was called while that feature is disabled for the workspace. The response names the feature: { "error": "feature_disabled", "feature": "<name>" } where <name> is one of playground, agents, workflows, or scheduled_tasks (so the cause is actionable without checking each flag by hand). The user-facing inference and invoke paths return the feature-specific code instead (workflows_disabled, agents_disabled); the plain feature_disabled (now feature-tagged) form is what the feature's /admin/v1/... routes return. The internal email-ingest path (/internal/workflow/ingest-email) likewise tags its { ok: false, error: "feature_disabled" } result with "feature": "workflows". Enable the feature for the tenant. |
tenant_not_found |
404 | The tenant or gateway slug in the URL does not exist. |
endpoint_not_found |
404 | The URL names a real provider but a path that is not a servable endpoint on it. Only a fixed allowlist of whole paths is accepted — /chat/completions, /completions, /embeddings, /v1/messages, /v1/messages/count_tokens, and /v1/models (an optional /v1 prefix allowed). The entire path is matched, not a trailing suffix, so a hostile prefix (/internal-admin/embeddings) is refused too. For an explicit provider the check runs in the request's access phase before the token is read, so an invalid token on an unknown endpoint still returns this 404 rather than 401; for the /compat surface the same allowlist is re-checked after the provider is resolved from the model. This is an authorization boundary: each provider forwards the path verbatim into its upstream URL, so an unvalidated path would let a caller aim a request (carrying the gateway's or the platform's upstream credential) at an arbitrary upstream path. A missing/bare path (a bare …/{provider} URL, an empty suffix, or a bare /v1) is refused too — append an explicit endpoint such as /chat/completions; an absent path fails closed exactly like a present, unrecognised one (they share the reject outcome, never a permissive dial). See Only servable endpoints are accepted. |
conversation_not_found |
404 | An admin conversation route (GET /admin/v1/conversations/{id}, POST/PATCH/DELETE …/messages…) was called for a conversation the caller cannot see at all: unknown id, deleted (in another tab, on another device, or by retention), or never shared with them. One answer for all three — no existence oracle. Never returned for a conversation the caller can still read (those get 403 live_not_driver / forbidden on a write), so the app treats it as "this conversation no longer exists": it clears the thread, drops it from the list, shows a non-error notice and returns to the start page, keeping any draft. The error prose differs per route ("not_found" on the GET, "conversation not found" on the message writes); clients key on the code. See Getting a conversation and Messages. |
row_too_large |
413 | A message write (PATCH …/messages/{id} with content or corrected_content, the legacy POST …/messages commit) would make the message row larger than the database can read back in one packet — its large text columns (content, corrected_content, and a web-search turn's sources / search-steps lists) may total ~15.9 MB. The row is unchanged; shorten the text. A deterministic refusal of the caller's input, never a server fault: the app shows a notice and does not report a client error. See Messages. |
attachment_too_large |
413 | A conversation attachment upload (POST /admin/v1/conversations/{id}/attachments) whose inline stored row would exceed the single-packet database limit (16 MB) — the same ceiling as row_too_large, classified distinctly as a stored attachment. Nothing is persisted; use a smaller file. This is the handler refusal for a body that fit under the 16 MB admin body cap; a larger body is refused earlier by the edge as payload_too_large (below), so attachment_too_large covers the narrow band between the two limits. It replaced a raw 500; larger attachments are tracked internally. See Attachments. |
invalid_request |
400 | Malformed request body, missing required fields, or an unrecognised parameter value. Also returned when the request names a provider the gateway has not configured — e.g. a provider/ model prefix that resolves to a keyless provider this gateway did not opt into (no provider key / provider config for it). This is a client error, refused at the routing boundary before any upstream call — distinct from a provider the gateway did configure being down (all_providers_failed, 502). The equivalent rejection at save time (pinning that provider as a default, or summarizing onto it) is reported as provider_not_configured (400). |
method_not_allowed |
405 | The HTTP method is not accepted by this endpoint (for example a GET on a POST-only route). The Allow response header lists the permitted methods. |
rate_limited |
429 | Two origins share this code. The gateway's own sliding-window limit (per token or per gateway): the response carries X-RateLimit-Limit, X-RateLimit-Remaining and Retry-After; a token-scoped limit says Token rate limit: … in the message. A provider's final 429 relayed by the inference terminal (the last attempt answered 429 — an earlier 429 followed by a 5xx is not this): the response carries the provider's Retry-After when it sent one, the Anthropic-dialect window headers as X-RateLimit-Input-Tokens-Remaining / -Reset and X-RateLimit-Requests-Remaining / -Reset when present, and X-AIG-Rate-Limited: true (the gateway's own X-RateLimit-Limit / X-RateLimit-Remaining are also present when the gateway or token has a limit configured — they were set on the way in and describe that window, not the provider's). On a PII gateway whose stream is already open the terminal 429 arrives as an SSE provider_error frame with error_class: rate_limited and no headers. The chat UI renders one localized message for either — and for a 429 that reaches it without a gateway code at all (a non-JSON edge body, a provider body carrying only the provider's rate_limit_exceeded); when the response carries a positive Retry-After (delta-seconds or an HTTP-date), that message names the seconds to wait — a missing, zero, negative or malformed header is ignored (the plain message, never a "0 s"). The tool-loop / MCP path keeps its own upstream_rate_limited. |
spend_unverified |
503 | A Myra-provided (keyless) model request could not verify the workspace's spend against its budget (an authoritative spend read failed). The request is refused rather than routed on Myra's managed key with an unverifiable balance. Retryable — the balance is unknown, not known-exceeded. See Myra-provided models. |
scan_unavailable |
503 | An upload that must be virus-scanned could not be: the antivirus service was unreachable, timed out or errored — or whether scanning applies could not be determined (the tenant's gateway list was unreadable or a gateway config would not decode). The gateway fails closed rather than store an unscanned file. Carries Retry-After; retryable once the scanner / the config read recovers. See Upload malware scanning. |
quota_exceeded |
429 | The configured spend budget has been exhausted (token-level, tenant-level, or gateway-level). The response includes a Retry-After header (whole seconds) set to the time until the budget's period resets — e.g. seconds to the start of the next monthly/daily period, and (for self-serve workspaces) the seconds until the monthly allowance resets (~30 days after its last reset, on monthly and annual plans alike — see the two clocks in Billing). The header is omitted rather than set to 0 (which a client would read as "retry immediately") whenever no reset is scheduled: a lifetime (total) cap never resets (a retry can never clear it — the error message says so), and a self-serve subscription in grace, inactive, or cancelling at its period end with no allowance reset before that has no reset either. For self-serve workspaces with cap soft-degrade configured, this fires only past the secondary (extended) cap — at 100% of the base allowance the request instead continues on the plan's efficient models; without soft-degrade it fires at 100%. See Budgets. |
guardrail_blocked |
400 | A guardrail (regex, keyword, NLP PII Detector, Prompt Guard, PII Protector) matched with a block action. The message names the blocking guardrail and pattern/category. |
provider_error |
502 | The upstream provider returned a 5xx error and all retries were exhausted for that provider. Also returned for a mid-stream provider failure on a buffered/PII-forced tool-loop turn (the provider's stream emitted an error event or died before completing) when a transparent retry would be unsafe — server-side tools already executed — so the turn fails loud with this typed code instead of an HTTP 200 carrying an empty message. (When the retry IS safe and every attempt then fails, the terminal code is all_providers_failed, as for any exhausted routing chain.) On an already-open SSE stream the code arrives as an error frame plus an aig_status: "provider_error" event. A salvaged partial answer is delivered together with the explicit error surface of its failure kind — for a provider error event that same frame + banner pair; for a read error or a truncated stream (X-AIG-Error: stream_errored / truncated on a non-streaming response) the single terminator error frame (provider_error "connection failed" / stream_truncated) and the persisted interrupted note at the end of the visible answer, exactly as on the live streaming path (see Inference › Visible truncation note). |
all_providers_failed |
502 | All providers in the routing chain (primary + all fallbacks) returned errors or timed out. This is a genuine upstream fault for configured providers (it pages on-call). A request that merely names a provider the gateway never configured is instead an invalid_request (400, no page), classified before dispatch.The message depends on who owns the failing provider. For a provider you configured, it carries the raw upstream cause, which is the actionable detail for a key or endpoint you control. For a Myra-hosted model the cause is ours, so the message says the model is temporarily unavailable and suggests retrying or choosing another — the underlying cause is recorded on our side (request log, trace) rather than returned to you. Retry either case: this code is transient by definition.The gateway absorbs the transient class on the Myra-hosted leg before it reaches you: a pre-first-byte connection blip (the bursty upstream wobble) is retried on the same model with a small, bounded, jittered back-off — even on a gateway whose general retry budget is zero and on a Myra fallback leg — so an isolated blip that then succeeds is never surfaced. To keep the shared Myra circuit breaker from tripping on a wobble it recovers from, an intermediate transient attempt that will be retried does not count against the breaker; only the final, exhausted attempt does — so a genuinely sustained outage still opens it. This applies only to the transient pre-first-byte connection class: a provider 4xx/5xx, a read timeout, and any failure after streaming has begun are never retried this way (and never past the first byte), and a persistent transient failure still surfaces this code with the honest "temporarily unavailable" Myra message above. |
provider_quota_exhausted |
503 | The upstream provider's own account reached its usage or billing/credit limit — e.g. a serverless overflow group whose provider account has no balance (Together.ai), or an exhausted provider key. The gateway surfaces this clean, typed degraded error instead of relaying the provider's raw billing message, and triages it provider_fault (never user_fault) so operations sees the account needs attention. Switching to a different model/provider is the user remedy. Distinct from quota_exceeded (429), which is the workspace's own budget running out. Also returned before any upstream call on a self-serve plan when the managed Anthropic key pool is exhausted, its organisation suspended, or the pool key is not populated (the stored credential row is present but blank — e.g. after a sanitised environment re-seed) and the tier names no usable cross-vendor fallback: the provider account really is the thing that is unavailable, the condition is transient, and the remedy is the same (retry shortly, or pick another model). |
ask_docs_unconfigured |
503 | The public documentation assistant (Ask the docs) is unavailable because its backing gateway is not configured. The widget shows a try-again message rather than an error. See Ask the docs. |
request_timeout |
504 | The request took too long to complete. Try a simpler request, narrow the scope, or turn off web search. |
request_too_large |
413 | The request could not be processed by any provider. Large attachments are a common cause. Remove attachments or start a new chat. |
context_overflow |
413 | The estimated input exceeds the selected model's context window. The gateway's pre-flight gate now auto-compacts an over-window turn (summarizes older history, keeps recent + current, proceeds) instead of refusing, so this 413 fires only for a single message that alone exceeds the window even after clipping (or re-emitted after a provider rejects a model whose window the catalog does not know). The body carries estimated, limit, first_turn (exactly one user turn — the message then says the input is too large rather than suggesting a new chat), can_compact_now (its negation), and a suggested_model that is only ever a model the caller can route to, so acting on it never produces a model_not_found. Distinct from request_too_large (a terminal re-label after every provider genuinely failed) and context_length_exceeded (a provider-side 400). See context_overflow — the pre-flight context gate. |
context_length_exceeded |
400 | The prompt exceeds the context limit of the model. Shorten the prompt or remove attachments. |
structured_output_failed |
422 | An agent with a response_schema produced an answer that could not be extracted, parsed, or validated against its schema. The gateway fails closed rather than deliver non-conformant output as success. See Structured output. |
agent_not_found |
404 | The agent id in the URL does not exist, is not visible to the caller, or has been deleted. |
agent_misconfigured |
422 | The agent's configuration is incomplete (for example no resolvable model) and it cannot run. Fix the agent configuration and retry. |
malware_detected |
422 | An uploaded file (a chat file or conversation attachment, a project knowledge file, a personal My-PII keyword list, or an app-feedback evidence image) was rejected because the antivirus scanner flagged it as malware. The file is not stored; the signature is logged operator-side only. (The anonymous workflow-form upload reports the same token in its flat error field rather than as a code.) See Upload malware scanning. |
invalid_body |
400 | The request body's required text field is not a non-empty string (absent, null, an object, a boolean, a number, or ""): data on the conversation-attachment upload (POST /admin/v1/conversations/{id}/attachments) and the My-PII keyword-list upload (POST /admin/v1/me/pii-keywords/upload), text on the My-PII paste (POST /admin/v1/me/pii-keywords/batch). Answered before anything is read or stored. See Attachments. |
bad_base64 |
400 | An upload body's data string does not decode as base64 (standard alphabet, no line breaks — an embedded newline is enough). Returned by the conversation-attachment upload and the My-PII keyword-list upload. Nothing is stored; the body is never passed through unscanned as opaque text. |
pdf_password_protected |
422 | A chat PDF attachment is password-protected (encrypted with a user password) and cannot be read without the password. Remove the password, or attach an unlocked copy. Only a real user-password PDF is rejected — an owner-password-only PDF (which opens without a password) is accepted and read normally. The same password cause is surfaced on the project-upload path (a 422 "corrupt or password-protected") and, for a connector-synced PDF, in the document's ingest trace (the async worker terminal-fails it with a password-protected reason rather than a generic extraction failure). |
ingest_busy |
503 | The shared upload-processing pool (chat attachments and project knowledge uploads draw on one bounded slot budget) is momentarily saturated. Transient — the response carries Retry-After; retry the same upload shortly. The web chat shows a "busy — try sending again" banner and keeps the attachment in the composer (in a multi-attachment send the other files' finished extractions are kept and only the failed file is re-submitted on the next Send); the project knowledge upload retries a bounded number of times. |
agent_owner_inactive |
403 | The agent's owner account is deactivated or removed. An agent runs as its owner, so it can no longer execute — this includes scheduled and webhook-triggered runs. |
web_search_not_supported |
400 | The selected model rejected web search. Disable web search for the chat or pick a different model. |
model_capability_mismatch |
400 | The selected model cannot run this request. Returned when the model rejects a gateway-injected capability (tools / web search / URL fetch / vision), or when a native document/PDF block is sent to a model that cannot read documents (only Anthropic, and Bedrock for anthropic-family models, accept native document blocks). The gateway returns this typed error instead of relaying the raw provider status. Pick a tool-capable model, or disable web search / URL fetch and remove attachments. Note: for a model the gateway KNOWS cannot do function calling (e.g. a Perplexity Sonar model), the gateway now PROACTIVELY auto-disables the injected tools and the run PROCEEDS without them (a tools_skipped notice is surfaced — see below) instead of failing. This typed error is now only the residual safety net for a tool-incapable model the gateway did not recognize up front — and it is the last resort rather than the first: when the backend refuses the tools the gateway injected (web search, URL fetch, file access, connectors), the gateway strips its own injected set and re-sends the turn once as a plain answer, so the user gets a served reply plus a tools_skipped notice instead of a dead end. model_capability_mismatch is returned only when that stripped retry also fails, or when the refusal is about a capability stripping cannot fix (an image, a document). A caller that supplied its own tools array is never touched — its provider's error is passed through verbatim. |
provider_key_missing |
424 | The selected model routes to a provider that requires an API key, but no key is stored for this gateway (and alias). This is a configuration gap, not a gateway fault — an admin must add a provider key in the gateway's provider settings. Returned before any upstream call. |
managed_model_not_enabled |
424 | The selected model is a Myra-provided (Platform Key) model this gateway has not enabled (no managed grant, or granted without the required workspace budget), and no own provider key covers it on the default alias either. Distinct from provider_key_missing so the actionable hint is the right one: enable the model under Add Model → Use Platform Key (a workspace budget must be set), or add your own provider key. Returned before any upstream call. (A request that names a non-default x-aig-byok-alias keeps provider_key_missing — that alias signals an own-key intent.) |
connector_not_found |
404 | A tools:[{type:"mcp"}] connector reference names a connector that does not exist or is not accessible with this token (another tenant's, or another member's private connector). Fail-closed: the request is aborted before any model call. See MCP connectors from the API. |
connector_credential_required |
424 | The referenced MCP connector requires a per-user credential that the API token's user has not connected (or it expired — the detail names which). Connect or reconnect the connector in the UI, and call with an API token created by that user. Also returned by PUT /admin/v1/mcp/{id}/credential when the pasted token is rejected by the provider at save time, and it is the code a per-user token's runtime provider rejection (a 401/403 on the connector handshake) now maps to — previously a misfiled 502. The rejection also flips the connector to "Reconnect needed". |
connector_upstream |
502 | The referenced MCP connector's server did not answer tool discovery (unreachable, protocol error, or the ~20s discovery deadline was exceeded). A caller-side connector fault — check the connector's server_url and server health. Also returned by PUT /admin/v1/mcp/{id}/credential when the connector could not be reached to verify a pasted credential — a connectivity problem, not a bad credential; nothing is stored and the caller may retry. |
configuration_error |
500 | The gateway configuration could not be loaded or parsed (malformed config, or a route that resolves to a disallowed/unsafe upstream). Distinct from provider_key_missing, which is a missing credential (424). Check the gateway configuration and error log. |
internal_error |
500 | An unexpected error occurred inside the gateway. Check the gateway error log for details. |
service_unavailable |
503 | A transient backing-store or dependency read failed on a non-inference lane (for example the document-AI Copilot action or the Entra sign-in exchange), so the request was refused rather than served from an unknown state. Retryable. |
service_starting |
503 | A transient, retryable state during a deploy/restart window: the worker has started but its schema migrations have not finished applying yet, so it refuses DB traffic until the schema matches the code (rather than returning an inconsistent per-column error). The response includes a Retry-After header (seconds) — retry after that delay and the request succeeds once migrations complete. /healthz (liveness) stays 200 throughout. If it persists well beyond a normal restart, a boot migration has failed (fail-closed) — check the gateway error log. A separate /readyz endpoint (JSON) reports true readiness: it returns 200 only when the schema is ready and no enabled host-runtime dependency is broken, and 503 + Retry-After otherwise, with a body naming each capability's status (e.g. the code-interpreter sandbox's mTLS/HMAC and the blob store's writability). A boot-time preflight validates those host deps once per (re)start and logs CRIT naming any broken one, so a provisioning miss (a bad cert file mode, a non-writable bind-mount, an empty session secret) is loud and visible at /readyz instead of silently failing closed per request. Repoint a load balancer at /readyz to gate routing on the deps actually being usable. While the schema is not ready the Retry-After is the same value service_starting carries; while the preflight has not published yet (the first moments after boot) or its published status is unreadable, /readyz also answers 503 (preflight_published: false); a 503 caused only by a broken capability carries Retry-After: 5 and persists until the host dependency is fixed and the backend is restarted (the preflight runs once per boot). A third probe, /readyz/schema (JSON, API host), answers readiness of the schema only: 200 {ready: true, schema_ready: true, scope: "schema"} once the boot migrations have finished and the serving gate is open, 503 + the service_starting Retry-After until then — and indefinitely if a boot migration failed. It is the backend container healthcheck in the shipped Docker Compose files (the container reports starting/unhealthy instead of a green /healthz over a gateway that refuses every request); a broken host capability deliberately does not mark the container unhealthy. All three probes take no input (any method answers the same; a query string or body is ignored), and read only in-memory boot state — they never touch the database, so they answer during the very window they report on. |
Notes on specific codes
context_overflow — the pre-flight context gate
Before calling the upstream provider the gateway estimates the request's input token count
(including the server-composed system prompt the client never sees) and, when it exceeds the
selected model's advertised context window, it now automatically recovers rather than refusing:
the gateway summarizes the older conversation history into a compact preamble (reusing the same
bounded summarizer the /summarize compaction route uses — each message is head+tail-clipped and
the summary input itself can never re-overflow), keeps the most recent turns and the current
message, re-measures, and proceeds with the turn. The user gets a real answer, not a wall; a
subtle aig_status: "compacted" status line is surfaced. The original transcript is untouched —
only the model-visible context is compacted. This is always-on for every tenant and gateway: it
is a reliability floor, independent of the context_compaction setting (which governs only
Anthropic's native long-context compaction, a separate sub-window mechanism). A single oversized
message is first head+tail-clipped to fit; an inline data-file preview (a user text block of the
gateway's [File: <name>] format carrying a spreadsheet extract — see the inference API reference)
is likewise shrunk server-side to fit the window (with a visible truncation notice inside the
block).
The pre-flight HTTP 413 with the body below therefore fires only in the one case automatic
recovery cannot save — a single message that alone exceeds the window even after clipping
(architecturally impossible for any real routable model, whose window far exceeds one ~14 KB clip),
or the post-provider backstop for a model whose window the catalog does not know (no window is known
there, so nothing can be fitted). When it does fire, estimated reflects the ORIGINAL, un-shrunk
request.
{
"error": { "code": "context_overflow", "message": "This message is too long for the selected model. Switch to a model with a larger context window, remove an attachment, or start a new chat." },
"estimated": 142000,
"limit": 32000,
"suggested_model": { "provider": "anthropic", "model": "claude-sonnet-4-6", "max_input_tokens": 800000 },
"can_compact_now": true,
"first_turn": false
}
estimated— the gateway's conservative estimate of the request's input tokens.limit— the selected model's advertisedmax_input_tokens.nullwhen the model advertises no window (in that case the code is emitted as a backstop, after the provider itself rejects the request for context length, so the client still gets this graceful shape instead of the raw upstream error).suggested_model— the smallest catalog chat model with at least 50% headroom overestimatedthat the caller can actually route to. It is filtered by exactly the checks the dispatch path enforces: data-residency region, the gateway's provider allowlist, and a routable credential (a configured provider key, the keyless EU fleet — for self-serve workspaces, the platform Anthropic pool — or a tenant's Myra-provided managed grant) at model granularity: a grant-only gateway's grantedclaude-sonnet-5/claude-haiku-4-5can be suggested, but never the rest of the anthropic catalog it isn't granted — the same offer==route fold the model pickers use, so the suggestion never names a model that would then 424 for a missing key. This fold also withholds a managed grant when the deployment's platform model pool is not wired (its leg could not route), so an unconfigured pool never surfaces a Platform-Key model that would fail; and, for a self-serve plan, the model must also be inside that plan's model list. It isnullwhen no routable model has enough headroom. Because of this filter, switching to the suggested model never produces amodel_not_found, a residency/allowlist block, or aplan_model_not_allowed.first_turn—truewhen the request carries exactly one user turn (the prompt is the conversation): themessagethen says "Your input is too large for this model's context window. Shorten it, remove attachments, or pick a model with a longer context." instead of suggesting a new chat. A user turn is arole: usermessage with text or media content; an Anthropic-wiretool_result-only user message is not one.can_compact_now— kept for compatibility; always the negation offirst_turn(a first turn has no earlier history to summarise).
If the SSE stream has already opened (for example a PII/guardrail check flushed an early
progress event on a masking gateway), a raw 413 body would corrupt the event stream — so the
same overflow is delivered over the open stream as a terminal event on an HTTP 200
response, followed by data: [DONE]. A client that opted into the gateway's aig_* side
channel (x-aig-turn-id, or x-aig-extensions: 1) receives the provider_error shape below;
a plain OpenAI-compatible client receives the same failure as
{"error":{"message":"…","type":"gateway_error","code":"context_overflow","estimated":…,"limit":…,"suggested_model":…,"first_turn":…}}
— the recovery fields ride along either way; the envelope differs, and the plain shape carries
the actionable sentence in error.message rather than user_message.body (see
Inference — what a plain client receives):
data: {"aig_status":"provider_error","error_class":"context_overflow","message":"…","user_message":{"body":"…"},"estimated":148000,"limit":200000,"suggested_model":{"provider":"anthropic","model":"claude-sonnet-4-6","max_input_tokens":1000000},"first_turn":false}
data: [DONE]
The event carries the SAME estimated, limit (may be null), suggested_model (may be
null) and first_turn recovery fields as the HTTP 413 body (can_compact_now is not
forwarded on the stream — it is the negation of first_turn), so a client renders the identical
shorten / switch-model / new-chat recovery banner on either delivery path. Clients MUST treat
these fields as untrusted (validate types; a non-numeric estimated or malformed
suggested_model degrades gracefully rather than being trusted verbatim).
A turn whose egress is buffered (a code-interpreter / tool-loop leg) takes the same contract:
when its final leg overflows, the gateway delivers the honest context_overflow message and the
recovery fields exactly as above (over the open stream if one is already open, else as the 413
body), and persists the honest sentence as the turn's stored content — never the opaque
"The request could not be completed." A reloaded thread therefore shows the same honest text,
and the stored content returned by the conversations API / exports carries it too.
The stored content follows a strict at-rest policy: the gateway persists only its OWN
gateway-authored overflow sentence for a context_overflow turn — never a provider's raw error
prose. A context_overflow code on the wire is not sufficient to store the body's message: an
upstream (e.g. a BYO / proxied endpoint) that returns {"error":{"code":"context_overflow",…}}
with arbitrary prose is treated as an ordinary provider error at rest (the generic sentence is
stored), because a wire code the gateway did not author is untrusted input. The honest message is
stored only when the gateway itself produced the overflow body.
guardrail_blocked — streaming vs. non-streaming
In non-streaming mode, a guardrail block returns HTTP 400 with the guardrail_blocked error JSON.
In streaming mode ("stream": true), the guardrail check runs before the provider call. When a block occurs the gateway returns HTTP 200 with a synthetic SSE (Server-Sent Events) stream containing an error message chunk followed by data: [DONE]. This is necessary because some streaming clients do not gracefully handle a non-200 HTTP status on a streaming response.
data: {"id":"...","choices":[{"delta":{"content":"[Blocked: S1]"},"finish_reason":"stop"}]}
data: [DONE]
💡 Note: Your client should inspect the chunk content for the block message if it processes streaming responses. The log entry for the request will have
blocked: trueandblocked_by: "guardrail"regardless of the HTTP status returned.💡 Note (first-party turn-owned streams): In addition to the content chunk above, a turn-owned stream (the web app) receives a gateway-extension event
{"aig_status":"guardrail_blocked","error_class":"<class>"}so the app can render a distinct policy-block notice.error_classis a stable slug —guardrail_blocked(content policy),guardrail_pii_mandatory,guardrail_pii_media,guardrail_role_model(the selected model is not in the user role's allowlist), orguardrail_unavailable— and never carries the specific policy sub-category or request content. On a turn-owned stream the refusaldelta.contentis a generic, sub-category-free message; external API clients see the unchanged detailed prose and never receive this event. See Inference — streaming.💡 Note (response-phase blocks): A guardrail configured with
target: "response"runs on the model's buffered answer (after the provider call) rather than on the request. For a first-party turn-owned stream, a response-phase block is delivered exactly like the request-phase block above — the same{"aig_status":"guardrail_blocked","error_class":"<class>"}event followed bydata: [DONE](the gateway opens the stream if it was not already open) — so the app renders the same policy-block bubble whether the block happened before or after generation. The blocked model output is discarded and never reaches the client. The turn is persisted as a durable assistant row stamped with the block class, so on reload the web app reconstructs the identical policy-block bubble instead of showing a lone user message. External API clients receive the typedguardrail_blockederror (HTTP 400, or — if a stream was already open — a terminal SSE event:provider_errorwhen the client opted into theaig_*side channel, otherwise the OpenAI-shaped{"error":{…}}frame) unchanged, and nothing is persisted for them.💡 Note: The
Retry-Afterheader contains the window duration in seconds (e.g.60), not an absolute timestamp. It represents the maximum time before the window resets — retrying afterRetry-Afterseconds is guaranteed to succeed if no new requests have been made.
rate_limited — response headers
When the gateway's own sliding-window limit returns rate_limited, the response includes three headers:
| Header | Description |
|---|---|
X-RateLimit-Limit |
The configured request limit for the window. |
X-RateLimit-Remaining |
Estimated remaining requests in the current window (0 when blocked). |
Retry-After |
The window duration in seconds — the minimum time before retrying. |
A provider-origin rate_limited (the inference terminal relaying the provider's final 429) carries a different set: X-AIG-Rate-Limited: true always, the provider's Retry-After when it sent one, and — for providers on the Anthropic dialect — X-RateLimit-Input-Tokens-Remaining / X-RateLimit-Input-Tokens-Reset / X-RateLimit-Requests-Remaining / X-RateLimit-Requests-Reset. The gateway's own X-RateLimit-Limit / X-RateLimit-Remaining may ALSO be present on such a response (they are set on the way in whenever the gateway or token has a limit configured and describe that window — Remaining can even read 0 on the window's last admitted request); branch on X-AIG-Rate-Limited to tell the two origins apart, never on the presence of X-RateLimit-Limit.
all_providers_failed — when to expect it
This error is returned only when all of the following are true:
- The primary provider returned 5xx errors on every attempt (up to
retry_count). - Every fallback provider in the
fallbacksarray of the routing rule also failed.
4xx responses from a provider are not retried against a different model and are returned to the caller immediately (the error of the provider is forwarded, not wrapped in all_providers_failed). The one exception is a small set of transparent, gateway-side 4xx self-recoveries that fix our own request body and re-send the same model once before giving up: a deprecated request parameter is stripped, a forced tool_choice the backend cannot guided-decode is dropped, the gateway's own injected tools are removed when the backend refuses them (see model_capability_mismatch), and — since — an unsupported request parameter is stripped. Each is bounded to a single resend per turn and never fires after streaming has begun; if the resend also fails, the original error is surfaced.
⚠️ Caution: A
provider_erroron a single provider with no fallbacks configured behaves identically toall_providers_failed— both return502. Configure fallbacks in your routing rules to avoid single-provider outages surfacing as errors to your end users.
invalid_request — common causes
The following conditions commonly cause this error:
- Missing required fields in a POST body (e.g. no
slugwhen creating a tenant) - Unknown
bucketvalue inGET /stats/timeseries noutside the range 1–168
(A body that is not valid JSON is not an invalid_request: it is refused at the body read, with its own code — see the next section.)
Request body shape
Every route under /admin/v1/… and /admin/auth/… that takes a JSON body (POST / PUT / PATCH, and the one body-carrying DELETE, /me/device-token) reads it through one parser, so the rules below hold uniformly and are not re-implemented per route. (The SAML assertion consumer POST /admin/auth/saml/{tenant}/acs takes a form-encoded body and is outside these rules. The SCIM endpoint reads through the same parser in its stricter object-only mode and renders the two rejections below as SCIM Error documents.)
Accepted shape: a JSON object. Two other cases are read as "no fields present" and left to the route:
- an absent or empty body, or one containing only JSON whitespace (space, tab, CR, LF — e.g. a trailing newline from a shell) — legitimately "nothing sent";
- a well-formed non-object JSON value (
null,42,true,"text", or an array) — it parsed, so it is not malformed, but it exposes no object fields.
What the route then does with "no fields present" is the route's own contract: a POST missing a required field returns that route's 400 (e.g. slug on tenant create); a PATCH applies the recognised fields it finds — so a PATCH whose body carries no recognised field (an empty object, a scalar, a misspelt key, a key at the wrong nesting depth) is a successful no-op and answers 200 without storing anything. Unrecognised fields are ignored, never rejected.
Rejected — fail closed, at the body read (before any field is consumed). Absent and malformed are different answers: a body that was sent and could not be understood is refused rather than degraded to "nothing sent", so a save can never be reported as succeeded when the request itself was not understood. These come back in the flat admin shape with a sibling machine code:
| Code | HTTP status | When |
|---|---|---|
malformed_body |
400 | The body is not valid JSON ({, a truncated payload, a trailing token, binary garbage, a UTF-8 byte-order mark before the object, a vertical-tab/form-feed "whitespace" that JSON does not allow). No field is read from it and nothing is written; the parser's error position is logged server-side, the body content itself is never logged. {"error": "request body is not valid JSON", "code": "malformed_body"}. |
body_unreadable |
503 | A body was sent but the gateway could not read it (a large body spooled to a temp file that then failed to read). Not a client error — nothing about the request was wrong — so it carries Retry-After: 10. It only ever occurs on large bodies, so a client must not auto-retry in a tight loop; wait out the header. Distinct from an absent body, which is legal here (on the SCIM plane an absent body is itself a 400 invalidSyntax). |
Both rejections are raised only when a handler is about to read the body, i.e. after the readiness gate and — on /admin/v1/… and PATCH /admin/auth/me — after the session check and the demo-user write block (the other /admin/auth/… routes are anonymous by design, or — impersonation start, the CI session grant — run their own credential check before the body read): a demo user still receives 403, an unauthenticated caller 401, and a mid-deploy request the readiness 503. (Routes whose own permission check runs after the body read answer 400 before 403 for a malformed body; this leaks nothing a well-formed body would not also reveal.) A malformed body never yields a 500.
Unsupported-parameter strip-on-400
Some Myra-hosted OpenAI-compatible backends hard-400 a request that carries a reasoning-control
parameter their model does not accept — most often reasoning_effort on a non-reasoning model
(litellm.UnsupportedParamsError: openai does not support parameters: ['reasoning_effort'], for
model=…). Because the gateway attaches such a parameter (the caller never sent it), the turn used to
die outright. The gateway now strips the offending parameter and re-sends the same model once; the
original 400 is surfaced only if the resend also fails.
This is a fail-closed control, because the 400 body is untrusted (a provider, or a per-gateway
provider_base_urls endpoint, controls its text):
- Accepted (stripped + retried): a
400whose body is a litellmUnsupportedParamsError(or carries the exact phrase "does not support parameters:") and whose named parameter list contains only parameters on a fixed, hand-audited allow-list of pure reasoning/verbosity controls whose removal can never change the response contract — currentlyreasoning_effort,reasoning, andthinking. Only the named parameters actually present on the request body are removed, and the retry is bounded to one per turn. - Rejected (no strip, original
400surfaced): any body that is not that error; a body that names the error but from which no parameter list can be located or that lists no quoted identifier (absent and malformed both surface the original error — a garbled body can never trigger a blind retry); and — the load-bearing guard — a list that names any parameter not on the allow-list (e.g.tools,response_format,stream,temperature). One off-list name blocks the entire strip: the gateway never silently drops a parameter that carries safety or response-shape meaning.
The retry never fires after the first response byte has reached the client (streaming safety).
chat PDF/image OCR — the 422 extraction failure
When a chat upload needs OCR (a scanned/text-less PDF page, or an image), the gateway runs an
on-prem extractor that calls the OCR vision model. If that call fails the upload returns 422
with a detail such as OCR could not extract text from PDF: <reason>. The <reason> is now
actionable rather than a bare HTTP status:
(type=bad_request_error)/(type=validation_error)— the OCR model rejected the document (unsupported or unreadable page content). These two content-reject classes are the only provider error types surfaced to the caller. This is a terminal reject — a re-upload of the same page will not succeed.- No
typesuffix — a transient fleet condition (the extractor already bounded-retried and gave up) or an unrecognised error. Retry later. pdf_password_protected(a distinct 422code) — the PDF is encrypted with a user password. The extractor detects this up front (before any page read) and rejects it with a stable, machine-readable reason, so the chat upload returns the specific code above (and the worker/connector path records apassword-protectedreason) instead of a generic extraction failure. An owner-password-only PDF (readable without a password) is not rejected.
Accepted vs. rejected shape (the provider response is untrusted — principle 11): the OCR model's
error body is parsed defensively (size-capped, JSON-only, fail-closed). Only a known, bounded
content-reject type token is placed on the user-facing message; the model's free-text message
and any internal-infrastructure type (DB/auth/quota/model-access classes) are operator-log only
(the operator-facing server log), never the user response — so a malformed, oversized, or hostile error body can never
leak backend internals or forge a control line into the reply. A control-plane failure is classified
transient (retried); a genuine content reject stays terminal. Separately, each OCR page/image
strip is downscaled to the OCR model's pixel budget before the call, so an abnormally large
rendered page cannot itself provoke the reject.
Request body size limits (413 payload_too_large)
Every admin endpoint has a maximum request-body size. A body larger than its route's limit is rejected by the gateway edge with HTTP 413 before the handler runs — so an oversized upload never spools or parses on a worker. This bounds the pre-authentication surface (login, one-time codes, SAML sign-in, anonymous crash reports are all a few kilobytes) and the authenticated one.
The 413 body matches what the client asked for (the Accept header):
- An API / XHR client (
Accept: application/json,*/*, or noAccept) receives JSON:
- A real browser navigation (
Accept: text/html— e.g. the SAML identity provider's form POST) receives a small static HTML "Request too large" page instead of raw JSON.
The exact byte figure is not repeated in the 413 body; the browser SPA reads the per-surface
limits from GET /admin/auth/me (the limits object) and pre-checks a file's size before
sending, so the user sees a clear "too large" message rather than a failed request.
Per-route limits (the values that matter; everything else rides the general admin cap):
| Route | Limit | Notes |
|---|---|---|
POST /admin/v1/chat/files (file upload) |
~140 MB | base64 of a 100 MiB document |
POST /admin/v1/projects/{id}/knowledge/upload |
~140 MB | base64 of a 100 MiB document |
POST /admin/v1/app-feedback |
~88 MB | up to 6 evidence images inline |
all other /admin/v1/... (general cap) |
16 MB | chat export, transcription, messages, config, … |
POST /admin/v1/client-errors (anonymous) |
64 KB | crash reports; truncated further server-side |
/admin/auth/... except SAML (login, one-time codes, OIDC callbacks) |
16 KB | small JSON only |
POST /admin/auth/saml/{tenant}/acs (SAML ACS) |
1 MB | carries the SAMLResponse |
Rejected vs accepted. A body up to the route's limit reaches the handler and is validated
there (the handler may still reject it for content reasons — e.g. an unsupported file type is a
400, an image over the per-image cap is a handler 413). A body over the route's limit is the
edge 413 payload_too_large above and never reaches the handler.
Example error responses
⭐ Example: The following examples show the JSON body returned for common error scenarios.
401 — authentication failures
The first three share "code": "unauthorized"; the message tells them apart. An expired token is reported separately as token_expired, an explicitly revoked one as token_revoked (both below).
No token sent:
{
"error": {
"code": "unauthorized",
"message": "No API token provided. Send your gateway token as 'Authorization: Bearer <token>', or the 'x-aig-token' / 'x-api-key' header."
}
}
A valid token used on the wrong gateway (it belongs to a different gateway of the same tenant — fix the {gateway} path segment; do not regenerate the token):
{
"error": {
"code": "unauthorized",
"message": "This token is recognized but is scoped to a different gateway in this tenant. Check the {gateway} segment of your URL path /v1/{tenant}/{gateway}/{provider}/... ."
}
}
An unknown token, or a token from another tenant (generic — the two are indistinguishable by design, so the response never confirms a token's existence in a tenant you have not named):
{
"error": {
"code": "unauthorized",
"message": "Missing or invalid gateway token for gateway '<gateway-slug>'. Tokens are scoped to a single gateway — verify the {gateway} segment of /v1/{tenant}/{gateway}/{provider}/... matches where the token was created."
}
}
An expired token (its own code — the credential was real and simply lapsed, so a client may re-mint and retry, which it must not do for unauthorized):
{
"error": {
"code": "token_expired",
"message": "Gateway token expired. Regenerate it under Gateways -> Tokens and update your client/agent."
}
}
A revoked token (its own code — the credential was real and was withdrawn on purpose; the date is present when the revocation is on record, and the caller needs a new token, not a retry):
{
"error": {
"code": "token_revoked",
"message": "This gateway token was revoked on 2026-09-13T16:22:07Z. Mint a new token and update the caller."
}
}
429 — rate limited
429 — quota exceeded
{
"error": {
"code": "quota_exceeded",
"message": "Gateway budget $200.0000 exceeded (spent $200.0019). Adjust budget_usd in the gateway config (PATCH /admin/v1/gateways/{id}) or reset spend (DELETE /admin/v1/gateways/{id}/budget)."
}
}
400 — guardrail blocked
{
"error": {
"code": "guardrail_blocked",
"message": "Request blocked by content policy (block-pci): cc – Credit/Debit Card Number"
}
}
502 — all providers failed — returned when every provider in the fallback chain has been exhausted. The JSON envelope is identical to the above; only code and message differ.
Endpoint-specific error families
Some endpoints define their own machine-readable codes documented alongside the endpoint rather than repeated here (one source of truth). Notably the conversation summarize / compaction family — corpus_too_large (413), span_not_found (404), span_inverted / span_content_invalid (400), conflict_supersedes_stale (409), and the streaming-phase 502 codes — is defined in Conversations. SCIM uses the RFC 7644 error envelope (SCIM).