Skip to content

Inference API (/v1)

The inference API is the public, OpenAI-compatible surface that applications call to run model requests through a gateway. Every inference path is served under /v1/. The gateway accepts the OpenAI chat completions request body for all providers and translates it to each provider's native wire format automatically.

Data-file preview bounding (request mutation). A user-message text block in the gateway's inline data-file format — first line exactly [File: <name>] where <name> ends in a spreadsheet extension (.xlsx, .xlsm, .csv, .tsv, .ods), followed by a blank line and the extracted text — is treated as a shrinkable preview, not as prose. Before dispatch the gateway MAY truncate such blocks (appending a visible notice inside the block): down to the routed model's context window when the request would otherwise be refused with context_overflow — for self-hosted models the preview budget is computed worst-case (digit characters count as one token each, since local tokenizers split digits individually) with the full potential output budget reserved, because those models share one window between input and output and enforce it with their own tokenizer — and down to a ~16 KiB head when the raw file is available to the code interpreter's staged-input channel for this conversation (the sandbox reads the full file; the preview is orientation only). Blocks in any other format, assistant/tool messages, and nested tool-result content are never touched. API clients that need their text verbatim must simply not use this exact block format.

Inline text attachment size. Text attachments (Markdown .md, HTML .html, plain text .txt) are inlined into the turn as ordinary message text. Accepted: up to 512 KB of decoded text per such attachment. Rejected: the web app refuses an over-cap text attachment before the request is sent, with a visible message — it never inlines it (an oversized inline text would blow the model's context window and dead-turn with no answer). Server-extracted text that is inlined via POST /chat/files (.docx, .doc legacy binary Word, .pdf, .pptx, image OCR, and a spreadsheet routed through the inline-extract path) is instead truncated at the same 512 KB ceiling with a model-visible notice inside the text — not rejected. (A spreadsheet uploaded to the provider's Files API rides as a raw file_id, not inline text, so it is neither inlined nor truncated.) Note there is no dedicated server-side byte cap on inline text that arrives as plain message content (it is indistinguishable from a paste): the server's size backstop is the window-relative context gate, which now auto-compacts an over-window turn (summarizes older history, head+tail-clips the current message, and proceeds — see context_overflow) rather than refusing, so an oversized paste no longer dead-ends the turn; the context_overflow 413 fires only for a single message that alone exceeds the window even after clipping (and is skipped when the routed model has no known context window), plus the upstream provider's own limits. The 512 KB attachment cap is therefore a client-side bound; direct API callers should size their own request bodies.

Inline binary (image / PDF) and total request-body size — client-side CDN guard. A vision-capable or PDF-capable (Anthropic) model receives images and PDFs as native inline base64 (image_url / a document block) carried in this request body, not via POST /chat/files. Because the whole request body — the base64 media plus the entire conversation history — travels to this inference host, whose edge caps the request body at 11 MB (a body above it is dropped with a 413; behind the CDN, that 413 historically carried no CORS headers, blinding the browser's fetch to an opaque network error), the web app enforces two client-side guards before sending: (1) a per-attachment limit of about 7.4 MB of raw file bytes (~10 MB base64, conservatively under the 11 MB inference edge) — an oversized inline image is first downscaled + re-encoded in the browser to fit this budget (the pinned vision models accept far higher resolution than the edge allows — measured HTTP 200 up to 9.4 MP, with the edge, not the model, being the binding limit — so a fitted image is equivalent), and only a non-image or an image that still cannot fit is rejected with a specific "…is too large to upload (max ~7.4 MB)…" message; (2) an assembled-body preflight — if the fully-serialized request body (media and history) would exceed that budget, the send is refused with a "This message is too large to send… remove or shorten large attachments, or start a new conversation." message. Both turn the CORS-blinded 413 dead end ("the gateway is unreachable, keep retrying") into an actionable message. These are UX guards, not part of the endpoint contract — the endpoint itself has no application-layer body cap here beyond context_overflow; direct API callers are not bound by these browser guards, but a body above the edge cap will still be dropped at the edge. The admin POST /chat/files upload path is a separate host with its own, larger cap (140 MB edge → a ~110 MB client guard), so a large document uploaded for extraction is bounded independently from an inline attachment — see Client-side upload size limit in the Conversations reference.

Base URL: https://<your-gateway-host>


Endpoint paths

Every inference request follows the same four-segment shape:

POST /v1/{tenant}/{gateway}/{provider}/chat/completions
Segment Meaning
{tenant} Tenant slug.
{gateway} Gateway slug within the tenant.
{provider} A registered provider name, or the pseudo-provider compat.

The {provider} segment selects how the request is routed:

  • Pinned provider — naming a real provider (for example openai, anthropic) and ending the path with /chat/completions pins that exact provider with no model-name guessing.
  • Compat (unified) — compat is a pseudo-provider; the gateway guesses the provider from the request body's model field and always returns an OpenAI-shaped response.
POST /v1/{tenant}/{gateway}/compat/chat/completions

For the full list of supported providers, the compat model-resolution tiers, and the native endpoint table, see Providers overview.

Anthropic-native passthrough

Provider paths that do not end in /chat/completions are forwarded to the upstream provider in its native wire format. The supported case is the Anthropic Messages API, used by Anthropic-native clients such as Claude Code:

POST /v1/{tenant}/{gateway}/anthropic/v1/messages

💡 Note: On the native passthrough path the request and response bodies are the provider's own format, not the OpenAI chat-completions shape. Use the /chat/completions suffix whenever you want the OpenAI-compatible translation.

Only servable endpoints are accepted

The entire provider path is validated against a fixed allowlist of servable endpoints before the request is authenticated. Any other path is rejected with 404 endpoint_not_found — the request never reaches the provider and nothing is dialed upstream. The accepted paths (an optional /v1 version segment is allowed on each) are:

Path Endpoint
/chat/completions OpenAI chat completions
/completions OpenAI legacy (text) completions
/embeddings OpenAI embeddings
/v1/messages Anthropic Messages
/v1/messages/count_tokens Anthropic token counting
/v1/models list models (read-only OpenAI-SDK helper)
POST /v1/{tenant}/{gateway}/{provider}/v1/anything-else         → 404 endpoint_not_found
POST /v1/{tenant}/{gateway}/{provider}/internal/x/embeddings    → 404 endpoint_not_found

The match is on the whole path, not a trailing suffix: a path such as /internal-admin/embeddings that merely ends in a servable name is rejected — an attacker must not be able to smuggle an arbitrary upstream path in front of a known endpoint. A missing path (a bare …/{provider} URL) and a bare /v1 are rejected with 404 endpoint_not_found — append an explicit endpoint such as /chat/completions. An absent path fails closed exactly like a present, unrecognised one: the two share the reject outcome, never a permissive dial (before, an absent path was forwarded to the provider's bare base URL and 502'd for every path-forwarding provider). Providers that build their endpoint from the model rather than the path — Gemini, Bedrock, Cohere, Vertex, Azure — previously tolerated a bare URL; they now require an explicit endpoint too, for one uniform rule. This is an authorization boundary, not a convenience check: every real provider forwards the path verbatim into its upstream URL, and the platform-managed myra provider attaches Myra's platform credential, so an unvalidated path would be a server-side request forgery. The allowlist is exactly the set of priced/known endpoints, so no accepted request escapes usage metering.

💡 Note (model-list probes). The read-only …/models list endpoint stays supported on both the explicit-provider path (/v1/{tenant}/{gateway}/{provider}/v1/models) and /compat (…/{gateway}/compat/models), so an OpenAI-SDK client that auto-probes /models on start-up keeps working. Only genuinely unknown paths are rejected.

The Messages API is not available under /compat

The /compat surface speaks the OpenAI schema in both directions — requests are read as OpenAI chat-completions and responses are always OpenAI-shaped. An Anthropic Messages request sent to a compat sub-path is therefore rejected with invalid_request (400) rather than partially processed:

POST /v1/{tenant}/{gateway}/compat/v1/messages        → 400 invalid_request

The same rejection applies to /compat/…/v1/messages/count_tokens. Use the native route for an Anthropic-native client, which is where the Messages wire format is actually served:

POST /v1/{tenant}/{gateway}/anthropic/v1/messages     → Messages in, Messages out

This also rejects the common SDK misconfiguration of appending /v1/messages to a base URL that already ends in /compat/chat/completions.

Usage metering on the native Messages route

/v1/messages is the Anthropic Messages endpoint, and the gateway meters it as such whichever provider serves it. A provider whose upstream forwards the native path — the self-hosted myra fleet (served by LiteLLM) is the practical case — relays the Anthropic response verbatim to the client. The gateway selects its usage reader by the endpoint it actually dialed and confirms it against the response's own envelope (a message_start or error frame on a stream, a type: "message" object on a stream: false turn); an upstream that answers that path in another dialect keeps that provider's own reader. On a confirmed Anthropic response the recorded leg carries:

  • input_tokens / output_tokens from message_start and message_delta (or the Messages object's usage), plus the cache_creation (5m / 1h) and cache_read buckets — the same counters a Claude-served turn on this route records, read with the same rules, so spend accrual against the gateway/tenant budget and cap enforcement follow the same path. Prefix-cache parity with /chat/completions: a prefix-cached prompt is booked as input = prompt − cached, cache_read = cached on both routes. On /v1/messages the relaying LiteLLM adapter nets the cached share out of input_tokens and surfaces it as cache_read_input_tokens, which this reader books as cache_read. On /chat/completions the gateway reads the OpenAI-compat usage.prompt_tokens_details.cached_tokens (a subset of prompt_tokens) and performs the same split itself (cached_tokens is netted out of input and booked as cache_read). So the same prefix-cached prompt books the same buckets on either route. Two contingencies, both fleet-side (not the gateway): the self-hosted fleet only reports a cached share at all when vLLM prefix-token-details reporting is enabled, and /v1/messages parity additionally requires the fleet's LiteLLM to map cached_tokens → cache_read_input_tokens (a stock adapter nets the share out but drops it, leaving cache_read = 0 on that route). When no cached share is reported, both routes book input = prompt, cache_read = 0. Pricing note: a model with no configured cache_read_per_1k bills the cache_read bucket at the input rate (an explicit 0 means free), so splitting the bucket never turns cached tokens free by default;
  • request_log.model_version = the served model the response announced;
  • on a gateway with payload logging and PII masking, the audit response_raw of the streamed answer.

The relayed bytes are never altered by this accounting. Two related rules on this route:

  • /v1/messages/count_tokens ({"input_tokens": N}, no content — a success by spec) is recognised by the dialed endpoint on every provider whose upstream endpoint carries the path, so it is never filed as an empty answer.
  • The gateway's server-side tool loop on the native Messages route runs only where the provider can read the Anthropic dialect it serves: anthropic itself, or a routed provider that dials its own endpoint. A provider that forwards the path to an upstream answering in the Anthropic dialect through an OpenAI-compatible reader (myra, groq, together, …) gets a transparent passthrough instead: gateway tools are not injected, and a request carrying MCP connector references is rejected with invalid_request (400) naming the reason. The rule is decided by the request's primary provider; a fallback provider in the routing chain that would relay the dialect is skipped at dispatch on a tool-loop turn (the primary keeps its tools), never folded with the wrong reader.
  • The gateway's own inner legs — the tool loop's compaction summary (the max_rounds_reached summarize-and-continue step) and the fetch-and-read helper the agentic fetch tool runs — are OpenAI-shaped stream: false chat completions whatever the parent turn's wire was. They never dial a parent's native /v1/messages path: on a path-forwarding provider that path is dropped for the inner leg in favour of the provider's canonical chat endpoint (e.g. a native-Messages parent on anthropic, cohere, gemini or bedrock whose agentic fetch runs on the myra default model), while a parent that is already on a chat-completions path keeps it (the inner leg dials exactly the parent's endpoint). The parent's path is restored once the inner leg returns, on every exit.
  • A tool loop that the gateway has to stop never dead-ends the turn with a bare note. Before giving up it makes ONE final no-tools model call (tool calling disabled) and delivers that answer as normal assistant content, followed by an honest note:
  • repeated_tool_call (a tool kept returning the same or an empty result and the model re-issued an identical call) → the model answers from its own knowledge, with a note saying so. It does not tell the user to rephrase.
  • max_rounds_reached (the 25-round per-turn tool-call cap) → the gateway first tries one conversation compaction (summarize_and_continue); when that cannot finish the task it delivers a best-effort answer and emits the typed tool_loop_terminated event with can_continue: true, and the note is actionable ("Ask me to continue …"). The SPA turns can_continue into a one-click Continue affordance that resumes the same task in a fresh turn; the per-turn 25-round ceiling is unchanged and the fallback adds at most one bounded model call per turn.
  • wall_clock_exceeded keeps its existing partial-synthesis behaviour. Every synthesis path fails closed: if the extra call errors or produces nothing, the turn ends on the honest stop note rather than a fabricated answer.
  • On a stream, the final message_delta usage snapshot is cumulative and authoritative: a positive input_tokens / cache figure there supersedes the message_start snapshot (on an Anthropic server-tool turn the cumulative input is the billed one); a zero or absent figure never overwrites a value message_start already reported. Usage values are untrusted provider output: a numeric string is coerced, and a non-finite, negative or non-numeric value reads as "not reported".

Accepted compat sub-paths. /compat/chat/completions (and its /compat/v1/chat/completions spelling) for chat, plus the non-chat OpenAI endpoints such as /compat/embeddings and /compat/models, which are forwarded to the provider's corresponding OpenAI endpoint. The Anthropic token-counting endpoint (/v1/messages/count_tokens) is a separate, non-streaming endpoint and is unaffected by the rejection above.

⚠️ Response-phase guardrails on a native stream: true request. When the gateway has a response-stage detector that cannot be applied incrementally — a content-safety block, a regex/presidio scrub, or a flag (i.e. anything other than the inline PII token-restore performed by pii_protector / custom_pii) — a native stream: true request is internally buffered: the gateway collects the full upstream answer, runs the detector, and then re-emits it as native Anthropic SSE events (message_start → content_block_* → message_delta → message_stop; no OpenAI [DONE]). The first byte therefore arrives only after the full generation completes (bounded by timeout_ms), exactly as for a native stream: false request — the latency cost of response-stage scanning. A gateway whose only response detector is an inline PII masker keeps streaming incrementally. A response-stage block on a native stream is delivered as an HTTP 400 guardrail_blocked error (the answer is never sent). SSE clients that enforce a short inter-event idle timeout should allow for this buffered window.

⚠️ Mid-stream provider failures on a buffered turn. When a PII-masking or guardrail configuration forces a stream: true turn onto the internally buffered path (or a stream: false tool-loop turn runs buffered by nature) and the upstream provider's stream fails mid-generation — a mid-stream error event such as Anthropic overloaded_error, a read error, or a truncated stream — the gateway never answers with a clean HTTP 200 carrying an empty message. Instead it: (1) retries the attempt transparently (same provider, then the fallback chain) when nothing was delivered and no server-side tool has executed; (2) fails loud with the typed provider_error / all_providers_failed error (as an error frame + aig_status: "provider_error" event on an already-open SSE stream) when retries are exhausted or unsafe; or (3) salvages a partial — any text, files, or images the turn already produced are delivered and persisted together with the explicit error surface, which is the same one the live streaming wire carries for that failure kind: a provider error event (provider_error) is the error frame plus the aig_status: "provider_error" banner event; a read error or a truncated stream (stream_errored / truncated) is the single terminator error frame (provider_error "connection failed" / stream_truncated), no banner, and the visible answer ends with the persisted "The response was interrupted before it completed. Please try again." note described under Visible truncation note — so the salvaged turn reloads exactly as it rendered live. (A leg that reached a natural end and then hit a read error keeps the frame and gets no note — the answer is complete.) A non-streaming JSON response carries the same content (the note included) plus an X-AIG-Error response header whose value is the failure kind — provider_error, stream_errored, or truncated, the same lowercase vocabulary the typed error codes use. A machine client that consumes content as the bare answer must read X-AIG-Error to tell a died-mid-answer turn from a completed one; the note is English-only gateway prose, not localized.


Authentication

Send a token on every request unless the gateway has authentication disabled. The token may be a gateway token or a personal access token.

The gateway reads the token from the first header present, in this order:

Order Header Form
1 x-aig-token Raw token value.
2 Authorization Bearer <token>.
3 x-api-key Raw token value.
curl -s -X POST "https://<your-gateway-host>/v1/myapp/prod/compat/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <token>" \
  -d '{
    "model": "claude-opus-4-6",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

A missing or invalid token returns 401 unauthorized; an expired token returns a distinct 401 token_expired and a revoked token a distinct 401 token_revoked (each with an actionable message naming where to regenerate it). Branch your refresh logic on the stable code, not on the shared 401 status. See Authentication and Error codes.

Inference is refused for viewers

A request bound to a user whose role is viewer is refused 403 forbidden ("Viewer role cannot make inference requests") — before any model is called, so no credit is spent. This holds on both user-bound auth paths: a raw /v1 call and the chat / /easy playground path (the busier channel, where a viewer would otherwise burn trial credit). The check fails closed on an unconfirmable role too: a user whose role cannot be resolved to a known string (a legacy role-assignment gap) is barred exactly like a viewer, never allowed through by default. A GDPR Art. 18 processing restriction is evaluated first, so a restricted viewer receives 403 processing_restricted (the legal-hold code), not this one. Other roles — member, demouser (limited inference via its demo quota), ki_manager, tenant_admin, admin — are unaffected. The mirror control on the admin session plane (a viewer's chat writes) is Viewers cannot spend credit or reshape a conversation.


Request body

All providers accept the OpenAI chat completions body. The common fields are:

Field Type Notes
model string Required. Must be a non-empty string. On compat, this also selects the provider.
messages array Required. A non-empty array of OpenAI message objects (role, content).
stream boolean When true, the response is a Server-Sent Events stream.
max_tokens integer Upper bound on generated tokens. Optional — see the default policy below.
tools array OpenAI tool/function definitions.

max_tokens default policy

max_tokens is a ceiling, not a target — the model may stop earlier. The gateway applies a single catalog-driven policy at dispatch:

  • Omitted (or ≤ 0) on a self-hosted model (our own vLLM fleet): the gateway fills it from the model's catalog output ceiling (model_price.max_output_tokens). This prevents the serving stack's own tiny per-model default (a thinking model can otherwise spend its whole budget on hidden reasoning and return an empty answer). The injected budget is bounded to a sane output ceiling and window-fitted against the estimated prompt; when only a sliver of room remains (fewer than ~512 tokens) the ~512-token useful floor is injected instead of a starvation budget, and when no safe room remains at all nothing is injected and the serving default applies. If the catalog has no ceiling for the model, nothing is injected and the serving default applies.
  • Omitted on a third-party provider (Anthropic, OpenAI, Bedrock, OpenRouter, …): the gateway does not inject a value — the provider applies its own default and enforces its own minimum/maximum. (Anthropic, which requires max_tokens, is filled from its own per-model ceiling — on both the OpenAI- compatible and the native /v1/messages paths.)
  • Larger than the model's real ceiling (any provider): the value is clamped down to that ceiling so the provider does not reject the request. Exception: on the native Anthropic /v1/messages passthrough path an over-large value is forwarded verbatim (a genuine Messages client may know a newer model than the gateway's catalog, so clamping to a possibly-stale ceiling would truncate its intended budget); Anthropic then applies its own limit.
  • Within the ceiling: an explicit value is sent verbatim — except for the self-hosted window budgeting below.

Native context management / compaction (context_management)

Anthropic's native context-management field on /v1/messages — context_management: {"edits": [{"type": "…"}]} — is untrusted client input and is sanitized at the gateway boundary. The compact_20260112 compaction strategy is only valid on models that declare support for it (sonnet-class; haiku does not), and Anthropic returns a 400 if it is sent to an unsupported model.

  • Accepted: a context_management whose edits is a JSON array. A compact_20260112 edit is forwarded only when the resolved (post-routing) model supports it. Other strategies (e.g. clear_tool_uses_20250919) are separate features and pass through untouched.
  • Rejected / stripped (fail closed): when the resolved model does not support compact_20260112, the gateway strips that edit from context_management.edits (and removes the context_management field entirely if no edits remain) and removes the compact-2026-01-12 anthropic-beta token — so the gateway never forwards a request Anthropic would 400. The routed model, not the client's model, is authoritative (the gateway may route a compaction-capable request to an incapable model), so this adapts rather than erroring. A context_management whose edits is not an array (an object, a scalar, or absent) is left untouched — a malformed body is the client's own 400.
  • /v1/messages/count_tokens: this sub-endpoint accepts no context_management (any model, any strategy); the gateway strips it wholesale.
  • Gateway-injected compaction merges, never clobbers: when the gateway itself enables compaction (context_compaction config) for a supported model, it appends its compact_20260112 edit to the client's existing context_management.edits array (deduped), preserving any client edits — it does not replace the field.

Self-hosted window budgeting (our own vLLM models)

A self-hosted model enforces a hard rule: input_tokens + max_tokens must not exceed the model's context window (max_model_len). A large max_tokens on top of a large prompt therefore overflows the window and the request is rejected, even though the input alone fits. For our self-hosted models only (never third-party providers, whose windows are much larger and are not policed this way), the gateway keeps input and output inside the window automatically:

  • Pre-flight clamp. Before dispatch, max_tokens is reduced so that the estimated prompt plus the output budget fits the window. The result is only ever smaller than the value you sent — a request that already fits is untouched. If the input is so large that no useful output budget remains (below ~512 tokens), max_tokens is left as-is rather than shrunk to a useless near-zero value. Inline data-file previews are additionally window-fitted pre-flight in a worst-case token frame with the full potential output budget reserved (see Data-file preview bounding above), so a near-window preview is shrunk before the serving stack can reject the request.
  • Overflow repair (one resend). If a request still overflows (the estimate under-counts, or a multi-step tool loop grew the prompt), the model returns a context-window error. When that error shows the input still fits the window, the gateway resends the request once with a reduced max_tokens; your full input is preserved and the answer is generated. The budget depends on what the error actually reports:
  • a measured input count (the serving stack printed its real tokenizer count): the remaining budget minus a 128-token safety margin (window − measured input − 128) — the margin absorbs re-count drift between serving layers;
  • an "at least N input tokens" figure: this is a derived lower bound (the smallest input consistent with the rejection), not a measurement, so no remaining-budget arithmetic is trustworthy — the gateway resends with the minimum useful output budget (~512 tokens) instead, which succeeds whenever a useful rescue is possible at all. The answer may be shorter than usual on this last-resort path.
  • Conservative fallback when the counts are unreadable. If the overflow error is recognisable as a context-window rejection but its token counts cannot be safely parsed (missing, malformed, or implausible numbers), the gateway still resends once, with a budget computed purely from its own prompt estimate and the model's catalog window, reduced by a further 4096-token margin and always strictly below the budget that just failed. No number from the error body is used on this path. If no useful budget remains (below ~512 tokens), the request is not resent.
  • Genuine input-overflow is not repaired here. If the input alone exceeds the window (no output budget to reclaim), the request is not resent — it surfaces the standard context_length_exceeded error (see Error codes).

The window-error bodies the gateway parses are untrusted model/serving output: only the specific captured vLLM/LiteLLM phrasings are recognised, and any implausible value (a count ≤ 0 or > 10,000,000, an input that meets or exceeds the window, or a missing input count) is rejected — those bodies take the conservative-fallback path above (which consumes no numeric value from the error body) or, failing its guards, the original error is returned unchanged. In every case the resend happens at most once — a second overflow surfaces the standard error.

Validation

The gateway validates the request body at the trust boundary before any inference runs. A request is rejected with 400 invalid_request when:

  • the body is empty (Empty request body) or is not valid JSON (Invalid JSON body);
  • the body is valid JSON but not a JSON object — a bare scalar, boolean, or null top-level body (e.g. 123, true, null) is rejected (Request body must be a JSON object). (A top-level JSON array is likewise not a valid request; it is rejected on the required model field below.);
  • messages is present but is not an array ('messages' must be an array);
  • messages is an empty array [] ('messages' must not be empty) — an empty array is a malformed request, not an empty conversation, and never reaches a provider;
  • any element of messages is not an object (each 'messages' entry must be an object);
  • model is missing, empty, or not a string (Missing or invalid 'model' field).

A JSON null messages value is treated as absent (equivalent to omitting the field), not as an empty array.

Trailing assistant turns are normalized (Anthropic)

A well-formed chat request ends with a user turn. For Anthropic models on the compat path, if the conversation ends with one or more trailing assistant turns (an accidental assistant-message prefill — never intended in a chat flow), the gateway drops the trailing assistant turn(s) so the array ends with the preceding user/tool-result turn before the request is sent. This is required because the Anthropic 4.x reasoning family rejects a last-assistant-turn prefill with a provider 400 ("this model does not support assistant message prefill; the conversation must end with a user message"). The normalization is trailing-only — interior assistant turns and tool-call/tool-result pairings are untouched — and a conversation that is entirely assistant turns (no user turn at all) is left unchanged and rejected by the provider. It applies only to the compat path; the anthropic native passthrough forwards the body verbatim (where an intentional prefill is the caller's own choice).

Per-provider nuances — system-prompt extraction, extended thinking, prompt caching, model name prefixes — are covered in Providers overview and the individual provider pages.

chat_template_kwargs — the adaptive-thinking flag

Accepted shape: an optional object {"enable_thinking": true|false} (the vLLM/Qwen chat-template knob; the Chat sends it to every non-Anthropic model). Only the strict boolean false is read as "thinking off"; the gateway translates or drops the field per provider and never forwards it where it would be rejected. Strict cloud providers (Mistral, Groq, OpenAI, Cohere) never receive the field itself; where the model has a reasoning knob the OFF signal is translated instead — Magistral and Groq's gpt-oss/qwen3 get reasoning_effort, Cohere's command-a-reasoning-* gets thinking: {type: "disabled"}, Gemini 2.5 gets thinkingBudget: 0, OpenRouter gets reasoning: {effort: "low"}; models without a knob (OpenAI direct, Gemini 2.0/3.x, Mistral Large, …) get no translation. A Myra-served Mistral-tokenizer model (mistral-small-4) receives neither the field nor a translation (see Myra). Qwen and Gemma fleet models receive it unchanged. Rejected/ignored: a null, non-object or otherwise malformed value is treated as absent by every adapter (one shared null-safe test — no translation, and the field is still dropped where the provider would 400 on it); it never causes a gateway error. A reasoning_effort you set yourself always wins over the translation.

Retired model ids upgrade to their successor (Myra fleet)

When the Myra fleet retires a self-hosted model in favour of a newer one, the retired id is deprecated in the catalogue (GET /model-prices shows deprecated_at) and would normally be refused at serve time with 400 model_not_found (see Models). For the fleet's own retired ids the gateway carries a successor map, so an integration that pinned the old id keeps working instead of hard-failing on the day of the rollover:

Requested model Served as
qwen3.6-27b, qwen3.6-35b-a3b qwen3.8-27b
gemma-4-26b-a4b-it gemma-4-31b-it

Accepted shape. The rewrite applies to the model field of an inference request — bare (qwen3.6-27b) or provider-prefixed (myra/qwen3.6-27b), on the provider-native myra route and on /compat/... — once the id has been normalised exactly as any other id is (provider resolved, prefix stripped), to a routing rule's actions.model and each fallbacks[].model when the rule remaps onto a retired id (the same resolver; an upgraded fallback target is only logged — the response headers and the request log always describe the model that served the turn; a rule whose model condition names a retired id simply no longer matches, see Routing rules), and to a gateway's agentic_fetch.model override (the inner fetch agent runs on the successor; no response signal, the inner leg does not serve the turn). It happens before cost, capability checks and the request log read the model, so every downstream surface names the model that was actually served: request_log.model, the cost figures and the analytics rows carry the successor; the requested id is kept for attribution in request_log.meta under the gateway-owned key aig_model_upgraded_from (a client cannot set it — x-aig-meta-aig_* headers are dropped like every other reserved key).

Stored model no longer permitted. On a self-serve workspace, a run whose model comes from stored configuration — an agent's model (/agents/{slug}/invoke, scheduled runs, workflows, delegated sub-agents), a scheduled prompt task's model, the Copilot document-AI model — that the plan or the EU-Gov add-on no longer permits runs on the pinned gateway's Auto choice instead of failing (see Tenant entitlements). request_log.model and the cost name the model that ran; the stored model is kept in request_log.meta under the gateway-owned key aig_substituted_from. A model the caller sends is never substituted.

What the response says. Every response to a rewritten request — buffered or streamed, and including any error raised after the rewrite (a guardrail block, a plan or gateway refusal of the successor, a provider error) — carries the headers below. Responses produced before the model is read — authentication, rate limiting, the tenant/token budget 429, the IP allowlist, and an exact-match cache hit — never do.

Header Value
X-AIG-Model-Upgraded-From the retired id the request — or the matching routing rule's rewrite — named (qwen3.6-27b)
Deprecation RFC 9745 form @<unix-seconds> — when the retired row was deprecated. Sent only when the client's own model was retired (a routing rule's rewrite and an agent's stored model — on /agents/{slug}/invoke and scheduled runs — are the tenant's configuration, which the caller cannot act on: those carry X-AIG-Model-Upgraded-From only), and absent when the catalogue no longer has the old row at all. The RFC defines the field for the resource; here it qualifies the requested model, and it only ever appears together with X-AIG-Model-Upgraded-From — the endpoint itself is not deprecated
X-AIG-LLM-Model (buffered) / X-AIG-Model (streamed) the successor that served the turn

Both new headers are CORS-exposed to browser clients.

What is NOT rewritten (each fails closed toward the existing behaviour — the deprecation gate answers 400 model_not_found, or the block 403 model_disabled_on_gateway):

  • an id with no entry in the map — the gateway never guesses a successor from the name (qwen3.5-27b stays a 400);
  • the same id on a non-Myra provider (openrouter/qwen3.6-27b is a passthrough serve, not the fleet);
  • a retired id the tenant admin disabled on this gateway (Disabling a Myra model) — a rename never sidesteps an admin's block;
  • a retired id whose catalogue row is live again (an operator un-deprecated it, or the fleet re-serves it) — the operator's decision wins, the request goes to the id as named;
  • a successor that is itself not a live catalogue row (the map is single-hop; chains are not followed);
  • platform-owned targets applied after routing: a self-serve plan's fallback / degrade_model naming a retired id is still refused at the serve-time gate (400 model_not_found) — platform operators update those at a rollover;
  • a catalogue read that fails — the request proceeds unchanged (and the deprecation gate, which fails open on its own read, decides).

Two consequences of "the successor is what the request now names": a plan's model allowlist and a gateway's disabled-model list are checked against the successor — a request for qwen3.6-27b on a gateway where the admin disabled qwen3.8-27b is answered 403 model_disabled_on_gateway naming qwen3.8-27b (with X-AIG-Model-Upgraded-From on the response), and a plan that lists only the retired id refuses the successor by name. At a rollover, update plan allowlists and gateway blocks to the successor id.

The exact-match response cache is keyed on the id as requested, so qwen3.6-27b and qwen3.8-27b are separate cache entries; a cache hit on the old id serves the stored answer without dispatch and therefore without the two upgrade headers (the entry records the model that produced it).

Stored picks are covered too: a project default model, an agent's model or a conversation's pinned model that names a retired id resolves to the successor at dispatch time through the same resolver (the stored value itself is not rewritten — the conversation keeps showing the retired id until the user picks again, while every turn it sends is served on, logged as and billed to the successor), and /compact (its summary is recorded under the model that wrote it), the save-time pin checks and a scheduled task's plan check use that resolver too, so they agree with what an inference request would serve. A soft pin — a project default model without a gateway, or a user's personal default model — that names a retired id is resolved to the successor when a new conversation is created (a retired Gemma pin lands on the Gemma successor, not on the Auto pick); that ladder honours a block recorded under the old name on a gateway (a sibling gateway without the block serves it — gateways stay independent; with no such sibling the pin degrades to the Auto pick, as an unmapped pin does) as well as the successor's own availability. Note for the chat UI: a conversation whose stored pin is a retired id shows the picker label Auto (the catalogue has no row for it) while its turns are served on, and labelled with, the successor; picking a model again stores the live id. The pre-rollover model unavailable recovery bubble does not appear for a mapped id — the transparent upgrade replaces it.


Empty responses & inline dead-turn retry

A model can finish a turn cleanly (finish_reason: stop) yet produce no visible answer — 0 output tokens, a few tokens of internal reasoning with nothing user-facing, or a handful of whitespace-only tokens (the instant-EOS "dead turn" some self-hosted models emit on certain prompts: a \n\n preamble, then stop, before any real token). A visible answer that is empty or contains only whitespace (matched against the same whitespace set as JavaScript's String.trim(), so a \n\n or an ideographic-space answer counts as blank) is treated as no answer; any substantive character (e.g. a terse 4 or Ja.) is a real answer and is delivered unchanged. The gateway substitutes a short fallback message in content and records the event for the per-release empty-rate metric. The wording is chosen on the response's output_tokens (the completion-token count, which for a reasoning model includes its reasoning tokens): > 0 → an empty-answer notice; == 0 → "The model returned an empty response. Please try again or rephrase your prompt." For the > 0 case the notice is further split by whether the turn actually produced reasoning content: a reasoning turn that spent its budget on hidden reasoning gets "The model produced internal reasoning but no visible answer. Try rephrasing your prompt or turning off Reasoning…", while a thinking-off turn that simply emitted no visible content (the instant-EOS dead turn — no reasoning at all) gets "The model ended the turn with no visible answer. Please try again or rephrase your prompt." — the "turn off Reasoning" advice is dropped because reasoning was never on. On a streamed OpenAI-compat response the count is only final after the trailing usage chunk, so the classification is made once the whole response has been read. The fallback applies whether or not tools were offered on the turn: a first-party conversation turn (where the file tools are offered whenever the model can call tools) that finishes with a whitespace-only answer gets the same fallback as a tool-free request. It is the leg's terminal fallback — it is not added to a leg the turn is about to continue or recover: a tool call the loop will run, a length-cap continuation, the broken-promise nudge (a file was asked for and none produced — the nudged leg answers; a second blank leg, or a nudge that never runs because its dial failed or the loop's budget ended the turn first, gets the one fallback — after any earlier leg's text), or a turn that already produced a deliverable (the bare re-download card, or a file / image / code-interpreter artifact earlier in the loop). On a streamed turn a blank leg whose read failed after its finish chunk still gets it (nothing else covers that bubble); on a buffered turn that cut attempt is re-dialled or reported like any other failed attempt (see the buffered-failure rules below), and a leg that ended with a provider error event gets that banner, not this fallback. Three tool-related shapes get something else: a leg whose visible content is empty because an inline tool call was stripped and recovered continues the loop into the real answer; a leg whose only content was a tool tag the gateway could not turn into a call gets the "unparseable tool" retry hint below; and a leg whose only content was a stripped non-tool tag (a <memory> proposal, which was delivered as its own event) gets no wording at all — the model did produce content, so "no visible answer" would be false (the gateway logs the strip). A turn the client aborted is recorded as aborted, not as empty.

A third case takes priority on a max_tokens (length-capped) empty turn: when the prompt itself ~filled the model's context window (prompt_tokens ≥ 90% of the routed model's known context window), the turn produced no answer because there was no output budget left. The gateway returns a distinct fallback whose wording follows the turn count of the request — on a first turn (exactly one user message in the request, so the prompt is the conversation) "Your input is too large for this model's context window. Shorten it, remove attachments, or pick a model with a longer context."; on any later turn "This conversation is too long for the model's context window. Start a new chat or shorten the conversation, then try again." — and, because a same-prompt retry would re-send the near-window prompt and deterministically fail again (and re-bill the input), suppresses the retry — the inline dead-turn retry below never fires for this class, and the chat UI hides its regenerate affordance on the bubble. When the routed model has no known context window (the catalog lacks max_input_tokens) this class is not selected — the turn falls back to the reasoning-only / empty wording above.

For self-hosted (inline) models a "dead turn" — an empty-visible clean finish with ~0 output tokens (an instant-EOS glitch) — is transparently retried once on a buffered (stream: false) request before the fallback is returned. If the retry produces a real answer you receive it normally; if the turn is still dead the fallback is returned and a provider fault is recorded (and, for a sustained storm of dead turns, the provider's circuit breaker opens so traffic fails over). The retry is capped at one attempt and never applies to third-party providers. A dead turn is an instant end of turn (finish_reason: stop): a max_tokens (length-capped) empty turn is not one — the model was still generating when the cap cut it, typically because the caller asked for a tiny max_tokens (a one-token health probe) or, on a model with no known context window, because the prompt filled it. Such a turn is delivered once with the fallback above, is never retried (the identical request would end identically), and counts as neither a success nor a failure for the provider's circuit breaker. Live-streaming (stream: true) turns are past the point of no return once bytes are on the wire and are not retried.

Broken-promise recovery (Workspace file writes)

On a Workspace turn where the server-side file tools are offered, a model can narrate an action it never performs: you ask it to create or write a file, and it replies "I'll now create the file…" and stops — a clean finish_reason: stop with prose but no write_file call and no file produced. Because a clean stop with visible text is the tool loop's normal successful exit, the turn would otherwise end there with nothing written. When your typed request was to write a file and the turn produced no file and called no tool, the gateway re-drives the turn with a short instruction to use the tool now. The write intent is read from the text you typed — the contents and filename of an attached document are excluded, so attaching a file whose text happens to contain a word like "generate" does not trigger the recovery. Re-drives stop as soon as either a file is written or a nudged attempt again makes no progress (answers with no tool call), so a turn that has genuinely finished is never re-driven repeatedly; no answer field of a standard response is altered. This recovery applies only to Workspace/chat turns (identified by their conversation and turn ids); a plain OpenAI-compatible client is never re-driven and always receives exactly one finish_reason.

Streaming

Set "stream": true to receive an SSE response. The gateway emits OpenAI-shaped chunks and terminates the stream with a data: [DONE] sentinel.

data: {"id":"...","object":"chat.completion.chunk","created":1753718400,"model":"...","choices":[{"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"...","object":"chat.completion.chunk","created":1753718400,"model":"...","choices":[],"usage":{"prompt_tokens":676,"completion_tokens":127,"total_tokens":803}}

data: [DONE]

What a plain client receives

A plain OpenAI-compatible client is one that sends neither x-aig-turn-id nor x-aig-extensions: 1. Every data: frame it receives is either [DONE] or a JSON object carrying choices (an array — empty on the final usage chunk) or an error object. Nothing else. This is a hard contract: clients that validate each chunk against the OpenAI schema (the Vercel AI SDK's @ai-sdk/openai-compatible provider and everything built on it) abort the stream on the first non-conforming frame.

The gateway's own aig_* extension frames — aig_status (all values, including thinking), aig_tool_call, aig_tool_input, aig_sources — are suppressed for such a client. They are a first-party side channel for the chat UI, described throughout this page; a client opts into them explicitly:

Signal Who sends it Effect
x-aig-turn-id: <id> the chat UI, every turn turn ownership and the side channel
x-aig-extensions: 1 a first-party client with no turn (agent Try-It, the playground) the side channel only

Accepted values for x-aig-extensions are 1, true, yes, on — matched exactly, so TRUE and On do not count. Any other value is legal to send and is treated as absent (fail closed): 0, an empty value, an unknown word, a non-string, or a repeated header whose first value is not one of the four. The stream then stays pure OpenAI.

Turn persistence is tenant-fenced

A turn-owning request (x-aig-turn-id) that also names a Workspace conversation with x-aig-conv-id has its assistant answer persisted into that conversation. Both headers are client-supplied, so the conversation named by x-aig-conv-id is not trusted as an authorization grant: a conversation is pinned to the tenant of the gateway that owns it, and the gateway persists an answer into it only when that conversation belongs to the same tenant as the gateway the request routed through (the tenant is derived server-side from the gateway, never from the request body or a header).

  • Accepted — the conversation is in the caller's tenant. This includes a conversation reached through a different gateway of the same tenant (a platform admin routing across the tenant's gateways); the answer persists normally.
  • Rejected — the conversation belongs to another tenant. The turn still runs and the answer is returned to the caller (it was a legitimate request on the caller's own budget), but nothing is written into the foreign conversation: no assistant message and no streaming placeholder row. From the caller's tenant that conversation does not exist, so the gateway treats it as a ghost — no error status, the answer simply is not saved anywhere the other tenant can see. This closes a cross-tenant write (an attacker cannot inject a chosen answer, or a stalling placeholder, into someone else's thread).

Within a tenant, a second check — the same write-permission rule the Workspace UI applies — decides whether the answer is saved into the named conversation:

  • Accepted — the caller is the conversation's owner, or a member of the project it is shared into who currently holds the live-driver turn (the collaborative co-editing lock). This also covers a server-produced artifact (an image or generated file) persisted under an absent or non-conforming turn id.
  • Rejected — a same-tenant conversation the caller may not write: not the owner and not the current driver (for example a read-only shared member, or someone with no access at all). As with the cross-tenant case the answer is still returned but nothing is written — no assistant message and no streaming placeholder — so an attacker cannot inject an answer into, or stall the reload of, a colleague's private (or read-only) conversation.
  • A service token with no user identity persists into no conversation on this path.
  • The driver lock is evaluated at the start of the turn (when the placeholder is claimed), not re-evaluated at save time — so a long-running turn whose lock lapses mid-answer is still saved, never dropped. A transient database error while checking permission is surfaced as a retryable failure (the answer is reported as not-yet-saved), never silently discarded.

Consequences worth planning for as a plain client:

  • Errors arrive as frames, not statuses, once the stream is open. A failure detected after the response head is committed is delivered as data: {"error":{"message":"…","type":"gateway_error","code":"<code>","param":null}} followed by [DONE] on an HTTP 200 — the full official OpenAI error shape (message, type, code, and the nullable param; the gateway never attributes an error to a request parameter, so param is always null). Before the head is committed a failure is still an ordinary HTTP error response. The terminal chunk of such a stream stays inside the finish_reason enum (stop), so treat the error frame, not the terminal, as the failure signal.
  • A long reasoning or tool phase carries no data: frames. The thinking frames that used to fill those gaps are extension frames. To keep the connection alive the gateway writes an SSE comment line (: hb) after 15 s of silence — never while content is flowing. A comment is ignored by every conformant SSE reader (EventSource, @ai-sdk/openai-compatible, openai-python); a reader that JSON.parses whole blocks instead of following the SSE framing must skip lines that do not start with data:.
  • Exactly one data: [DONE]. The stream is terminated once; nothing follows it.
  • Exactly one non-null finish_reason — even across a server-side tool loop. When the gateway runs a tool for you (an MCP connector, web search, the code interpreter, file writes), it may call the model several times in one response. Only the final leg's terminal reaches the wire; the intermediate finish_reason: "tool_calls" legs, where the gateway executes the tool and continues, are suppressed. So a client that treats the first non-null finish_reason as the end of the turn still gets the whole answer. Server-side tool execution is the only thing that runs several legs for a plain client — the length-cap auto-continue of a long answer (below) is a first-party feature a plain client never triggers, so it never adds a second finish_reason here. This is distinct from a turn where you supplied the tools (tools in the request body): there the gateway forwards finish_reason: "tool_calls" to you as the legitimate, single terminal — you run the tool and continue the conversation with a new request.

created and system_fingerprint

Every chat.completion.chunk carries created — the completion's creation time as a unix timestamp in seconds. Every chunk of one response carries the same value (it stamps the completion, not the frame), and the buffered chat.completion body carries it too, so a re-serialising proxy (LiteLLM and similar) accepts either shape.

A response served from the gateway's response cache replays the body it stored, so its created (like its id) is the original completion's — not the time of the cache hit. The X-AIG-Cache: HIT response header tells you which case you are in.

system_fingerprint is deliberately never sent. It identifies the backend configuration a completion ran on; a gateway routes one model id across heterogeneous providers, regions and adapters, so there is no value it could report that would stay true. The field is optional in the OpenAI schema — treat its absence as intentional rather than as a gap.

The final usage chunk

The last chunk before [DONE] carries token usage with an empty choices array:

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":1753718400,"model":"…","choices":[],"usage":{"prompt_tokens":676,"completion_tokens":127,"total_tokens":803}}

One deliberate difference from OpenAI: OpenAI sends this chunk only when you ask for it (stream_options.include_usage: true); the gateway sends it by default and honours stream_options.include_usage: false as an opt-OUT.

A turn can run several upstream legs — a server-side tool loop, or an auto-continued long answer. A plain client receives exactly one usage chunk per response, carrying the turn total summed across every leg, as the last frame before [DONE]. You may assign it or add it; both give the same number. Because the gateway sums it rather than relaying one leg's block, provider-specific breakdowns that do not add up across legs (prompt_tokens_details, reasoning_tokens) are not carried; Anthropic's cache_creation_tokens / cache_read_tokens / cache_deletion_tokens are, when non-zero. The upstream's prompt_tokens_details.cached_tokens is consumed, not forwarded: the gateway nets it out of prompt_tokens and books it as cache_read_tokens instead (so the summed chunk reports prompt_tokens = the uncached prompt plus a cache_read_tokens bucket, the same shape an Anthropic-origin turn already reports on this envelope). Only a finite non-negative cached_tokens no greater than prompt_tokens is accepted; a null / non-numeric / NaN / ±inf / negative value is treated as no cache hit (cache_read = 0), and a value exceeding prompt_tokens is clamped to it.

A client that opted into the gateway extensions (x-aig-turn-id / x-aig-extensions) instead receives one usage chunk per leg and must SUM them for the turn total — that cadence exposes the leg structure first-party tooling reports on. It is the only place the two audiences see different bytes at all; finish_reason, created and every other standard field are identical for both.

The gateway normalises each provider's stop reason to a consistent OpenAI-compatible finish_reason:

finish_reason Meaning
stop Normal completion (Anthropic end_turn/pause_turn, Gemini STOP, Cohere COMPLETE also map here), and the fallback for anything the gateway cannot classify.
length Output truncated at the token limit (OpenAI length, Anthropic/Gemini/Cohere/Bedrock max_tokens/MAX_TOKENS, Mistral model_length map here).
tool_calls The model requested a tool call (provider tool_use).
content_filter The model refused or was blocked (Anthropic refusal, Gemini SAFETY/RECITATION/PROHIBITED_CONTENT/BLOCKLIST/SPII, Cohere ERROR_TOXIC, Bedrock guardrail_intervened/content_filtered).

That enum is closed, and it is the same for every client and every egress — the streaming wire and the buffered chat.completion body, whatever headers you send. An unrecognised or future provider stop reason is mapped, never relayed verbatim, and a stream aborted mid-flight reports stop with an accompanying error frame (below) rather than inventing a value. function_call is accepted as the legacy enum member but the gateway never mints it.

This matters most for agents that continue automatically when a response was cut off: they key on length. A value outside the enum does not make a strict client abort — it makes such an agent stop silently, with a half-finished answer and no error.

Internally the gateway keeps a richer terminal vocabulary (that is what drives server-side auto-continue, the truncation indicator, and the rule that an abnormal turn is never persisted as a durable summary). It is deliberately not on this field: a standard OpenAI field carries the OpenAI value, and the gateway's own reason travels on the aig_* side channel, where extensions belong.

This normalisation now covers every provider. Gemini/Vertex, Cohere and Bedrock previously reported no terminal reason at all, so a truncated or content-filtered turn on them was mislabelled a clean stop: the "response was truncated" indicator never appeared, auto-continue never armed, and a truncated compaction summary could still supersede the conversation history. Their parsers now surface the real terminal, so these behave like OpenAI/Anthropic.

On the Anthropic-native response shape (/anthropic/v1/messages), the stop_reason field is always one of Anthropic's own enum values (end_turn, max_tokens, stop_sequence, tool_use, refusal) even when the request is routed to a non-Anthropic model — a foreign provider's terminal (Gemini MAX_TOKENS, Cohere COMPLETE, Bedrock guardrail_intervened) is translated to the matching Anthropic value, and an unrecognised foreign reason resolves to a clean end_turn rather than being emitted verbatim. A genuine Anthropic leg's stop_reason (including pause_turn and any newer value) is relayed unchanged.

A long answer that fills the model's context window is auto-continued across legs; if a continuation leg cannot proceed (the input plus the answer so far already exceeds the window), the turn ends truncated, and the gateway terminates the stream with finish_reason: "length" — never a clean "stop" — so the client can surface that the response is incomplete. Route context-heavy jobs to a larger-window model to avoid truncation.

A turn truncated at the model's own output limit is reported as length on every egress — the live SSE stream, the buffered response body, and the buffered re-emit (below). Web-search turns are included: a search-augmented answer that hits the limit is labelled truncated, not clean.

Buffered re-emit turns carry the same tool telemetry

A gateway whose only response detector is an inline PII token-restore masker (pii_protector / custom_pii) streams incrementally on both the native and the compat (OpenAI-format) wire: PII tokens are restored on the live wire (visible answer, reasoning, tool-call args, grounded-excerpt quotes, filenames, web-search queries all carry real values), so these turns are not internally buffered and their tool-loop telemetry streams live. (On a wholly-local Myra/EU model leg no PII scan/restore happens — masking is skipped; see PII Protector — Local model legs are not masked.)

A "stream": true turn is still force-buffered internally when a response detector the live stream cannot satisfy is active — a content-safety block, or a regex/presidio scrub (mask) — or when the turn runs buffered by nature (a stream: false tool loop). On that path the whole answer is produced, scanned, and then re-emitted to the client as a normal SSE stream. That re-emit replays the turn's tool-loop side-channel events — aig_status values tool_call, tool_result, tool_deferred, tool_error, and ephemeral_file — in their original order, before the answer content. A PII-protected turn therefore carries the same tool telemetry as a live stream — to a client that opted into the aig_* side channel; a plain OpenAI-compatible client receives it on neither path. With that telemetry a client can tell a genuinely computed code_interpreter answer from a model estimate, render tool chips, show file-write notes, and receive a live-delivered generated file.

Bounds (the replayed frames originate from model-driven tool calls and are treated as untrusted): at most 256 events are replayed per turn, and an unencodable frame is dropped. An event whose encoded size exceeds 64 KiB is replayed with its args removed and "args_truncated": true set. ephemeral_file is exempt from both the size check and the encodability check — its payload is the user's file and every field of it is gateway-constructed; its size is bounded at the source instead (see below). Statuses the buffered tail already reconstructs from its own state (thinking, pii_masked, tools_skipped, aig_sources) and mid-stream control signals (provider_error, compacted, pause_turn) are not replayed — no event is ever delivered twice, and no stale control signal outlives the turn it steered. Plain non-streaming requests ("stream": false) are unaffected: their single JSON response never contains side-channel frames.

When the buffered answer is a tool call — the model returned tool_calls (for example a client-provided tool, or a web-search direct answer that called the client's own tool) with no visible text — the re-emitted stream carries those tool_calls as a normal OpenAI delta.tool_calls chunk followed by a terminal finish_reason: "tool_calls", exactly as a live stream would. A streaming client is never handed a silent empty turn on a tool-call answer. (If the tool call's arguments were truncated by the token limit the terminal is length, not tool_calls — the truncation signal wins. An empty or non-array tool_calls is not emitted.)

ephemeral_file — a generated file delivered live and stored nowhere

When the code interpreter produces a file in a conversation that cannot store files — a private (ghost) chat — including one opened inside a project, where the gateway refuses to write to the project's knowledge base — or any request without a conversation to write to — the file is not persisted and there is no download route to point at. Instead the gateway sends the bytes once, inline on the turn's own stream:

data: {"aig_status":"ephemeral_file","name":"report.xlsx","mime_type":"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet","size":8134,"data":"UEsDBBQ..."}
Field Type Meaning
name string Bare filename, at most 100 characters. Never a path, never a leading dot, never a Windows-reserved character (: * ? " < > |) and never a control byte.
mime_type string One of the 11 accepted types: text/csv, text/plain, application/json, the .xlsx workbook type, image/png, Word (.docx), PowerPoint (.pptx), PDF (application/pdf), and OpenDocument (.odt, .ods, .odp).
size integer Size of the raw (pre-base64) bytes.
data string Base64 of the raw bytes.

Nothing is written server-side for such a file: no conversation row, no knowledge row, no stored blob, and no id — the event is the file's only existence, and it is gone once the client discards it.

Opt-in. The event is sent only to a client that identified its turn with the x-aig-turn-id header (the chat UI always does) or opted in with x-aig-extensions: 1. A plain OpenAI-compatible client that sends neither never receives it, and the generated file is reported as not shown instead. This is a delivery contract, not an authorization boundary.

Bounds. A turn delivers at most 4 MiB (4 194 304 bytes) of raw bytes this way by default — the effective per-turn cap is the tenant's code-interpreter artifact-size limit, which an admin can raise up to a 25 MiB ceiling — and at most four files plus four images (the same per-turn budgets the stored path uses). An artifact that does not fit, or that cannot be delivered (for example on a non-streaming request), is reported to the model as "could not be shown to the user" rather than silently dropped.

Client-side validation. A consumer must treat this event as untrusted, exactly like every other stream frame. The chat UI rejects — and renders nothing for — an event whose name contains a path separator, a control character, or a leading dot, whose mime_type is outside the list above, whose size is not a positive integer within the budget, whose data is not plain base64, or whose size and data length disagree (padded base64 length is a function of the raw length, so the two must match). The filename rule is the same list the gateway applies before emitting — a path separator, a Windows-reserved character, a control byte, a leading dot, or a name over the byte ceiling is REJECTED, never "repaired" into a different file, and yields no download card. Because both sides apply the same list, an artifact the gateway delivers always renders; one whose name would fail it is never delivered at all — the model gets the honest "could not be shown" note instead. Cosmetic trimming (collapsing runs of whitespace, stripping trailing dots) happens only at render time and changes how the name looks, never which bytes are written.

Tool-loop termination note on the buffered egress

When the gateway's own tool loop stops a turn early — the per-turn tool-call limit, a detected tool-call loop, an exhausted time budget (including the "answer is based on partial results" label when the gateway synthesized a partial answer over the results gathered so far), or a sub-agent deadline — it appends a short human-readable note to the answer. On a streaming turn that note is relayed live as its own content delta.

On the buffered compat egress (a "stream": false request, or a PII-force-buffered "stream": true turn re-emitted as SSE) the note is included in the delivered response body's message content — choices[0].message.content on a direct non-streaming response, and the re-emitted content on the PII path — so a buffered client sees the same stop/partial label a streaming client does, on the first response, matching the persisted turn (a reload shows the same text). The note is appended exactly once and never onto a tool_calls body, an error envelope, or a structured-output (response_schema) answer. A native /anthropic/v1/messages direct non-streaming response gets the note folded into its last text content block as well; only the native PII-force-buffered turn is excluded — it is re-emitted through the compat-shaped tail, an existing limitation outside this scope.

Malformed upstream streams (fail-closed)

The upstream provider's SSE bytes are treated as untrusted. A single Server-Sent Events line (the text between two newlines) has an internal size ceiling of 1 MiB. Legitimate event lines are far smaller — token deltas of at most a few kilobytes, and providers split even a large tool-call payload across many newline-terminated delta events — so this ceiling never truncates a well-formed stream regardless of the total response size (only the size of one un-terminated line is bounded). A broken or hostile provider that streams bytes without a terminating newline past the ceiling is cut off rather than allowed to grow gateway memory without limit: the affected leg fails closed. On the compat (unified) path the stream terminates with a final chunk (emitted when no terminal chunk had yet been sent) and then ends without a data: [DONE] sentinel — the sentinel is suppressed on an errored stream. A plain client receives the failure the way the OpenAI schema carries one: a data: {"error":{"message":…,"type":…,"code":…,"param":null}} frame before that final chunk, whose own finish_reason is the in-enum stop (error is not an OpenAI enum member, so it is never written to this field). The code is stream_truncated when the upstream closed early and provider_error when the connection itself failed; when the provider announced the failure with its own error event, the code is the gateway's classification of it (a rate-limit or overload class) and type is the provider's own error type, so an SDK's retry policy can read it. This frame is delivered to every client, first-party included — once the terminal has to stay inside the enum, the error object is the only thing distinguishing a truncated turn from a short one, so no audience may be left without it. A client that also opted into the extensions still gets its aig_status:"provider_error" banner event.

Such a leg is not re-dialled on the same provider. Both the per-line ceiling above and the per-leg total stream ceiling (64 MiB) are properties of the request/response pair, not of the moment: an identical retry downloads the same oversized response and fails identically, while the provider generates — and bills for — the output again. The gateway therefore skips the remaining same-provider attempts and moves straight on to the next provider in the routing chain, which may well be a model that answers. Only if the whole chain is exhausted does the request end in all_providers_failed, and the recorded cause names the ceiling that was hit rather than a bare stream error. Transient mid-stream failures (a dropped connection, a read timeout) are unaffected and keep their full retry budget. Models that return generated images inline in the stream are the common way to hit the per-line ceiling — their base64 payload arrives as one un-terminated line — and such a model cannot be delivered on this path at all; pick a text model, or one whose images are returned by reference.

One case is deliberately not treated as a truncation: an Anthropic-style two-phase terminal where the model's message_delta already carried a natural stop_reason (end_turn/stop_sequence/pause_turn) and only the trailing message_stop was lost to a socket close. That answer is complete, so the gateway emits neither the stream_truncated error frame nor the visible note below — only the finish chunk (still the in-enum stop). A genuine cut-off, where no terminal stop_reason ever arrived, is unaffected and still gets both. (A max_tokens terminal is a real length cap, not a natural end, so it keeps its length truncation signal.) On the native passthrough path the raw stream simply ends without a terminator (which conforming clients treat as an error). In both cases the upstream connection is closed rather than returned to the pool, and any pending tool call the truncated leg had begun is discarded (never executed).

Visible truncation note (persisted), so a cut-off is never silent to a human. Because the wire terminal reads the in-enum stop, a partial answer would otherwise render — live and on reload — as a clean, complete one. On the compat path the gateway therefore appends a one-line "The response was interrupted before it completed. Please try again." note to the visible answer — as its own content delta after whatever partial text had already streamed (one blank line between the partial and the note; when the partial already ends in a blank line the note follows it directly, and a dangling whitespace-only last line is closed first so the note is never indented into a code block), or as the whole bubble when nothing had — and records it into the persisted turn, so a reload shows the truncated turn the same way it rendered live. The note is gateway-authored English prose (like the tool-loop termination note above); it is not localized. The chat UI additionally raises a non-fatal response may be incomplete banner off the error frame's code. A turn that instead ends with an actionable provider-error banner (aig_status:"provider_error") or a context-overflow recovery banner gets neither the generic note nor the banner — those surfaces already name the failure. This note follows the same rules as the tool-loop termination note above (never onto a tool_calls body, an error envelope, or a structured-output answer). It is the same on the buffered / PII-force-buffered egress: a turn the gateway buffers internally (a PII-masking or guardrail configuration, or a stream: false tool-loop turn) whose upstream dies after partial bytes delivers the salvaged partial with the note and the same failure surface for its egress — the terminator error frame on the re-emitted SSE, X-AIG-Error on a JSON body (see the Mid-stream provider failures on a buffered turn callout under Anthropic-native passthrough) — and a conversation-owned turn persists exactly those bytes; it used to persist the bare partial as a clean-looking complete answer. The one verdict is shared: a leg that reached a natural end and only lost its message_stop is complete on both egresses (no note, no frame), and a provider_error event gets its banner and no note on both. A leg that will continue in the tool loop gets no note — and only a leg that finished cleanly (its finish event arrived — the finish_reason chunk on the OpenAI wire, message_stop on Anthropic — with no read error and no provider error) with pending tool calls continues. A leg that did not finish cleanly is dead: nothing that would make the turn continue is synthesized from it (no inline <tool_call>, no unwrapped file tool, no file-claim steer), structurally delivered tool calls are discarded (Anthropic tool_use blocks included — even after message_delta, if message_stop never came), the gateway logs and traces dead_leg_no_continue, the note is appended, and the error frame is terminal — Regenerate is the remedy. (A leg that reached a natural end and only lost message_stop on an otherwise clean close is complete for the wire and the row — no note, no error frame, its finish chunk emits — but equally never continues: nothing may follow that terminal. If such a leg carries a tool call its death cost it — one the gateway would have run had the finish event arrived: a recovered <tool_call> marker, an unwrapped <write_file> / <read_file>, a write_file sentinel, or a structurally delivered call (an Anthropic tool_use block, an OpenAI tool_calls delta) — it is not treated as complete: it gets the note and the error frame, because its prose is a preamble to work that will never happen. Text the gateway discards on a clean finish too (an anchored bare-JSON or pseudo-call naming write_file / read_file, which the inline channel never dispatches) does not change that verdict — the answer stays complete — but the discard is always logged, so a dropped write is never silent. A pending bare-re-download card is not such a marker — on a streaming turn, and on a PII-force-buffered turn (whose captured card the buffered re-emit replays), that card is re-served on any leg that neither hit a read error nor carries the interrupted note (a leg that ended in a provider error does get it), so a re-download does not go silently missing; a leg that carries the note re-serves no card on either egress — a card under an "interrupted" note would contradict it, and Regenerate re-serves the file — and on a non-streaming (stream:false) turn the card has no wire to ride, as before.) So a partial on a tool-capable turn (the common first-party conversation turn, where file tools are always offered) is a cut-off like any other and carries the note, live and on reload. The same holds for an API turn that supplies its own tools: if the provider dies after tool_calls deltas but before the finish chunk, those calls are discarded and the note is appended to the visible answer — the turn is a cut-off, not a tool turn, and no tool_calls body is emitted for it.

The client's own connection dies mid-stream (a laptop lid, Wi-Fi gone at 80 % of a long answer). The gateway sees the client abort, stops reading the provider at once and closes the provider connection (the model stops generating for a turn nobody will receive), then persists the partial answer it had already streamed once, synchronously, under the turn's id — nothing is lost server-side and nothing is doubled; the request is accounted as a client abort (request_log.aborted = 1, the leg status 499 / error_class cancelled / partial = 1). The chat app treats the read failure the browser raises for a dropped connection as exactly that — one predicate accepts each engine's wording for both moments, before a response and mid-body: Chrome Failed to fetch / network error, Firefox NetworkError when attempting to fetch resource. / Error in body stream, Safari Load failed / The network connection was lost. — and keeps the partial the user was watching on screen as the assistant turn, raises the same non-fatal response may be incomplete banner, re-enables the composer, and adopts the persisted row (which may carry a token or two more than the client received) through the same bounded reconcile a Stop uses — so live and reload show one partial, once. If the network is still down when the reconcile runs, the partial simply stays as rendered — no blank, no "no response received" row — and the next load shows the persisted partial once. (The same failure while nothing had streamed yet is a plain error banner — there is no partial to keep.)

The provider's HTTP body framing. Before any of the above, the gateway frames the provider's response body itself (its own coroutine-free reader over the connection — the receive belongs to the coroutine that is cancelled on a client abort, so an abort really closes the provider connection). Accepted: Transfer-Encoding: chunked with a hexadecimal chunk-size line (as tonumber(line, 16) reads it — the SSE case), a Content-Length body read to exactly that length, or — with neither header — a read-until-close body. Rejected, and surfaced as a read error that ends the leg (never a fabricated length, never a silent end): a chunk-size line that is not hex (a chunk extension such as 1a;name=value, a non-hex token) or is negative; a Content-Length that is not a single non-negative integer (4.5, a duplicated header) is ignored and the body is read until close, so a truncated read can never leave stray bytes on a pooled connection. Sizes are honoured in bounded pieces (at most 64 KiB per read without a caller limit), so a size the socket cannot honour ends the leg as a read error. HEAD / 1xx / 204 / 304 responses have no body and are never waited on.

Structured / malformed content values. A provider's delta.content (streaming) and message.content (buffered) are equally untrusted. Accepted shapes are a plain string, null, or an array of parts; the gateway flattens an array into the visible answer text (the text parts concatenated; the same flattener, with a separator between the parts, also joins a user message's text parts where a route needs a single string — the reply-language detection and the Mistral web-search fold). Array normalization is two-tier: recognized typed non-text parts (a part whose type is a string other than "text" — e.g. a thinking or citation part) are omitted from the visible text silently; garbage shapes (non-object/non-string elements; text-equivalent blocks — type absent, null, or "text" — with a non-string or absent text; blocks with a non-string non-null type; content values that are neither string, null, nor a table with a leading element sequence) are dropped fail-closed with a gateway WARN log (deduplicated per stream leg); an empty array yields empty text silently. The same fail-closed guards apply to the per-part text fields of the native Anthropic, Gemini, and Cohere parsers, and the compat streaming dispatcher additionally drops any non-string delta a provider parser might emit (one WARN per leg) instead of letting it abort the live stream. Before these guards, a structured content part could kill a streaming turn mid-leg with no finish event — the user saw a silently dead chat.

Source citations (aig_sources)

When an answer draws on an uploaded document, the gateway emits a structured, deterministic source list as a dedicated SSE chunk on the final leg of the turn — built from the turn's actual document-retrieval calls (reading a file by name or the search_knowledge hybrid search over the project corpus), not from model text:

data: {"aig_sources":[
  {"file":"budget-report.pdf","loc":"p.1","kind":"page"},
  {"file":"budget-report.pdf","loc":"p.2","section":"Section 3.1","kind":"page"}
]}
Field Type Meaning
file string The document the answer drew on (required, non-empty).
loc string The locator that was retrieved, e.g. p.2 (required, non-empty).
section string Section/heading within the document, when known (optional).
kind string Locator kind — page, slide, sheet, para, line (optional).

The list is deduplicated by file + loc + section, ordered by first retrieval, and reflects only the portions actually delivered to the model (a document truncated to fit the context window does not contribute citations for the pages that were cut). It is bounded — at most 100 entries, each field at most 300 characters — so a many-page document cannot produce an unbounded list. A turn that retrieved nothing emits no aig_sources chunk. Consumers must treat every field as untrusted input and ignore malformed entries. Like the other aig_* gateway-extension events, aig_sources is emitted only in streaming mode; a non-streaming ("stream": false) response returns a plain chat.completion JSON object with no citation channel.

Only strict-verifiable retrieval is cited. Segments extracted by OCR (confidence='ocr') are served to the model as document text but are not cited — an OCR page may be mis-located, so the gateway never emits a deterministic page/section claim it cannot stand behind.

Citations persist. The same list is stored on the assistant message, so the citations rehydrate when a conversation is reloaded (not just for the live stream). The persisted list is re-validated through the same bounds above before it is returned, and a turn that produced no answer stores no citations (a citation attributes an answer).

fetch_url / agentic_fetch pages range (untrusted model input). For a long PDF truncated to its first pages, the model may re-fetch the same URL with an optional pages argument to read a specific window. Accepted shape: a string — a 1-based range "START-END" (e.g. "110-125") or a single page "117"; surrounding whitespace is tolerated. The window is [START, END] and its span is clamped to the 200-page extraction cap. Rejected / ignored (→ read from the first page, never an error): a non-string value (number/table/boolean), an empty or non-numeric string, a reversed range ("10-2"), or a zero/negative page ("0-3"). ABSENT (omitted) and every malformed shape share this identical first-page fallback — the parse is fail-closed and never widens the read. pages applies to PDFs only (DOCX/XLSX/web pages ignore it). A range that begins beyond the document, or hits only text-less pages, returns a served notice ("Requested pages start beyond … the N-page document"), not an extraction failure.

💡 Note: A guardrail block in streaming mode returns HTTP 200 with a synthetic SSE error chunk (the refusal text as a normal assistant delta.content chunk) followed by data: [DONE], rather than a non-200 status. This wire shape is unchanged for external API clients — integrate against the assistant message as before. See Error codes — guardrail_blocked.

First-party turn-owned streams (the web app) additionally receive a {"aig_status": "guardrail_blocked", "error_class": "<class>"} gateway-extension event so the app can render a distinct policy-block notice instead of a normal reply. Like every aig_* event it is emitted only in streaming mode and only on a turn-owned stream; a plain API client never sees it (see What a plain client receives). error_class is one of guardrail_blocked (content policy), guardrail_pii_mandatory, guardrail_pii_media, guardrail_role_model, or guardrail_unavailable (treat the set as open) — a stable, non-sensitive slug (it never carries the specific policy sub-category or any request content). On a turn-owned stream the refusal delta.content is a generic, sub-category-free message (the specific category stays in the gateway logs); external clients keep the detailed prose.

🔒 File attachments on a PII-protected gateway: because PII guardrails scan content as text, a gateway with an active PII protector rejects a request that embeds inline base64 file media (an image_url/video_url data: URL, an input_audio/file block, or an Anthropic-native image/document with a base64 source) with the block reason pii_media_unmaskable — such binary would otherwise reach the provider unmasked. Attach files through the Chat view, which extracts each PDF/image to text server-side first so the text is masked before it leaves the gateway. Text-extracted blocks and Files-API file_id references are accepted. A request served wholly by a first-party local (Myra/EU) model is exempt from this block — there is no third-party egress to fail closed for. See Data protection — file attachments on a PII-protected gateway.

♻️ Model-unavailable recovery (extensions-enabled streams): when the pinned model is no longer routable (deprecated/renamed/deleted upstream), the gateway ends a STREAMING turn from a client that opted into the side channel with HTTP 200 and a single aig_status recovery event instead of a cryptic provider error (a plain OpenAI-compatible client, and ANY non-streaming request, gets the typed model_not_found HTTP error instead — see the closing sentence of this note): {"aig_status": "model_unavailable", "old_model", "old_display", "fresh_model", "fresh_display", "fresh_gateway_id"} when a replacement default exists, else {"aig_status": "no_default_available", "old_model", "old_display"}. The suggested (fresh_gateway_id, fresh_model) is policy-legal by construction: it is validated against the conversation's permission tier and PII posture with the SAME predicate the conversation-route PATCH guard enforces (see Route policy validation), so the server never proposes a route the app would then be refused. If the only available default would be illegal for the conversation (a cloud model for a local_only project, or a non-scrubbing gateway under a PII mandate) — or the tier can't be proven during a transient fault — the event fails closed to no_default_available rather than offering an illegal target. Like every aig_* event it is emitted only to a client that opted into the side channel (x-aig-turn-id / x-aig-extensions: 1) and asked for a stream; every other caller — a plain OpenAI-compatible client, or any "stream": false request — receives the typed HTTP error model_not_found (400) instead of an SSE body it did not ask for. That error names the model id — e.g. the model '<id>' is not available on this gateway. (the exact wording varies by path) — so an integrator can see which id was wrong without reading the gateway's logs.

What counts as "no longer routable": any 404 from the upstream on the chat-completions call, whatever its body says. It used to depend on the provider's prose — the same fault classified three different ways depending on whether the upstream happened to name the model in its error text, so a {"message":"not found"} body was reported to the user with the generic "The request could not be completed.", the same sentence the product uses for a timeout. A 404 is now always a model failure: it is never retried against the same id (definitive, not transient), and it reaches this recovery. A 404 whose body describes a MORE specific fault (a capability gap, for instance) still classifies as that fault — the status is a floor, not an override.


Per-request header overrides

The provider pass-through supports an optional control header:

Header Effect
x-aig-byok-alias Selects which stored BYOK key alias to use for the resolved provider on this request (see Provider key management). Accepted shape: a label matching an alias already stored for the provider on the gateway. Absent or empty → the default alias. An alias that does not exist is rejected with a configuration error — there is no silent fallback to default, and no fallback to a tenant-level key. Applies only to the primary provider; a fallback provider always uses its default-alias key.
x-aig-extensions Opt into the gateway's aig_* SSE side channel (streaming only). Accepted: 1, true, yes, on. Anything else — including 0, an empty value, an unknown word, a non-string, or a repeated header whose first value is not accepted — is rejected and treated as absent, so the stream stays pure OpenAI. Implied by a valid x-aig-turn-id. This is a verbosity switch, never an authorization signal: it unlocks nothing about any turn but the caller's own. See What a plain client receives.
x-aig-image-gen Client opt-in for the generate_image tool (see Image generation). The tool is offered to the model only when the gateway has image_generation.enabled: true (see config reference) and this header carries a plain value other than 0 / false — the web chat always sends x-aig-image-gen: 1. Accepted: any single plain-string value except 0 and false — the match is exact and case-sensitive, so an empty value, FALSE, no or off all count as opt-IN (unlike x-aig-extensions, which is an allowlist). Absent, 0, false, or a repeated header (which arrives as a list, never as a plain string) → the tool is not offered; a non-string can never arm it. On a gateway with image generation disabled the header has no effect. Saved/scheduled agents never carry it.
x-aig-provider-* The x-aig-provider- prefix is stripped and the remaining header is forwarded to the upstream provider only if the stripped name is on a short allow-list — anthropic-beta, openai-organization, openai-project. Any other stripped name (a credential such as authorization, a request-framing header such as content-length / transfer-encoding / connection, or any unrecognised name) is dropped and logged at warn; it never reaches the upstream request. At most 16 overrides / 8 KB total are forwarded; the excess is dropped.
x-aig-knowledge-ids Per-request knowledge-source ALLOWLIST. A comma-separated list of project/conversation knowledge-document ids the assistant may consider for THIS request. Tri-state by presence: absent → all knowledge (the default); present → exactly the listed ids; present with no valid id (e.g. a single - sentinel) → no knowledge. At most 200 ids are read; a longer list is truncated (truncating an allowlist only narrows what the assistant may use). The selection is intersected with — never widens — the caller's already-authorized project/conversation scope, and is enforced for auto-injected knowledge, the read_file tool, and the search_knowledge hybrid-search tool (the model can neither read nor search a deselected document). It does not restrict loading a project file into the code interpreter's sandbox, which is a separate, deliberate by-name action gated only by the caller's access to the file (see Code interpreter). search_knowledge itself accepts a single query string (natural language or exact terms such as a section/form number); a blank query is rejected, and the search is always confined to the caller's own project/conversation and tenant — a document from another project or tenant is never searched or returned. A file the assistant creates during this request (via write_file or the code interpreter) is always readable by that same request's tools, whatever the selection says — the selection describes the pre-existing knowledge, not the request's own output. The selection is part of the cache key, and the semantic cache is skipped when it is present. Saved/scheduled agents do not honour caller headers, so they always use all knowledge.
x-aig-knowledge-exclude-ids Per-request knowledge-source DENYLIST — the complement of the header above, and what the chat UI sends when you uncheck a file. A comma-separated list of knowledge-document ids to EXCLUDE from this request; everything else in the caller's authorized scope stays in, including documents added after the client last looked. Absent → nothing excluded. Unlike the allowlist this header is honoured in full or not at all: if any token is not a document id (internal whitespace included), if the header carries no usable id at all (an empty value, or only the - sentinel), or if more than 200 ids are listed, the request is rejected with 400 invalid_request — silently dropping an id would feed the model a document the caller asked to exclude. There is no "exclude everything" value; use x-aig-knowledge-ids: - for "no knowledge". Both headers may be sent together, in which case the allowlist is applied first and the denylist removes from it; neither can ever widen the caller's authorized scope. Same enforcement points as the allowlist, and the same carve-out for a file created during the request. A request carrying this header is not cached (see Caching). Saved/scheduled agents do not honour caller headers.
x-aig-leg Marks a request the chat interface made on its own behalf, rather than a message the user sent — currently followup_suggestions (the follow-up chips under an answer) or title_generation (the automatic conversation title). Both are ordinary completions and are guardrailed like any other, but they follow the user's turn and would otherwise be indistinguishable from it in the logs: an operator inspecting "the newest request" was inspecting one of these. Log-only — it never affects routing, guardrails, billing or any policy decision, and a value outside the two above is dropped (the request is logged exactly as it would be without the header). The label appears as the leg kind on request_log_legs.kind and on the log row, where the Logs list badges it. Saved/scheduled agents do not honour caller headers.
X-Request-Id Your own correlation id for this request. The gateway always mints its own request id — returned in the response X-Request-Id header, and the id of the request's log entry; a value you send is never used as that id, so it can neither collide with nor suppress the logging of any request. Accepted shape: a single header whose value matches ^[A-Za-z0-9._\-:+]{1,64} anchored at the absolute end of the value — a trailing newline is rejected (the server anchors with \z, not a $ that would also match before a final \n) — it is recorded on the log entry as client_request_id (list and single-entry responses, the CSV export, the tenant export and JSON SIEM events). Rejected (dropped — not recorded, not echoed, the request is served normally and still logged): an empty value, more than 64 characters, any other character, or the header sent more than once. To the caller "sent nothing" and "sent something unusable" look the same — deliberately so, since the header unlocks nothing; the drop is logged server-side at info level (visible only with the error log at info; the shipped configs run at notice). Like the other correlation headers it is read from the first 100 request headers — placed beyond them it reads as absent. Saved/scheduled agents and delegated agent runs never carry it.

Sending the same header twice

HTTP allows a header to appear more than once, and some SDKs do send anthropic-beta on two lines. The gateway's rule depends on what kind of value the header carries, and it is never "whichever copy arrives first wins by accident":

Header kind Repeated Why
List-valued — anthropic-beta, x-aig-provider-*, user-agent The copies are combined into one comma-separated value (a, b), exactly as RFC 7230 §3.2.2 permits and as the front-end proxy already does for the same header. The header's own grammar is a list, so nothing is lost. Dropping the second copy would silently discard a beta flag and surface as an opaque provider 400.
Boolean opt-in — x-aig-web-search, x-aig-image-gen Treated as not set. An explicit : 0 sent twice keeps the capability off. The value is unusable, and an unusable value must never be more permissive than an absent one. In particular, an intermediary that appends the header cannot defeat your opt-out. (x-aig-extensions, a first-party verbosity switch, instead takes the first value on a duplicate.)
Scoping / identity — x-project-id, x-aig-conv-id, x-aig-skill, x-aig-thinking-budget Treated as not asserted, and the request is not cached. These select a permission and residency scope. A duplicated value names no single scope, so the gateway declines to guess: the request runs without that scope rather than with a guessed one, and its response is never written to (or served from) the shared cache.
Correlation only — X-Request-Id, x-aig-turn-id, x-aig-parent-msg-id, x-aig-regen-of-turn Treated as not sent and logged server-side (X-Request-Id at info, the x-aig-* ids at warn); the request is served and logged normally, with the gateway's own request id. These name nothing that decides routing, permissions or caching — a duplicated value is simply unusable as a correlation key, and an unusable value must never be recorded as if it were the one you meant.

A duplicated header never returns 500. Before this rule a repeated anthropic-beta produced an internal error on the inference path that looked like a provider fault.

💡 Note: To pin a specific provider, address it in the URL path (/v1/{tenant}/{gateway}/{provider}/chat/completions) instead of the compat pseudo-provider. The compat endpoint resolves the provider from the model field, not from a request header.

x-aig-provider-anthropic-beta: interleaved-thinking-2025-05-14
is forwarded to the provider as:
anthropic-beta: interleaved-thinking-2025-05-14

See Providers overview — Compat model resolution and Provider header pass-through for the full behaviour.

💡 Note: Additional x-aig-* control headers are accepted for specific features (web search, projects, MCP tools, conversation threading). They are documented alongside the feature they configure.

🔒 Scope of caller control headers. On this OpenAI-compatible endpoint, the caller-supplied x-aig-* control headers are honored as sent — the gateway does not silently drop them here (unlike the saved-agent invoke path, which is owner-scoped and clears caller control headers). This is safe because the token is bound to a single user, and every header resolves against that identity: x-aig-conv-id, x-aig-policy-conv-id and x-project-id are workspace-scoped: they resolve only within the caller's own workspace (an id from another workspace returns nothing), and what they resolve is a policy tier, never message content — no conversation text is loaded or returned through them. MCP connector references (tools:[{type:"mcp",connector_id}]) can invoke only connectors the caller owns or is granted (using the caller's own credentials), and x-aig-web-search merely toggles a capability the gateway must also have enabled and provisioned. No control header can disable PII protection. The client is never the authorization boundary — these headers convey the caller's own intent within the caller's own scope.

🔒 How the egress tier is resolved (residency / PII). A request's data-residency tier (local_only — no external tools/models; pii_mandatory — masking required) governs whether the turn may egress externally, and it is resolved only from the authoritative source, never from a value the client can vary: - A request that carries x-aig-conv-id (a conversation) resolves its tier solely from that conversation's committed project binding, server-side. A value the gateway cannot use — one that does not match the id shape, or the header sent more than once — is not treated as "no conversation was asserted": the turn fails closed, because otherwise a control could be defeated by sending the header twice. The x-project-id header is not consulted for a conversation-bearing request — so a stale or mismatched x-project-id can neither loosen a restricted conversation (no residency leak) nor spuriously restrict a plain chat. If the binding cannot be confirmed for the request, the turn fails closed (external tools/models blocked) and returns the retryable project_tier_unresolved (503), never a false permanent "only local models" — retry resolves it. - An id that names no conversation at all is not an unconfirmed binding, and is treated exactly as if the header had been omitted: the tier comes from x-project-id as below. Conversations are created only through the conversations API, never by an inference turn, so an id with no conversation behind it is a caller's own grouping key (it also keys the PII salt and the prompt cache). This grants nothing — omitting the header reaches the same place — and a restricted conversation is unaffected, because it has a real binding and takes the rule above. - A conversation whose project no longer exists (deleted) keeps failing closed, but the answer is permanent, not retryable: conversation_project_unresolved (403). The retryable 503 above is reserved for a binding that genuinely might resolve on a retry — promising "try again" for a project that is gone is a promise nothing can keep. - A request with no conversation (a raw /v1 caller acting within a project) resolves its tier from x-project-id, server-side and tenant-scoped — the header names a project; the gateway looks up that project's real tier. A cross-tenant or unknown x-project-id yields no tier. - A request that carries x-aig-policy-conv-id resolves its tier from that conversation's committed project binding, exactly as x-aig-conv-id would — but it does nothing else. It exists for the chat surface's own post-turn helper calls (auto-title, follow-up suggestions), which carry real conversation text and so must obey the same tier, but which cannot send x-aig-conv-id because that header also switches on server-side system-prompt composition and would replace the helper's instruction with the conversation's persona, memories and project knowledge. The header names a conversation; the gateway looks up that conversation's real binding, tenant-scoped, and fails closed when it cannot. Accepted shape: the same id shape as x-request-id (^[A-Za-z0-9._\-:+]{1,64}, end-anchored — a trailing newline is rejected). It is ignored whenever x-aig-conv-id resolved a conversation for this request, so it can never override a real conversation's binding. Anything it cannot use — a malformed value, an empty one, or the header sent more than once — is refused (the turn fails closed), never silently skipped: a header the caller deliberately sent must not be treated as one they never sent. - A conversation that provably has no project is a plain chat and is unrestricted; the header is ignored there too.

If the database is briefly unreachable. Tier resolution always reads the database on the happy path, so a tier change takes effect on the very next request — the gateway keeps no read-through cache. It does keep a last-known-good value per gateway node, consulted only when that read fails, so a momentary database blip degrades to the last confirmed answer instead of blocking every project conversation. If nothing is cached, the request fails closed.

That fallback is an availability measure, not a guarantee, and it is worth stating what it does not promise. Moving a conversation between projects, or changing a project's access tier, clears the value on the node that handled the change — but each node keeps its own, so for up to 5 minutes another node can still answer a failed database read from what it last confirmed. Where the tier was loosened in the meantime that merely over-restricts. Where it was tightened — a conversation filed into a restricted project, or a project moved to a stricter tier — that node can treat the conversation as it was before, during a database outage, for up to that window.

Use Local only where that window is unacceptable: it constrains model routing itself, so a turn cannot reach a non-local provider regardless of what any cache answered.

Every such decision is recorded. Whenever a tier is served from the last-known-good value because the authoritative read failed, the request log carries meta.residency_tier_from_cache set to the value that was served — a tier, or no_project for a conversation last confirmed to have none. It is visible on GET /admin/v1/logs/{id} and accompanied by a [residency_tier_from_cache] warning in the gateway log. A healthy resolution is not marked, and neither is a request that failed closed: the marker means exactly "this answer may predate a change we could not read", so a compliance review can enumerate what was let through during an outage instead of inferring it.

MCP connectors from the API

Registered MCP connectors can be used from the inference API by reference — the gateway resolves the connector's tool list server-side, offers the tools to the model, and executes every tool call inside its own tool loop with the caller's stored connector credential (which never leaves the gateway):

{
  "model": "…",
  "messages": [{ "role": "user", "content": "What does ticket ABC-123 say?" }],
  "tools": [{ "type": "mcp", "connector_id": "<connector id>" }]
}

Accepted shape per entry: exactly { "type": "mcp", "connector_id": "<id>" }, plus an optional server_label (accepted for OpenAI-shape compatibility and ignored — gateway tool names are not label-prefixed). connector_id must be a non-empty string of at most 64 characters. Duplicate ids are de-duplicated; at most 8 distinct connector references per request.

Authorization. The API token's user is the credential identity: a user-scoped connector uses that user's connected credential (connect it in the UI first), and a private connector is visible only to its owner. A tenant-level token (no associated user) can use shared connectors; user-scoped ones return connector_credential_required and private ones connector_not_found. The connector's server-side tool_policy filters the offered tools and is re-enforced on every call; connector tool names that collide with built-in gateway tools or the reserved agent__ prefix are silently dropped from the offer.

Per-tool shape narrowing. Every tool definition that reaches the model — whether discovered from a connector's tools/list or carried in the gateway-internal x-aig-mcp-tools body field — is treated as untrusted and narrowed fail-closed before it is offered: an entry must be a JSON object with a function object and a usable name (non-empty string, valid UTF-8, at most 200 bytes — resolved from the entry's name or function.name), or the entry is dropped; duplicate names keep the first entry only; at most 5000 entries per request are considered. Within each kept tool, non-conforming fields are repaired rather than failing the request — see MCP connectors — tool shape narrowing for the exact field rules.

Rejected inputs (all 400 invalid_request unless noted):

Input Behaviour
unknown / foreign / other member's private connector_id 404 connector_not_found (fail closed, request aborted)
connector requires a credential the token's user has not connected (or it expired) 424 connector_credential_required
connector server unreachable / tool discovery failed / discovery deadline exceeded 502 connector_upstream (fail-closed; but see x-aig-mcp-best-effort below — the chat surface drops the connector and proceeds instead)
mixing {type:"mcp"} references with client function tools in one tools array 400 — one loop owner per request
server_url, allowed_tools, headers, authorization, require_approval, or any other key on an mcp entry 400 naming the key — register a connector; restrict tools via its server-side tool_policy
more than 8 distinct references 400
combining references with the gateway-internal x-aig-mcp-tools body field 400 — one MCP source per request. (The legacy x-mcp-tools header was retired and is now ignored, never an error.)
combining references with legacy functions / function_call 400
a forcing tool_choice ("required", a named function/tool, {"type":"any"}, allowed_tools mode required, …) 400 — the gateway drives the tool loop. Non-forcing forms pass: absent, null, "auto", "none", {"type":"auto"}, {"type":"none"} (shape-agnostic; the provider still enforces its own wire shape)
references on a non-chat endpoint (embeddings, …) 400
references on count_tokens 400 — connector tools are never included in a token count (a refs request cannot pre-count)
references with a model whose provider has no gateway tool loop (e.g. Gemini, Vertex, Bedrock) 400
references on native /v1/messages with stream: true 400 — connector references on /v1/messages require stream:false in this version
references on native /v1/messages with a provider that relays the Anthropic response through an OpenAI-compatible reader (the self-hosted myra fleet, groq, together, …) 400 — that provider dials /v1/messages but cannot read the Anthropic dialect, so no gateway tool loop can run there; without references the request is a transparent passthrough. A provider that dials its own endpoint (a routed cohere model) keeps the loop.
references in a local-only project / egress-blocked context 403 forbidden — "MCP connectors are not available for this request (egress blocked)"

Operational notes.

  • Tool discovery (tools/list) runs server-side with an overall deadline of ~20 seconds across all referenced connectors (plus at most one in-flight operation); results are cached for up to 5 minutes per connector + user, so a connector-side tool change can take up to 5 minutes to appear. A failed discovery (server unreachable / deadline) is briefly negative-cached (~15s) so a persistently-down connector is not re-probed on every request.
  • A connector that (legitimately) exposes zero offerable tools does not fail the request — the turn proceeds with the gateway's built-in tools only.
  • tool_choice: "none" with references still performs the (billable) tool discovery for tools the model can then never call — omit the references instead.
  • MCP-carrying requests are never served from or written to the response cache.
  • A member-created private connector may only reach public addresses; a private-range/internal server_url surfaces as connector_upstream.

Connector provenance notice in the model's context

When at least one connector-backed MCP tool is injected, the gateway merges one system-level notice (sentinel [aig:mcp-connectors]) into the outgoing request that enumerates each attached connector's display name and its tool names, and instructs the model that these connectors are available — call their tools instead of denying access. (Without provenance, some models deny having an attached connector even though its tools are offered.)

Behaviour and guarantees:

  • Merged at most once per request — the sentinel makes re-injection across retries and tool-loop continuation legs a no-op.
  • Absent when no MCP connector tools are attached (built-in-tools-only turns, and agent-as-tool-only turns) — the notice never promises tools the model does not have.
  • Untrusted input is sanitized before embedding (fail-closed): connector display names and tool names are stripped of ASCII control characters and structure punctuation (quotes, brackets, braces, backticks, <>, |, backslash), whitespace-collapsed, and capped at 64 bytes (UTF-8-safe — accent and umlaut characters survive). A display name that sanitizes to nothing falls back to the connector id; if that is unprintable too, the connector is omitted from the enumeration entirely. Raw values are never embedded.
  • Flood-bounded: at most 20 tool names are listed per connector (then "and N more tools") and at most 16 connectors are enumerated (then "and N more connectors").
  • The internal provenance fields (connector_id, connector_name) never reach the provider's tools array.

x-aig-mcp-best-effort request header. Accepted value "1" opts the request into best-effort connector-reference resolution: a connector that fails discovery (connector_upstream / discovery deadline) is dropped and the turn proceeds with the remaining tools, and the drop is surfaced to the SPA via the tools_skipped progress notice instead of failing the turn. The same opt-in also governs the egress-policy drop: in a Tier-1 local_only project the referenced connectors are dropped (token connectors_policy) and the turn proceeds, where a fail-closed caller gets 403 forbidden. A run marked no-egress by its profile (a webhook-triggered agent run) stays fail-closed regardless of the header. Any other value, a repeated header, or its absence keeps the default fail-closed behaviour (a discovery failure aborts the request with the typed connector_* error above). The web chat sends x-aig-mcp-best-effort: 1 on every send — an auto-attached connector whose server is momentarily down must not break the chat; the raw API omits it so an explicitly referenced connector still fails loudly. Under best-effort, any discovery failure for a connector (including connector_not_found / connector_credential_required) drops that connector and proceeds; per-connector authorization is still enforced, so a dropped connector's tools are simply never offered.

The turn_committed frame (turn-owned streams)

A turn-owned stream (one sent with x-aig-turn-id) ends with one aig_status: "turn_committed" frame — the gateway's word that the answer is durably persisted (or that there was nothing to persist). Its fields:

Field Meaning
turn_id the turn this frame closes
message_id the committed assistant row's id, or null when nothing was persisted (an empty turn, or a ghost conversation — then ghost: true is also present)
files[], images[] artifacts the turn produced and claimed onto the row
usage input_tokens, output_tokens for the whole turn
finish_reason the turn's final stop reason
superseded_message_ids[] always present (an empty array when nothing was superseded): the message ids this commit replaced — the whole trailing run of the regenerated question: its prior answer(s) and the error rows among them (see x-aig-regen-of-turn below). The app removes exactly these rows in the same update that shows the new answer.
data: {"aig_status":"turn_committed","turn_id":"…","message_id":"…","files":[],"images":[],"usage":{"input_tokens":12,"output_tokens":48},"finish_reason":"stop","superseded_message_ids":["<prior answer id>"]}

A failed turn (a provider fault, or a guardrail block on the buffered response path) persists its failure row without a turn_committed frame; the app learns the outcome from the error frame and the next conversation load. (A request-phase guardrail block on the streaming path ends with a normal turn_committed — carrying the error/guardrail rows it superseded.)

How large an answer can be committed. Two ceilings, and anything that can be written can be read back. The row's own bound: an assistant message's large text columns — content, any later corrected_content or edit, and the sources / search-steps lists a web-search turn stores — may total at most ~15.9 MB (the largest payload one database packet can carry back — the wire protocol's own single-packet limit — less a reserve for the rest of the row); a write over it is refused before anything is written, atomically with the row it would have grown. The statement's bound: one row is written in a single statement whose packet may not exceed 16 MB. Text is escaped on the way in, so the worst case — an answer made entirely of characters that need escaping — leaves room for about 8 MB; ordinary prose grows by a few percent, so in practice an answer of ~12 MB still fits (the same statement also carries the turn's sources and search steps). A multi-megabyte answer — a long analysis, a big pasted document folded back into the reply — commits normally and reloads byte-identical. Beyond either ceiling the commit FAILS LOUDLY rather than silently dropping the answer: the turn takes the persist-failure path (aig_status: "persist_failed"), and the client sees the answer it streamed but is told it was not saved. Nothing is truncated behind your back, and no half-row is left in the conversation. On the read side the database driver is capped at the wire's single-packet limit, so a row that somehow outgrew it (written before this rule) fails to load with a clear error instead of loading as truncated phantom rows.

x-aig-regen-of-turn — regenerating a tool-backed answer: forced tool, and the prior answer superseded on commit

When the chat UI regenerates an assistant answer, it sends x-aig-regen-of-turn: <original-turn-id> — the logical-turn id of the answer being regenerated.

  • Accepted shape: a string of at most 36 characters matching the request-id pattern (the same rule as x-aig-turn-id). Anything else — a non-string, an over-length value, or forbidden characters — is rejected (dropped, logged at WARN) and treated as absent. The value is used only as a parameterized query parameter, scoped to the caller's tenant (the leg lookup) and to the request's own conversation (the supersede), so a forged id cannot read or touch another conversation's data.
  • Effect 1 — tool forcing: if that original turn ran a model-driven tool loop (its persisted request legs include a tool_loop leg), the gateway forces the model to call a tool on the regenerate (tool_choice: "required") so it cannot decline and answer "no access". If the target provider/model does not support a forced tool choice, the gateway falls back to offering the tool (logged at WARN) rather than failing. When the header is absent, or the original turn used no tool, nothing is forced. (A forked copy of an answer carries a fresh turn id whose legs belong to the source turn — no forcing on it.)
  • Effect 2 — supersede on commit. The app no longer deletes the prior answer before regenerating; it stays visible while the replacement streams. When the replacement commits, the gateway — in the same database transaction — soft-deletes the prior answer and any error rows trailing it, and lists them on the turn_committed frame as superseded_message_ids. Rules:
    • Who: the conversation's owner only — the same predicate the message DELETE route's storage applies (a project member deletes nothing there, and so supersedes nothing here — not even as the current driver of a live session; their regenerate is appended beside the owner's answer). Decided once when the turn starts, never on the commit path. A non-owner, an API-key turn (no user), an invisible conversation, or a gateway-side fault ⇒ nothing is superseded, the new answer is simply appended (both remain visible), and the refusal is logged at WARN + recorded as the trace step regen_supersede_refused with a reason (not_owner, no_user, no_turn_id, not_found, db_error:<err>; and from inside the commit: no_anchor, not_trailing, self_anchor — the header named the turn itself). Note that the turn itself is not refused on ownership grounds: the streaming path checks the live-session lease, not conversation ownership (a pre-existing gap, tracked separately) — so a member's regenerate on a shared conversation still lands its answer in that conversation; only the supersede is withheld.
    • What: the answer with that turn id is the anchor (resolved even if already soft-deleted, so a stale tab regenerating twice still ends with ONE live answer — last-commit-wins); the whole trailing run is superseded — every assistant row after the question the anchor answers (the last user message at or before the anchor; with no such message, the anchor itself and everything after it) — but only if no live user message follows the anchor (a stale tab regenerating an older turn never deletes another question's answer: refused, logged at WARN, traced as not_trailing). So an earlier answer stacked under the same question (an appended non-owner regenerate, a rolling-deploy gateway) goes with the swap too. Live streaming claim rows of a sibling turn and the new row itself are never touched; the rows soft-deleted are exactly the live subset of the ids on the frame (an already-superseded row is listed so a stale tab removes it too, but is never deleted twice).
    • A failed regenerate never removes a real answer. A provider failure or a guardrail block persists a failure row that supersedes only the error/guardrail rows of that run — so the prior answer stays, the failure lands beneath it (and repeated failures never stack), and a later successful regenerate replaces both.
    • A stopped regenerate (the client aborts after tokens streamed) commits its partial and supersedes the prior answer like a full answer would.
    • Idempotent and atomic: a retried commit of the same turn re-runs the rules with no effect on already-deleted rows; the supersede and the new row commit together or not at all.
    • Rolling deploys: a gateway container still on the previous image supersedes nothing (both rows stay visible until the fleet has rolled); its turn-owned answers carry a turn id like any other, but rows it writes through the legacy paths (POST /messages, a bridge pair, a fork copy) have none and cannot be superseded later. Every assistant row written by a current gateway carries a turn id on every path (migration 0283 backfilled the historical ones).

Download-request tool forcing (server-side, no header)

When a chat turn's latest user message asks for a downloadable file — e.g. "gib mir die Datei zum Download", "erstelle mir eine Excel-Datei zum Download", "give me the file to download", "export this as a PDF" — the gateway forces a file tool on that turn so the model produces a real file instead of narrating a phantom "here is the file to download: X.xlsx" with no tool call (the prod class: no file, no download card). This needs no header; it is inferred from the user's message.

  • When it fires: the latest real user message is classified as a download request (a deliver-verb — gib mir / give me / erstelle / export / … — plus a file/format token, OR an explicit "zum download" / "as a download" phrase) and at least one file tool (write_file or code_interpreter) is resolved for the turn and the provider honors a forced tool_choice and the request carries no client/gateway tool_choice already.
  • What it does: for that first leg only, the offered tool set is narrowed to the resolved file tools {write_file, code_interpreter} and tool_choice is set to "required", so the model must call one of them (it still chooses write_file for text/HTML/office vs code_interpreter for computed/formatted output). Continuation legs are unforced (they answer from the tool result). A provider that cannot guided-decode "required" is not forced (logged at WARN); a rare 400 on the forced leg strips the force and retries once.
  • Precision (fail toward NOT forcing): requests about a file — "gib mir eine Zusammenfassung / Analyse / einen Überblick der Datei", "give me a summary of the file" — and how-to / question / inability phrasings do not force (the user wants an inline answer, not a file). A genuine download ask this misses is caught by the honesty backstop below.
  • Backstop: if a turn still ends with the model claiming a download it did not produce, the claim-steer (detect_file_claim) re-legs it into a real write. The download card itself always renders from the durable chat_message_file relation, never from model prose.

Bare re-download → server re-serve of the existing file

A bare re-download — the user asks for the file they already have back with no change: "lade die Datei herunter", "download it", "gib mir die Datei nochmal" — is handled by re-serving the file that already exists, not by forcing a fresh creation. Forcing a new creation here is both wasteful and unsafe on a weaker model: at 1000-case scale gemma4's forced regeneration produced accepted=0 and then simply stopped, leaving the user with no download card even though a perfectly good file was sitting in the conversation.

  • When it fires: the turn is classified as a download request (as above) and carries no modifier — no format/conversion token (als PDF / as pdf / convert), no version verb (aktualisier / updated / upgedatet / Ändere / Überarbeite), no create verb (erstell / mach mir / create), no daraus / draus, and no redirect (email / slack / @). Any modifier → this does not fire and the normal download-force creation path runs, so "als PDF" still produces a real new PDF. (The modifier list dual-cases sentence-initial umlauts — Ä/Ü do not ASCII-lower to ä/ü — so "Ändere die Datei und lade sie herunter" is correctly treated as a modify, not a bare re-download.)
  • What it does (two coordinated steps, model-INDEPENDENT): at turn setup the gateway looks up the conversation's most recent generated file (chat_project_knowledge WHERE conversation_id = ? AND chat_generated = 1 ORDER BY (filename = ?) DESC, created_at DESC — a file the user named wins over the newest) and, if one exists, skips forcing the model to regenerate and caches the row. Then at the streaming terminal it re-emits that file's download card deterministically — whatever the (now-unforced) model says or doesn't say, including an empty turn or an honest decline. It is the same card the original write_file/code_interpreter turn produced, reconciled into chat_message_file so it survives a reload. No new model creation leg runs, so a weak model can neither botch nor corrupt it, and no duplicate file is minted.
  • Neutral re-download card, not "Updated": the re-served link reuses the original file's file_id in a new message, so without a marker the client would read it as a same-name rewrite and label the card "Updated" (and mark the original "superseded") — implying a change that never happened. The re-serve therefore persists chat_message_file.reserved = 1 (migration 0257; from ctx.reserved_file_ids), and the conversation-message files[] relation carries a reserved: true boolean on that entry. A reserved link renders as a plain re-download: the client ignores it when inferring Created/Updated, so the original card stays Created and the re-download card carries no revision label. Every genuine write_file link is reserved: false.
  • Trust boundary (owner-gated, fail-closed): the lookup is keyed on the authenticated user_id via get_conversation_route_scalars (the owner gate, WHERE user_id) — never on the untrusted x-aig-conv-id header or any history marker alone. A nil/empty principal, a foreign/shared-reader caller, or a conversation with no owned generated file each yield no card (no metadata leak, no cross-tenant IDOR). A re-served id is tagged in ctx.reserved_file_ids so the artifact-drop guard does not misfire on it, and set in ctx.written_file_ids so the branch-e file-claim and promise steers do not double-fire.
  • Precision: if no owned generated file exists (e.g. the first turn is "download the file"), nothing is re-served and the normal download-force creation path runs — a bare re-download only short-circuits when there is genuinely a validated file to hand back.

x-aig-datafile-auto — permit a per-turn model upgrade for data-file turns

The web chat sends x-aig-datafile-auto: 1 when a conversation routed by Auto (no explicit model pick) sends a turn that carries an inline data-file block (an attached spreadsheet/CSV preview). The header is a permission hint, never an authorization: it only allows the gateway to consider upgrading this single turn to a more capable model; every decision input is validated server-side.

  • Accepted shape: exactly the string "1". Any other value — absent, empty, another spelling, an over-long value, or a repeated header — is treated as absent (no upgrade consideration; the request proceeds unchanged).
  • Effect: when the request also actually carries an inline data-file block (verified server-side — a hint without a file is inert), the resolved model is not reliably tool-capable per the gateway's curated capability registry (provider/catalog self-claims do not count) with a context window covering the request, staging the raw file is possible on this gateway, and the tenant is not in a reduced-model (over-cap) window, the gateway rewrites this turn only to the best eligible model: providers the gateway can actually dispatch to (configured or managed fleet, intersected with data-residency and provider-allowlist enforcement) ∩ the plan's model allowlist ∩ usable chat models with registry-proven function-tool support and an adequate context window — ranked sovereign-first, then by price. No eligible model → the request proceeds unchanged. The upgrade is never persisted to the conversation; a later turn without a data file routes by the conversation's own model again.
  • Disclosure: an upgraded turn reports the actually served model in the X-AIG-Model response header on streamed responses (X-AIG-LLM-Model on non-streamed ones), and the assistant message row is persisted with the served model — the web chat renders a per-turn "Answered by …" notice from it.
  • Security: because the target set is confined to models the caller could already select explicitly on the same gateway, a spoofed hint cannot cross a residency, plan, or authorization boundary — at most it moves the caller's own turn to another model the caller is already entitled to use.

x-aig-auto-complexity — permit a per-turn Auto complexity upgrade

The web chat sends x-aig-auto-complexity: 1 on every turn of a conversation routed by Auto (no explicit model pick), to mark that the turn used the conversation's sticky Auto model rather than an explicit choice the gateway must never override. Like x-aig-datafile-auto, it is a permission hint, never an authorization: it only allows the gateway to consider upgrading this single turn; every decision input is validated server-side.

  • Accepted shape: exactly the string "1". Any other value — absent, empty, another spelling, an over-long value, or a repeated header — is treated as absent (no upgrade consideration; the request proceeds unchanged).
  • Effect: only on a gateway with per-gateway complexity routing enabled (gateway_config.laya_routing = shadow or live), the hint lets the gateway run the self-hosted Laya classifier on the turn's user text. In live mode a turn whose hard-probability mass clears the floor is rewritten this turn only to the best eligible frontier model — intersected with data-residency and provider-allowlist enforcement ∩ the plan's model allowlist ∩ registry-proven tool-capable chat models with an adequate context window, ranked quality-first. No eligible model, a below-floor signal, or an easy/medium tier → the request proceeds unchanged. In shadow mode the verdict is only logged; nothing is rewritten. The upgrade is never persisted to the conversation; a later turn routes by the conversation's own model again.
  • Disclosure: an upgraded turn reports the actually served model in the X-AIG-Model response header (X-AIG-LLM-Model on non-streamed responses), and the assistant message row is persisted with the served model.
  • Security: the target set is confined to models the caller could already select explicitly on the same gateway, so a spoofed hint cannot cross a residency, plan, or authorization boundary — at most it moves the caller's own turn to another model the caller is already entitled to use. A best-effort classifier failure (Laya unreachable) fails open: routing is unchanged.

x-aig-web-search-freshness — restrict web search to a recency window

When web search is enabled for the request (via x-aig-web-search: 1, or a gateway configured with mode: "always"), this optional header narrows the underlying search to a time window so the model grounds its answer in recent results.

  • Accepted shape: one of the discrete windows pd (past day), pw (past week), pm (past month), py (past year), or an explicit inclusive date range in the form YYYY-MM-DDtoYYYY-MM-DD (e.g. 2022-04-01to2022-07-30). Anything else — an unknown token, a partial range, wrong separators, extra query fragments, or a non-string — is rejected: the header is ignored (fail closed) and the search runs unfiltered rather than forwarding an unvalidated value to the search provider.
  • Effect: the validated window is passed to the web-search provider so only results within that period are returned. When the header is absent or rejected, web search behaves exactly as before (no recency filter). The filter applies to every cross-provider path that runs the gateway's own web-search backbone (buffered OpenAI-format providers and the streaming vllm/myra tool loop); providers that use their own native search (Anthropic, OpenRouter, Gemini, Mistral Agents) ignore it. Recency is a Brave-adapter feature: the EU linkup adapter (see below) currently ignores the window and searches unfiltered.
  • In the app: the chat composer's + menu exposes a Recency dropdown (Any time / Past 24 hours / Past week / Past month / Past year) whenever web search is on; picking a window sets this header for the next send. "Any time" sends no header (unfiltered).

tool_choice on a web-search turn is honored one-shot

When x-aig-web-search: 1 injects the gateway's web_search tool and the request carries a forcing tool_choice ("required", or a named function/tool), the gateway forces the model to call the tool on the first leg only, then frees it: the injected search runs, and the model is unforced on the continuation leg so it can produce a final answer. Without the one-shot drop, a forcing choice would be re-sent on every leg and the tool loop could never terminate. Non-forcing values (absent, null, "auto", "none") are untouched. (This differs deliberately from an MCP connector-references loop, which rejects a forcing tool_choice with 400 — there the tools are resolved server-side and invisible to the client, so forcing them is meaningless; a header-injected web_search, by contrast, is a tool the client explicitly opted into. Both paths share one fail-closed forcing definition.) A provider/model that does not support a forced tool choice is unaffected by web search; a forcing choice sent to such a provider follows that provider's own wire behaviour.

Web-search direct answer relays tool_calls and labels the terminal reason

When x-aig-web-search is on but the model answers directly on leg 1 (it does not call web_search), the gateway returns that leg as the response. On the compat (OpenAI) egress:

  • if the model made its own tool call on that leg (a client-provided tool), those tool_calls are relayed on choices[0].message.tool_calls (with finish_reason tool_calls) — the turn is not collapsed to an empty completion; and
  • otherwise the terminal finish_reason is routed through the gateway's single stop-reason authority: a length-capped answer surfaces as length (so a truncated web-search answer is labelled incomplete, matching the non-web-search paths) rather than a clean stop; every other terminal reads as stop.

(An empty tool_calls array or a non-array value is not relayed. On a PII gateway the relayed tool-call arguments are token-restored with the correct depth-2 escaping.)

If the web-search direct answer instead comes back with zero visible content (no text and no tool call — e.g. a reasoning-only turn), it is classified through the same shared empty-response path as the non-web-search buffered egress: the turn is counted as an EmptyResponse in the gateway's telemetry, and the compat egress delivers a short fallback bubble (a "produced internal reasoning but no visible answer" or "returned an empty response" note) instead of a blank message — so a web-search direct answer is never a silently empty turn, matching the other egress paths.

Web-search provider selection and EU residency

The gateway's own web-search backbone is provider-pluggable, selected per gateway by web_search.provider in the gateway configuration:

  • brave (default) — Brave Search (US). Used for every gateway that does not set the field.
  • linkup — linkup.so (EU, France; GDPR Art. 28 DPA, zero data retention). The EU-sovereign backend for tenants that must keep the model-generated search query in the EU.

The per-gateway search API key stays in the existing web_search.api_key field (it holds the key for whichever provider is selected). Web search is turned off per gateway with web_search.enabled: false.

Platform Linkup key injection at read

Every new production gateway — the pay-first default gateway, both trial gateways (internal and external), and a gateway created by an administrator via POST /admin/v1/tenants/{id}/gateways — is provisioned with web search on via Linkup — web_search: { enabled: true, provider: "linkup", max_results: 5 } — so a fresh organisation can search immediately with no manual setup (AGF-2874; before it, only trial gateways were seeded). The block is seeded only when the gateway has no web_search block at all: an explicit web_search.enabled: false, a block with its own api_key (BYOK), another provider, or any existing value is a deliberate choice and is never overridden. Gateways created with purpose test, benchmark or archived (E2E fixtures, rigs) are not seeded. Existing trial gateways were backfilled the same way (migration 0302); existing non-trial gateways that have no web_search block are not backfilled (a separate decision).

The Linkup api_key for these gateways is not stored in the gateway config. It is a shared Myra platform key held in the database settings row trial_linkup_api_key (the historical name — it applies to every plan) — a super-admin sets it live from the admin console (Feature Flags), taking effect within about a minute. It is injected into the effective config at read time, for any tenant plan, gated only on the block being enabled, the provider resolving to linkup, and the block carrying no api_key of its own. It is a low-value search key, not a credential: a plain setting — the settings GET returns it verbatim and it is admin-editable live (see Global settings API). Consequences:

  • The stored gateway config (and therefore the admin GET/list responses and the tenant config export) contains the web_search block but never the platform api_key — the shared key is injected only at read time and is not written into per-gateway config or exports. This holds for every plan, including callers holding GATEWAYS_MANAGE.
  • If trial_linkup_api_key is unset (or empty), new gateways are still seeded with the keyless block (so a later key set lights them up with no backfill), but they have no web search until then: they report web_search_setup_pending: true, the /easy composer shows a visible "not set up" notice, the Health dashboard flag web_search_platform_key_missing is true, the gateway logs ERR every 10 minutes and a [platform_alert] fires (at most every 6 h). Signup and gateway creation never fail because of it. Set it to a dedicated, rotatable Linkup key from the admin console.
  • A trial → paid conversion changes nothing for web search: the injection is plan-agnostic, so the converted organisation keeps searching. A customer who wants their own Linkup or Brave account sets web_search.api_key (and provider) on the gateway — an explicit key always takes precedence over the platform key.
  • max_results is honored by the Brave adapter but ignored by Linkup (kept for shape parity).

Regardless of plan, the server-computed web_search_configured boolean the gateways endpoint returns (and the composer web-search globe reads) reflects the effective config — so a keyless Linkup gateway reports web search as configured via the injected key, without the key ever reaching the client.

  • Fail-closed EU gate. When a gateway (or its tenant) enforces EU residency (eu_region_routing, including the inherited tenant_eu_region_routing floor), the model-generated search query may egress only to an EU-vouched search provider. A non-EU provider — including the brave default a sovereign gateway forgot to switch — is blocked: the query is never sent, and the model is told the search was withheld. A sovereign tenant can therefore turn web search off, but can never silently fall back to Brave (US). This also neutralizes the buffered path's wttr.in weather side-channel (which would otherwise embed the raw query in a URL to a US host). The query PII scrub (see guardrails.md) runs on top of this on every provider path.
  • The linkup response is untrusted third-party content and is validated fail-closed (principle 11): a non-2xx status, a non-JSON / non-object body, or a body with no results array yields a classified failure and zero results (never a partial or fabricated answer); every result url must be an absolute http(s) URL or it is dropped (a javascript:/data:/scheme-relative/empty URL can never become a citation or link). linkup's results[].{name,url,content} are normalized to the internal {title,url,snippet} shape at the single dispatch point, so downstream formatting/citation code is provider-agnostic.

Inline numbered citations

When the gateway's own web-search backbone runs (the cross-provider vllm/myra path — not a provider's native search), each search returns a numbered source list to the model, and the model is instructed to mark the claims it grounds with inline [1], [2] … references. The answer ends with a numbered Quellen: (sources) footer listing those URLs in the same order, so [N] in the prose maps to source N in the footer. In the chat UI this footer is rendered as a collapsed-by-default disclosure ("Quellen (N)") so a long source list does not flood the answer; the inline [N] markers remain direct links to the source, so a source is still one click away without expanding the list.

Citations inside a written file. The [N] markers belong to the chat answer only — a file produced with write_file has no footer to resolve them, so a document that carried them showed bare [1], [3]. The model is therefore instructed (in the write_file tool description, the file-tool rules of the composed system prompt and the file-first system notice — one shared rule, emitted only on turns where write_file is actually offered) not to use inline [N] markers inside a file and to put the sources in a final Sources / Quellen section (one line per source with title and URL), or to omit them when the user asks for a clean document. The gateway does not rewrite file content: what the model writes is what is stored and delivered.

  • Source URLs are untrusted third-party content. They are validated at two boundaries and fail closed: non-http(s) result URLs (e.g. javascript:/data:) are dropped when the provider (Brave or linkup) response is parsed and again when the footer is rendered (so Anthropic-native citations, which don't pass through the search path, are covered). Source titles are stripped of markdown-link metacharacters. The SPA re-validates every URL scheme before rendering a link.
  • Best-effort per claim. Inline [N] anchoring depends on the model following the instruction (as with Perplexity/ChatGPT Search). The numbered source footer is the guaranteed floor; models that don't emit [N] still get the footer, and an out-of-range or hostile-scheme [N] is left as plain text. Providers with native search cite in their own style and are unaffected.

Auto-disable of injected tools on a non-tool-calling model

Some routed models (for example the Perplexity Sonar family) expose no upstream endpoint that accepts a function-tools array; injecting the gateway's own function tools (web search, URL fetch, file access, knowledge, code, connectors) would make the provider reject the whole request (OpenRouter returns 404 "No endpoints found that support tool use").

For a model the gateway knows cannot do function calling (recorded in its capability registry), the gateway now proactively strips its injected function tools before the upstream call and lets the run proceed as a plain answer — instead of relaying the opaque 404.

The notice has a second, reactive trigger. The proactive strip only fires for a model the gateway has explicitly curated as non-tool, which cannot keep up with a catalog that grows (an OpenRouter image model, openai/o1-pro, whatever is added next). So when a model the gateway did not recognise as non-tool refuses the injected tools at the upstream — OpenRouter answers 404 "No endpoints found that support tool use" — the gateway now strips its own injected set, rebuilds the system prompt so it no longer offers tools the model cannot call, and re-sends the turn once as a plain answer. Same notice, same category tokens; the only difference is that the strip was learned from the refusal rather than known in advance. At most one extra upstream call per turn, and only for tools the gateway injected. A caller that named its own tools gets its provider's error passed through verbatim — both a plain tools array (the client is driving its own agentic loop, and silently removing its tools would leave it waiting for tool_calls that never come) and an array of {"type":"mcp"} connector references (those were explicitly requested, and the typed error names the actual lever: remove the references or pick a tool-capable model).

Scope, stated plainly: the retry is triggered by the upstream saying it cannot accept tools, so it covers the refusal phrasings the gateway recognises — OpenRouter's "No endpoints found that support tool use", and the "does not support tools / tool use is not supported / function calling is not supported" family. A backend that refuses tools in wording the gateway does not recognise still surfaces the typed error. That is a shorter and slower-moving list than one entry per model (openai/o1-pro is not an image model, which is why per-model curation could not keep up), but it is a list.

The client is told which tool categories were skipped:

  • Streaming (chat / playground): an additive SSE status event {"aig_status": "tools_skipped", "tools": ["web_search", "url_fetch", ...]} just before [DONE]. The SPA renders an informational notice on the assistant message. tools is a deduped list of coarse category tokens (web_search, url_fetch, files, knowledge, image, code, connectors, connectors_policy, or a generic tools). connectors means a connector was unavailable (transient — a retry may help); connectors_policy means the project forbids external egress, so the connectors could never have run (permanent for that conversation, and nothing is broken). When a dropped connector can be identified, the same event also carries an optional connectors array — [{"id","name","reason"}] — naming each dropped connector so the SPA can name it (instead of the bare connectors word) and, for a credential-class reason (credential_required / credential_expired), offer a Connect/Reconnect deep link to /connectors?focus=<id>. The coarse tools categories still cover nameless drops (policy / discovery-deadline). The gateway also injects a turn-scoped system note telling the model the named connector exists but is unavailable this turn (so it asks the user to reconnect rather than claiming the connector does not exist).
  • Buffered agent /invoke: a tools_skipped trace step (visible via the run's X-AIG-Trace-Id / run-detail) records the skipped categories, the model, and the provider.

Provider-native web search (OpenRouter plugins, Gemini grounding, Anthropic's native search) needs no function-calling capability and is injected in the provider's request build, so it intentionally survives this strip and is not listed as skipped — and that is true of the reactive strip too: the notice enumerates only what the gateway actually removed, so a turn whose OpenRouter web plugin still runs is never told it has no web search. A tool-incapable model the gateway does not recognize up front is not stripped here (this section describes the proactive strip, which fires only for a curated non-tool model) — it is stripped reactively instead, once the upstream refuses the injected tools, as described above. model_capability_mismatch (see error codes) remains for the case where that stripped retry ALSO fails, or where the refusal names a capability stripping cannot fix (an image, a document).

EU-residency exception. Auto-disable + proceed applies only on a non-EU gateway. On an EU-residency-enforced gateway (eu_region_routing), a definitively-non-tool online model (e.g. a Perplexity Sonar model) is not auto-disabled — it stays fail-closed. Such a model's intrinsic web search egresses outside the EU and cannot be stripped by the gateway, so proceeding would breach the tenant's data residency; rejecting is the safe outcome. In practice today's such models (Perplexity Sonar via OpenRouter, a non-EU provider) are refused by the residency dispatch gate with data_residency_blocked before any upstream call; skipping the auto-disable here simply avoids emitting a misleading tools_skipped notice ahead of that block. (A future definitively-non-tool model on an EU-vouched provider whose intrinsic search still egressed would instead be rejected with model_capability_mismatch.)

Web search enabled but not used / no results / failed

When web search is enabled for a turn the gateway keeps tool_choice: "auto" — it never forces the model to search (that would burn a turn on greetings, rewrites, and context-answerable follow-ups). A consequence: a weaker model may answer from its own training, or ask "should I search?", without running a live search — and a search that the model does run can still fail or return nothing. In all three cases the answer used to look normal with no signal that no trustworthy live search happened.

The gateway now emits an additive side-channel status event so the SPA can render a distinct chip on the assistant bubble:

data: {"aig_status":"web_search_notice","reason":"not_used"}
reason Meaning Chip
not_used Web search was enabled but the model answered without calling it (replied from training / asked whether to search). "Answered without web search"
no_results The search ran but returned zero hits. "Web search returned no results"
failed The search provider errored (quota / auth / network). "Web search failed"

reason is the only field; a value outside the three above is ignored by the SPA (the boundary rejects unknown values rather than rendering a raw slug). The event is additive — an older client ignores the unknown aig_status (no schema bump) — and is delivered only to a client that opted into the aig_* side channel. It is emitted on the two-leg web-search engine (not_used at the buffered leg-1 direct answer, re-emitted on the buffered→SSE tail; no_results/failed inline on the streaming leg-2 wire) and on the tool-loop web-search path. It is a notice, not an error — the turn proceeds and the answer is still delivered; the chip only tells the user the answer is not backed by a fresh search. Live-session only — like the tools_skipped and pii_masked notices, it is not persisted, so a reloaded conversation carries no chip.

Declining all tools with tool_choice: "none"

A request that carries tool_choice: "none" (the string, or {"type": "none"}) is a declaration that the model will call no tool this turn. On such a request the gateway injects no built-in tools — not web search, URL fetch, file access, knowledge, image generation, nor code interpreter — and strips the now tools-less tool_choice from the upstream body (a tool_choice with no tools array is rejected by several providers). This is honored uniformly across every provider, including the pass-through providers that have no gateway tool loop (Gemini, Vertex, Bedrock).

  • Only an explicit "none" opts out. Absent, null, and "auto" are not a decline — they receive the normal offered tool set. A forcing value ("required", a named tool, {"type":"any"}, …) is not a decline either (and is separately rejected on the MCP path — see the connector table above).
  • MCP connector references are the one exception: a request that ALSO carries tools:[{"type":"mcp", …}] still resolves those connectors and runs the loop with tool_choice:"none" (the model declines to call), exactly as the connector table above documents — a client that named connectors asked for that loop. The built-in-tool suppression above applies to the plain (no-refs) request.
  • The client is never the authorization boundary: tool_choice:"none" can only reduce the offered tools, never widen them.
  • The decline is snapshotted from the client's request before any gateway middleware mutates the body, so a gateway-internal tool_choice:"none" set on a later leg (e.g. the web-search two-leg synthesis leg) is never mistaken for a client decline.

This is how the /easy surface's post-turn utility completions (auto-title, follow-up suggestions) stay pure text: they send tool_choice:"none" so a code-interpreter-enabled gateway does not run the sandbox for a completion the user never sees.

Inline tool-call recovery (self-hosted vllm/myra models)

A self-hosted model (our qwen/vLLM fleet) occasionally emits a tool call — most commonly web_search — as inline text in the assistant content instead of a structured tool_calls field, when the serving-side tool parser fails to extract it. The gateway recovers these so the tool still runs. Recognised dialects:

  • the wrapped XML form <tool_call><function=web_search><parameter=query>…</parameter></function></tool_call>;
  • a whole-content JSON object {"name":"web_search","arguments":{"query":"…"}} (also accepted inside a <tool_call>…</tool_call> block, and when arguments is a JSON-encoded string);
  • a whole-content pseudo-call web_search(query="…") / {web_search(query="…")}.

The recovered call is untrusted model output and is validated fail-closed at this boundary — a call is recovered only when ALL hold, otherwise the text is left as-is (no tool runs):

  • Accepted: the JSON/pseudo call is the whole trimmed message OR the sole content of a <tool_call>…</tool_call> block (the wrapper is itself an anchor, so a wrapped call is honoured even with surrounding prose — consistent with the existing <function=…> behaviour); its name is a string on the request's resolved/allow-listed tool set; and its arguments is a JSON object with at least one string key.
  • Rejected (rendered as plain text, no dispatch): an un-wrapped JSON/pseudo tool call embedded in surrounding prose (a documentation example, a pasted snippet, "what does this JSON look like"); an unknown / non-allow-listed / hallucinated tool name; missing / empty / non-object arguments (incl. a JSON array); malformed or truncated JSON. write_file/read_file are never synthesized from this ambiguous channel (a mutating write is only dispatched from an unambiguous marker). When a write_file/read_file IS emitted as a bare-JSON/pseudo call and therefore not dispatched, the gateway does not lose it silently — it records a fail-loud silent_write_file_drop / silent_tool_call_drop trace step and an ERR log (so a new serving dialect is caught and patched), rather than letting the reply falsely claim success.
  • Recognised tool tag the visible stream removed but nothing synthesized: when the model's entire visible output is an inline tool tag the gateway strips from the bubble (a <tool_call> / <invoke> / <function_calls> wrapper, or a <write_file> with no parseable filename) but no dialect/allow-listed call is recovered, the assistant turn used to render as a silent empty bubble — the user had to send a bare ? for the model to produce the answer on the next turn. It now surfaces a short retry hint ("I tried to call a tool but the call wasn't in a format the gateway could parse. Please ask me to continue…") instead, and records an EmptySuppressedNoTool (gateway_fault) signal so the unrecognised dialect is caught and patched. (A <memory> tag stripped from an otherwise-empty turn is not a tool call and gets neither this hint nor the empty-response fallback — see Empty responses above; a whitespace-only turn does get the fallback, tools offered or not.)
  • Unwrapped <write_file> / <read_file> block bounding: in the gemma/qwen inline-XML channel the next tool opener is a hard wall for every block: a block ends at its own close tag when that lies before the wall, otherwise at the earliest of a Hermes terminator (</function> / </tool_call>), the wall, or the end of the message. A close tag beyond the wall belongs to that later block and never ends the earlier one, so two <write_file> blocks where the first was left unclosed yield TWO calls — the first marked torn (the dead-leg diagnosis reads it as such), the second whole — instead of one call for the first file with the second file's markup embedded and the second file silently never written. An open tag whose > lies past the next opener is unparseable and skipped — no call is yielded for it (the next opener is parsed on its own). Bounding is dialect-free (the block's own close tag before the wall ends any block; otherwise the earliest of a Hermes terminator, the wall, or the end of the message); whether a block counts as closed follows its dialect: the attribute form (<write_file filename="…">) closes only with its own </write_file> — a </function> / </tool_call> inside it is quoted content, which bounds the block (the two are indistinguishable on the wire) but leaves it torn; the parameter forms (<write_file> / <function=write_file> with <parameter=…> children) close with the Hermes terminators or their own </write_file>. A body in the attribute form that quotes a Hermes terminator keeps it when its own close tag arrives before the wall. The scan is linear in the message on every axis (blocks, body bytes, repeated <parameter=…> openers): every literal lookup is a plain find, memoised. The wall's inherent trade-off: a body that quotes another tool opener (a write documenting a tool call) is split at that opener — a nested quote and an interleaved torn block are indistinguishable without a depth model — so the quoting block comes out torn and the quoted call is dispatched.
  • web_search query typing: on the buffered path a non-string query is not recovered (left as plain text); on the streaming path the call is dispatched and the tool returns an "Error: query required (must be a string)" result (it never crashes).

Degraded write_file sentinel recovery + file-claim honesty guard (streaming)

A weaker model sometimes announces a file instead of calling write_file: it prints a pseudo-sentinel like DATEI_ERSTELLT:report.html, a bare write_file token, or narrates "the file is saved — click the file card" while the turn produced zero write_file artifacts. Previously nothing recovered, stripped, or corrected these — the sentinel leaked into the visible answer, no download card existed, and the user was stranded retrying. On streaming tool-loop legs the gateway now:

  • Sentinel recovery (whole-content anchored, like the JSON/pseudo dialects above). Accepted: the entire trimmed message is a single IDENT:<filename> line (IDENT all-caps, ≥4 chars; the filename a single token with a 1–8-char alphanumeric extension; at most one space after the colon — e.g. DATEI_ERSTELLT:Report_Q2.html, FILE_CREATED: notes.md), a bare write_file / write_file() token, or a contentless pseudo-call write_file(filename="…"). The blob is suppressed from the visible answer and ONE contentless write_file call is dispatched instead — the tool rejects it deterministically (content required …, before any persistence or cap tick, so the dispatch is provably side-effect-free) and the error text steers the model into a real write_file call on the next leg. Rejected (left as plain text / loud-guarded only): a sentinel embedded in prose or spanning multiple lines; lowercase or short identifiers; a post-colon token without an extension (WICHTIG: Bitte …); a pseudo-call with a content key (a mutating write is never synthesized from this ambiguous channel — it stays with the fail-loud guards); any sentinel when write_file is not among the resolved tools (fail-closed); a sentinel co-occurring with a structured tool call (silent_write_file_drop fires with file provenance instead of dispatching).
  • File-claim honesty guard (once per turn). When the leg's visible text claims a produced file while the turn delivered no downloadable artifact and nothing is dispatched, the gateway dispatches the same contentless steering write_file (hard one-shot per turn; a claim that merely references an earlier turn's file costs one steering leg and the error text tells the model to answer honestly). A claim is one of: gateway card vocabulary (Datei-Karte/file card) plus a named document file; a sentinel-shaped line mid-prose; a gateway delivery-card term (downloadable card/download card/herunterladbare Karte) on a line that is not a modal offer, a how-to/imperative, or 2nd-person attribution — scanned per line (a modal on another line cannot excuse a phantom line) and, uniquely, without a filename, since the prod repro (opus-5: "the CSV is attached above as a downloadable card") carried none; a verbatim echo of the gateway's own history CREATED_FILE frame ([In this reply you created these downloadable file(s)…: X]), a frame that is gateway scaffolding and never legitimate model prose (so, like the sandbox surface below, it needs no exclusion filter); a download-offer term (zum Download/to download/for download/as a download) on a non-excluded line plus a document filename anywhere (the offer and filename may be on different lines); or a hallucinated markdown download link with the sandbox: scheme ([name.xlsx](sandbox:/mnt/data/name.xlsx)) — the OpenAI/ChatGPT training-data artifact path that resolves to nothing here (no producer or resolver exists in the backend or SPA), which weak models (e.g. gemma-4-26b-a4b-it) emit instead of calling write_file, so the client renders a dead, non-clickable link and no download card. Untrusted-input shape (principle 11): the sandbox surface is anchored on the presented-link opener ](sandbox: — a rendered clickable link — so it is rejected/steered; a bare sandbox: prose mention or quoted path (no ]( opener) is accepted (not steered), as is a link to any non-sandbox: scheme. Because ](sandbox: has no legitimate producer it is not run through the modal/how-to/2nd-person exclusion filter (a dead link is a phantom even when framed 2nd-person, e.g. "Du kannst x öffnen"); the accepted residual is a user pasting ChatGPT output that the model then reproduces (pasted or illustrative) on a no-artifact turn → one bounded steer leg, since the honesty clause only forbids claiming a saved file; a folder claim — export folder / Export-Ordner / Export Ordner / Exportordner, per line, without a filename, since the prod repro ("Fertig — bitte prüfen Sie den Export-Ordner") carried none: excluded by a how-to, a filesystem path or menu crumb ("unter Einstellungen > Berichte") anywhere on the line, or by a first-person offer ("I can put it in the export folder"), a negation ("there is no export folder here") or a question mark in the sentence holding the term (so a courteous closer — "…Export-Ordner. Brauchen Sie noch etwas?" — never excuses the claim) — not by 2nd-person framing, because no export folder exists in this product (the same structural rule as the sandbox: link: the natural prod phrasings are 2nd-person); or a delivery claim without a card term — saved / gespeichert in the same sentence as an office-format token (PDF, DOCX, XLSX, PPTX, ODT, ODP, ODS — with or without a filename, so "in eine PDF-Datei umgewandelt und gespeichert" counts; the filename, when named, is the last office filename of the sentence, the delivery target — a heuristic: a source named after the target still wins, harmlessly, since the steering call is contentless), excluded by how-to words, a filesystem path or a menu crumb anywhere on the line, or on one of the three preceding non-blank lines when that line was instructional in shape — a numbered step, a bullet, a table row, a heading, a "…:" lead-in, or an explicit how to / Anleitung phrase ("Öffnen Sie Word, klicken Sie auf Datei > Exportieren. Die Datei ist dann als PDF gespeichert." is a how-to answer, and so is the same answer rendered as steps, bullets or a table, where the instruction words sit above the concluding step; a merely courteous line — "Use the button below." — does not silence what follows it, and a fenced code block consumes the carry like any other line; the same rule applies to the folder branch), and — in the claim sentence — by a present-passive description ("can be saved", "wird … gespeichert", "…, dass X gespeichert wird"), 2nd-person attribution of the save ("the report.pdf you saved", "Sie haben … gespeichert", and with words in between: "Die Datei, die Sie hochgeladen haben, wurde als angebot.pdf gespeichert"), a negation (an honest decline never steers — "wasn't saved", "nichts gespeichert"), a date/year (provenance: "wurde 2023 im SharePoint gespeichert"), or a final question mark ("Wurde report.pdf gespeichert?"), or a wie … herunterladen instruction idiom; or a readiness / delivery-availability idiom — a German <verb> bereit (liegt/liegen/steht/stehen, then whitespace, then bereit) in the same sentence as a document filename ("Outline_Template.md liegt bereit" — the confirmed claude-opus-5 prod escape, conv 87c8d619: a completion claim carrying no card term, no download-offer term and no gespeichert/saved participle, so every other branch misses it and the turn ended with neither a card nor an honest notice). It is German-only by design: the English "is ready" and German "ist fertig"/"ist bereit" are generic completion adjectives that routinely describe an operation over the user's OWN upload ("die Auswertung von umsatz.xlsx ist fertig") or advice ("make sure report.md is ready before you deploy"), so they are deliberately excluded (precision-first). The filename must be the subject of the idiom — it directly precedes <verb> bereit (markdown emphasis and whitespace allowed between), with a document extension from the default CLAIM_DOC_EXT (so a .md counts — unlike the delivery branch's office-only set). Subject-adjacency is what keeps it precise: it does not fire when a genitive/prepositional OBJECT marker sits before the filename — either directly or with one intervening noun — from a closed class (von/vom, the genitive/dative determiners des/der/dem/den/dieser…, the genitive-inflected possessives Ihrer/des/seiner…, and the object-prepositions zu/zur/zum/aus/mit/bei): "die Auswertung von umsatz.xlsx liegt bereit", "die Auswertung der Datei umsatz.xlsx liegt bereit", "die Zusammenfassung des Dokuments bericht.pdf steht bereit" — the analysis is ready, not a produced file. The NOMINATIVE determiners die/das are deliberately NOT markers, so apposition to a real subject still fires ("die Dateien report.md und notes.csv liegen bereit"). It also does not fire on a filename that sits after the idiom ("Folgende Vorlagen liegen bereit: vertrag.docx" — a RAG list of files that already exist), nor when excluded by a same-line 2nd-person / modal / how-to token or — in the claim sentence — by a negation (the contiguous idiom also self-excludes "liegt nicht bereit"), a first-person offer, a 2nd-person save-attribution ("Die von Ihnen bereitgestellte report.pdf liegt bereit"), or a final question mark. Named residuals (each costs one bounded steer leg): a bare possessive with no attribution verb ("Ihre umsatz.xlsx liegt bereit" — nominative Ihre, the file is the subject) FIRES; a genitive object reached by TWO+ intervening nouns ("die Auswertung der neuen Datei umsatz.xlsx liegt bereit") FIRES (a full fix needs German case parsing, which no branch here does); a non-contiguous idiom ("liegt zum Abruf bereit") and the German verb-BRACKET where the subject follows the verb ("steht die datei.pdf bereit") are MISSED. This is a BACKEND-only steer — unlike the SPA phantom-notice detector, it has no frontend twin: the lead-less idiom cannot be carried precisely by that detector's lead-anchored, exclusion-free design, so the authoritative backend re-drive is the fix. Fenced code blocks (or ~~~) are not scanned by any of the (g) sub-branches** — a quoted script or log line (*"`print('report.pdf saved')`"*) is not a delivery claim. A fence is three or more markers and closes only on the same character, so an inline-code span never opens one and a-fence is not closed by a ~~~ line inside it; an UNCLOSED fence suppresses every later claim (fail-safe). 4-space-indented code, inline single-backtick code, quoted log prose and a blockquoted fence are not fences (accepted residuals). URLs are removed before the term, format-token and filename scans of all three (g) sub-branches, so neither a vendor help link nor a foreign path can become a claim or the recovered name. All three (g) sub-branches run before branch (a) (the card-vocabulary + filename branch), so when a turn matches both, the recovered filename is the delivery sentence's own — or, when that sentence names none, no filename, and the steer lands in write_file's filename-required branch (all carry the same honesty clause; the readiness sub-branch always names one, since it requires a filename). The corrective write_file error text denies the invented location without denying the product's real delivery paths ("There is no export folder the user can browse: a file reaches them only as a download card from a write_file or code_interpreter run that succeeded (a file in another program is unaffected)." — stated about the product, never about this turn's deliverables: the same filename-validation branch is reachable after a successful write) — the model must not be told that a Word or Lexware export folder does not exist, nor that the sandbox's own /output path (which the code_interpreter tool description tells it to write to) is unavailable. Accepted residuals of the folder/delivery branches (each costs one bounded steer leg): a third-party-software answer naming an export folder without a path or how-to marker; an undated statement of where an existing office document is stored ("ist im SharePoint gespeichert"); third-party attribution ("Word hat die Datei als PDF gespeichert"); a memory/settings sentence carrying a bare format token; indented or inline-backtick code, quoted log prose and a blockquoted fence (only fenced blocks are skipped); a how-to answer under a heading that carries no how-to vocabulary ("## Wie Sie ein PDF erzeugen"), which therefore cannot arm the carry; drafted correspondence or a translation quoting a save; and a possessive 2nd-person provenance line ("Ihre Datei wurde als angebot.pdf gespeichert"). In the other direction (missed phantoms, equally deliberate): a claim within three lines of a how-to token, a claim sentence that happens to pair a 2nd-person pronoun with an attribution verb, and zwischengespeichert. Rejected (no steer): any turn that actually produced a downloadable artifact — a write_file card, a persisted written-file id, a code_interpreter output data file, a live ephemeral artifact, or a generated image (the airtight turn_produced_deliverable gate, tested by emptiness* so a failed image/CI attempt that only lazily-inits its registry still steers a phantom); a same-line modal / how-to / 2nd-person token; a bare "attached above" without card vocabulary (ambiguous with a user-upload reference). Detection runs on the visible text only — model thinking* about file cards never fires it. The delivery-card vocabulary and the saved participle are shared, test-tied (test_system_prompt_composer F6), with the system prompt's document-delivery wording and the write_file nudge so the three cannot drift; the folder vocabulary is detector-only (the no-file-tool prompt already denies an export folder). The full fire / no-fire / accepted-residual corpus is executable: test_inline_tool_calls C12–C17.
  • Observability: file_sentinel_recovery / file_claim_no_artifact WARN logs + trace steps (filenames gated behind include_bodies, as for the drop guards).

Errors

Inference errors use a structured JSON envelope and set the X-AIG-Error response header to the error code:

{
  "error": {
    "code": "invalid_request",
    "message": "Human-readable detail about what went wrong"
  }
}

Branch logic on the stable code, never on the human-readable message. The full list of codes (HTTP status, cause, and the rate_limited / quota_exceeded / guardrail_blocked specifics) is in Error codes.

An error raised after the stream has already opened

On a gateway with a PII detector, a streaming turn's response is committed early — the gateway emits a {"aig_status":"scanning_pii"} progress event before it starts scanning, so the client sees HTTP 200 and Content-Type: text/event-stream immediately.

If the request is refused after that point but before the model is called, a fresh HTTP status is no longer possible. The gateway therefore delivers the error over the open stream and ends it normally:

data: {"aig_status":"scanning_pii"}

data: {"aig_status":"provider_error","error_class":"forbidden","message":"…","user_message":{"body":"…"}}

data: [DONE]

So on a PII gateway a refusal that would otherwise be 403/413 arrives as HTTP 200 plus a terminal provider_error frame. Clients must treat a provider_error event as an error regardless of the HTTP status; error_class carries the same stable code the JSON envelope would have used.

The response status is not the whole story for such a turn, so the gateway records the logical outcome separately: the request log's status column carries the real typed status (e.g. 403), not the 200 seen on the wire, and meta.aig_stream_error holds the error code. Previously nothing was recorded at all and these turns were indistinguishable from successful ones.


See also