Inference API (/v1)
The inference API is the public, OpenAI-compatible surface that applications call to
run model requests through a gateway. Every inference path is served under /v1/.
The gateway accepts the OpenAI chat completions request body for all providers and
translates it to each provider's native wire format automatically.
Data-file preview bounding (request mutation). A user-message text block in the
gateway's inline data-file format — first line exactly [File: <name>] where <name> ends
in a spreadsheet extension (.xlsx, .xlsm, .csv, .tsv, .ods), followed by a blank
line and the extracted text — is treated as a shrinkable preview, not as prose. Before
dispatch the gateway MAY truncate such blocks (appending a visible notice inside the block):
down to the routed model's context window when the request would otherwise be refused with
context_overflow — for self-hosted models the preview budget is computed
worst-case (digit characters count as one token each, since local tokenizers split
digits individually) with the full potential output budget reserved, because those
models share one window between input and output and enforce it with their own
tokenizer — and down to a ~16 KiB head when the raw file is available to the
code interpreter's staged-input channel for this
conversation (the sandbox reads the full file; the preview is orientation only). Blocks in
any other format, assistant/tool messages, and nested tool-result content are never touched.
API clients that need their text verbatim must simply not use this exact block format.
Inline text attachment size. Text attachments (Markdown .md, HTML .html, plain text
.txt) are inlined into the turn as ordinary message text. Accepted: up to 512 KB of
decoded text per such attachment. Rejected: the web app refuses an over-cap text attachment
before the request is sent, with a visible message — it never inlines it (an oversized inline
text would blow the model's context window and dead-turn with no answer). Server-extracted text
that is inlined via POST /chat/files (.docx, .doc legacy binary Word, .pdf, .pptx, image OCR, and a spreadsheet
routed through the inline-extract path) is instead truncated at the same 512 KB ceiling with a
model-visible notice inside the text — not rejected. (A spreadsheet uploaded to the provider's
Files API rides as a raw file_id, not inline text, so it is neither inlined nor truncated.) Note
there is no dedicated server-side byte cap on inline text that arrives as plain message content
(it is indistinguishable from a paste): the server's size backstop is the window-relative context
gate, which now auto-compacts an over-window turn (summarizes older history, head+tail-clips the
current message, and proceeds — see context_overflow) rather than refusing, so an
oversized paste no longer dead-ends the turn; the context_overflow 413 fires only for a single
message that alone exceeds the window even after clipping (and is skipped when the routed model has
no known context window), plus the upstream provider's own limits. The 512 KB attachment cap is
therefore a client-side bound; direct API callers should size their own request bodies.
Inline binary (image / PDF) and total request-body size — client-side CDN guard. A
vision-capable or PDF-capable (Anthropic) model receives images and PDFs as native inline
base64 (image_url / a document block) carried in this request body, not via
POST /chat/files. Because the whole request body — the base64 media plus the entire
conversation history — travels to this inference host, whose edge caps the request body
at 11 MB (a body above it is dropped with a 413; behind the CDN, that 413 historically
carried no CORS headers, blinding the browser's fetch to an opaque network error), the web
app enforces two client-side guards before sending: (1) a per-attachment limit of about
7.4 MB of raw file bytes (~10 MB base64, conservatively under the 11 MB inference edge) — an
oversized inline image is first downscaled + re-encoded in the browser to fit this budget
(the pinned vision models accept far higher resolution than the edge allows — measured HTTP 200 up
to 9.4 MP, with the edge, not the model, being the binding limit — so a fitted image is equivalent),
and only a non-image or an image that still cannot fit is rejected with a specific "…is too large to
upload (max ~7.4 MB)…" message; (2) an assembled-body preflight — if the
fully-serialized request body (media and history) would exceed that budget, the send is
refused with a "This message is too large to send… remove or shorten large attachments, or start
a new conversation." message. Both turn the CORS-blinded 413 dead end ("the gateway is
unreachable, keep retrying") into an actionable message. These are UX guards, not part of the
endpoint contract — the endpoint itself has no application-layer body cap here beyond
context_overflow; direct API callers are not bound by these browser guards, but a body above
the edge cap will still be dropped at the edge. The admin POST /chat/files upload path is a
separate host with its own, larger cap (140 MB edge → a ~110 MB client guard), so a large
document uploaded for extraction is bounded independently from an inline attachment — see
Client-side upload size limit in the Conversations reference.
Base URL: https://<your-gateway-host>
Endpoint paths
Every inference request follows the same four-segment shape:
| Segment | Meaning |
|---|---|
{tenant} |
Tenant slug. |
{gateway} |
Gateway slug within the tenant. |
{provider} |
A registered provider name, or the pseudo-provider compat. |
The {provider} segment selects how the request is routed:
- Pinned provider — naming a real provider (for example
openai,anthropic) and ending the path with/chat/completionspins that exact provider with no model-name guessing. - Compat (unified) —
compatis a pseudo-provider; the gateway guesses the provider from the request body'smodelfield and always returns an OpenAI-shaped response.
For the full list of supported providers, the compat model-resolution tiers, and the native endpoint table, see Providers overview.
Anthropic-native passthrough
Provider paths that do not end in /chat/completions are forwarded to the
upstream provider in its native wire format. The supported case is the Anthropic
Messages API, used by Anthropic-native clients such as Claude Code:
💡 Note: On the native passthrough path the request and response bodies are the provider's own format, not the OpenAI chat-completions shape. Use the
/chat/completionssuffix whenever you want the OpenAI-compatible translation.
Only servable endpoints are accepted
The entire provider path is validated against a fixed allowlist of servable
endpoints before the request is authenticated. Any other path is rejected with
404 endpoint_not_found — the request never reaches the provider and nothing is dialed
upstream. The accepted paths (an optional /v1 version segment is allowed on each) are:
| Path | Endpoint |
|---|---|
/chat/completions |
OpenAI chat completions |
/completions |
OpenAI legacy (text) completions |
/embeddings |
OpenAI embeddings |
/v1/messages |
Anthropic Messages |
/v1/messages/count_tokens |
Anthropic token counting |
/v1/models |
list models (read-only OpenAI-SDK helper) |
POST /v1/{tenant}/{gateway}/{provider}/v1/anything-else → 404 endpoint_not_found
POST /v1/{tenant}/{gateway}/{provider}/internal/x/embeddings → 404 endpoint_not_found
The match is on the whole path, not a trailing suffix: a path such as
/internal-admin/embeddings that merely ends in a servable name is rejected — an
attacker must not be able to smuggle an arbitrary upstream path in front of a known
endpoint. A missing path (a bare …/{provider} URL) and a bare /v1 are rejected
with 404 endpoint_not_found — append an explicit endpoint such as /chat/completions.
An absent path fails closed exactly like a present, unrecognised one: the two share the
reject outcome, never a permissive dial (before, an absent path was forwarded to the
provider's bare base URL and 502'd for every path-forwarding provider). Providers that
build their endpoint from the model rather than the path — Gemini, Bedrock, Cohere, Vertex,
Azure — previously tolerated a bare URL; they now require an explicit endpoint too, for one
uniform rule. This is an authorization boundary, not a
convenience check: every real provider forwards the path verbatim into its upstream URL,
and the platform-managed myra provider attaches Myra's platform credential, so an
unvalidated path would be a server-side request forgery. The allowlist is exactly the set
of priced/known endpoints, so no accepted request escapes usage metering.
💡 Note (model-list probes). The read-only
…/modelslist endpoint stays supported on both the explicit-provider path (/v1/{tenant}/{gateway}/{provider}/v1/models) and/compat(…/{gateway}/compat/models), so an OpenAI-SDK client that auto-probes/modelson start-up keeps working. Only genuinely unknown paths are rejected.
The Messages API is not available under /compat
The /compat surface speaks the OpenAI schema in both directions — requests are read
as OpenAI chat-completions and responses are always OpenAI-shaped. An Anthropic Messages
request sent to a compat sub-path is therefore rejected with invalid_request (400)
rather than partially processed:
The same rejection applies to /compat/…/v1/messages/count_tokens. Use the native route
for an Anthropic-native client, which is where the Messages wire format is actually served:
This also rejects the common SDK misconfiguration of appending /v1/messages to a base
URL that already ends in /compat/chat/completions.
Usage metering on the native Messages route
/v1/messages is the Anthropic Messages endpoint, and the gateway meters it as such
whichever provider serves it. A provider whose upstream forwards the native path — the
self-hosted myra fleet (served by LiteLLM) is the practical case — relays the Anthropic
response verbatim to the client. The gateway selects its usage reader by the endpoint it
actually dialed and confirms it against the response's own envelope (a message_start or
error frame on a stream, a type: "message" object on a stream: false turn); an
upstream that answers that path in another dialect keeps that provider's own reader. On a confirmed Anthropic
response the recorded leg carries:
input_tokens/output_tokensfrommessage_startandmessage_delta(or the Messages object'susage), plus thecache_creation(5m / 1h) andcache_readbuckets — the same counters a Claude-served turn on this route records, read with the same rules, so spend accrual against the gateway/tenant budget and cap enforcement follow the same path. Prefix-cache parity with/chat/completions: a prefix-cached prompt is booked asinput = prompt − cached,cache_read = cachedon both routes. On/v1/messagesthe relaying LiteLLM adapter nets the cached share out ofinput_tokensand surfaces it ascache_read_input_tokens, which this reader books ascache_read. On/chat/completionsthe gateway reads the OpenAI-compatusage.prompt_tokens_details.cached_tokens(a subset ofprompt_tokens) and performs the same split itself (cached_tokensis netted out ofinputand booked ascache_read). So the same prefix-cached prompt books the same buckets on either route. Two contingencies, both fleet-side (not the gateway): the self-hosted fleet only reports a cached share at all when vLLM prefix-token-details reporting is enabled, and/v1/messagesparity additionally requires the fleet's LiteLLM to mapcached_tokens → cache_read_input_tokens(a stock adapter nets the share out but drops it, leavingcache_read = 0on that route). When no cached share is reported, both routes bookinput = prompt,cache_read = 0. Pricing note: a model with no configuredcache_read_per_1kbills thecache_readbucket at the input rate (an explicit0means free), so splitting the bucket never turns cached tokens free by default;request_log.model_version= the served model the response announced;- on a gateway with payload logging and PII masking, the audit
response_rawof the streamed answer.
The relayed bytes are never altered by this accounting. Two related rules on this route:
/v1/messages/count_tokens({"input_tokens": N}, no content — a success by spec) is recognised by the dialed endpoint on every provider whose upstream endpoint carries the path, so it is never filed as an empty answer.- The gateway's server-side tool loop on the native Messages route runs only where the
provider can read the Anthropic dialect it serves:
anthropicitself, or a routed provider that dials its own endpoint. A provider that forwards the path to an upstream answering in the Anthropic dialect through an OpenAI-compatible reader (myra,groq,together, …) gets a transparent passthrough instead: gateway tools are not injected, and a request carrying MCP connector references is rejected withinvalid_request(400) naming the reason. The rule is decided by the request's primary provider; a fallback provider in the routing chain that would relay the dialect is skipped at dispatch on a tool-loop turn (the primary keeps its tools), never folded with the wrong reader. - The gateway's own inner legs — the tool loop's compaction summary (the
max_rounds_reachedsummarize-and-continue step) and the fetch-and-read helper the agentic fetch tool runs — are OpenAI-shapedstream: falsechat completions whatever the parent turn's wire was. They never dial a parent's native/v1/messagespath: on a path-forwarding provider that path is dropped for the inner leg in favour of the provider's canonical chat endpoint (e.g. a native-Messages parent onanthropic,cohere,geminiorbedrockwhose agentic fetch runs on themyradefault model), while a parent that is already on a chat-completions path keeps it (the inner leg dials exactly the parent's endpoint). The parent's path is restored once the inner leg returns, on every exit. - A tool loop that the gateway has to stop never dead-ends the turn with a bare note. Before giving up it makes ONE final no-tools model call (tool calling disabled) and delivers that answer as normal assistant content, followed by an honest note:
repeated_tool_call(a tool kept returning the same or an empty result and the model re-issued an identical call) → the model answers from its own knowledge, with a note saying so. It does not tell the user to rephrase.max_rounds_reached(the 25-round per-turn tool-call cap) → the gateway first tries one conversation compaction (summarize_and_continue); when that cannot finish the task it delivers a best-effort answer and emits the typedtool_loop_terminatedevent withcan_continue: true, and the note is actionable ("Ask me to continue …"). The SPA turnscan_continueinto a one-click Continue affordance that resumes the same task in a fresh turn; the per-turn 25-round ceiling is unchanged and the fallback adds at most one bounded model call per turn.wall_clock_exceededkeeps its existing partial-synthesis behaviour. Every synthesis path fails closed: if the extra call errors or produces nothing, the turn ends on the honest stop note rather than a fabricated answer.- On a stream, the final
message_deltausage snapshot is cumulative and authoritative: a positiveinput_tokens/ cache figure there supersedes themessage_startsnapshot (on an Anthropic server-tool turn the cumulative input is the billed one); a zero or absent figure never overwrites a valuemessage_startalready reported. Usage values are untrusted provider output: a numeric string is coerced, and a non-finite, negative or non-numeric value reads as "not reported".
Accepted compat sub-paths. /compat/chat/completions (and its /compat/v1/chat/completions
spelling) for chat, plus the non-chat OpenAI endpoints such as /compat/embeddings and
/compat/models, which are forwarded to the provider's corresponding OpenAI endpoint. The
Anthropic token-counting endpoint (/v1/messages/count_tokens) is a separate, non-streaming
endpoint and is unaffected by the rejection above.
⚠️ Response-phase guardrails on a native
stream: truerequest. When the gateway has a response-stage detector that cannot be applied incrementally — a content-safetyblock, aregex/presidioscrub, or aflag(i.e. anything other than the inline PII token-restore performed bypii_protector/custom_pii) — a nativestream: truerequest is internally buffered: the gateway collects the full upstream answer, runs the detector, and then re-emits it as native Anthropic SSE events (message_start→content_block_*→message_delta→message_stop; no OpenAI[DONE]). The first byte therefore arrives only after the full generation completes (bounded bytimeout_ms), exactly as for a nativestream: falserequest — the latency cost of response-stage scanning. A gateway whose only response detector is an inline PII masker keeps streaming incrementally. A response-stage block on a native stream is delivered as an HTTP400guardrail_blockederror (the answer is never sent). SSE clients that enforce a short inter-event idle timeout should allow for this buffered window.⚠️ Mid-stream provider failures on a buffered turn. When a PII-masking or guardrail configuration forces a
stream: trueturn onto the internally buffered path (or astream: falsetool-loop turn runs buffered by nature) and the upstream provider's stream fails mid-generation — a mid-stream error event such as Anthropicoverloaded_error, a read error, or a truncated stream — the gateway never answers with a clean HTTP 200 carrying an empty message. Instead it: (1) retries the attempt transparently (same provider, then the fallback chain) when nothing was delivered and no server-side tool has executed; (2) fails loud with the typedprovider_error/all_providers_failederror (as an error frame +aig_status: "provider_error"event on an already-open SSE stream) when retries are exhausted or unsafe; or (3) salvages a partial — any text, files, or images the turn already produced are delivered and persisted together with the explicit error surface, which is the same one the live streaming wire carries for that failure kind: a provider error event (provider_error) is the error frame plus theaig_status: "provider_error"banner event; a read error or a truncated stream (stream_errored/truncated) is the single terminator error frame (provider_error"connection failed" /stream_truncated), no banner, and the visible answer ends with the persisted "The response was interrupted before it completed. Please try again." note described under Visible truncation note — so the salvaged turn reloads exactly as it rendered live. (A leg that reached a natural end and then hit a read error keeps the frame and gets no note — the answer is complete.) A non-streaming JSON response carries the samecontent(the note included) plus anX-AIG-Errorresponse header whose value is the failure kind —provider_error,stream_errored, ortruncated, the same lowercase vocabulary the typed error codes use. A machine client that consumescontentas the bare answer must readX-AIG-Errorto tell a died-mid-answer turn from a completed one; the note is English-only gateway prose, not localized.
Authentication
Send a token on every request unless the gateway has authentication disabled. The token may be a gateway token or a personal access token.
The gateway reads the token from the first header present, in this order:
| Order | Header | Form |
|---|---|---|
| 1 | x-aig-token |
Raw token value. |
| 2 | Authorization |
Bearer <token>. |
| 3 | x-api-key |
Raw token value. |
curl -s -X POST "https://<your-gateway-host>/v1/myapp/prod/compat/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <token>" \
-d '{
"model": "claude-opus-4-6",
"messages": [{"role": "user", "content": "Hello"}]
}'
A missing or invalid token returns 401 unauthorized; an expired token returns a
distinct 401 token_expired and a revoked token a distinct 401 token_revoked
(each with an actionable message naming where to regenerate it). Branch your refresh
logic on the stable code, not on the shared 401 status. See
Authentication and Error codes.
Inference is refused for viewers
A request bound to a user whose role is viewer is refused 403 forbidden ("Viewer role
cannot make inference requests") — before any model is called, so no credit is spent. This holds on
both user-bound auth paths: a raw /v1 call and the chat / /easy playground path (the busier
channel, where a viewer would otherwise burn trial credit). The check fails closed on an
unconfirmable role too: a user whose role cannot be resolved to a known string (a legacy
role-assignment gap) is barred exactly like a viewer, never allowed through by default. A GDPR
Art. 18 processing restriction is evaluated first, so a restricted viewer receives
403 processing_restricted (the legal-hold code), not this one. Other roles — member,
demouser (limited inference via its demo quota), ki_manager, tenant_admin, admin — are
unaffected. The mirror control on the admin session plane (a viewer's chat writes) is
Viewers cannot spend credit or reshape a conversation.
Request body
All providers accept the OpenAI chat completions body. The common fields are:
| Field | Type | Notes |
|---|---|---|
model |
string | Required. Must be a non-empty string. On compat, this also selects the provider. |
messages |
array | Required. A non-empty array of OpenAI message objects (role, content). |
stream |
boolean | When true, the response is a Server-Sent Events stream. |
max_tokens |
integer | Upper bound on generated tokens. Optional — see the default policy below. |
tools |
array | OpenAI tool/function definitions. |
max_tokens default policy
max_tokens is a ceiling, not a target — the model may stop earlier. The
gateway applies a single catalog-driven policy at dispatch:
- Omitted (or
≤ 0) on a self-hosted model (our own vLLM fleet): the gateway fills it from the model's catalog output ceiling (model_price.max_output_tokens). This prevents the serving stack's own tiny per-model default (a thinking model can otherwise spend its whole budget on hidden reasoning and return an empty answer). The injected budget is bounded to a sane output ceiling and window-fitted against the estimated prompt; when only a sliver of room remains (fewer than ~512 tokens) the ~512-token useful floor is injected instead of a starvation budget, and when no safe room remains at all nothing is injected and the serving default applies. If the catalog has no ceiling for the model, nothing is injected and the serving default applies. - Omitted on a third-party provider (Anthropic, OpenAI, Bedrock, OpenRouter,
…): the gateway does not inject a value — the provider applies its own
default and enforces its own minimum/maximum. (Anthropic, which requires
max_tokens, is filled from its own per-model ceiling — on both the OpenAI- compatible and the native/v1/messagespaths.) - Larger than the model's real ceiling (any provider): the value is clamped
down to that ceiling so the provider does not reject the request. Exception: on
the native Anthropic
/v1/messagespassthrough path an over-large value is forwarded verbatim (a genuine Messages client may know a newer model than the gateway's catalog, so clamping to a possibly-stale ceiling would truncate its intended budget); Anthropic then applies its own limit. - Within the ceiling: an explicit value is sent verbatim — except for the self-hosted window budgeting below.
Native context management / compaction (context_management)
Anthropic's native context-management field on /v1/messages — context_management:
{"edits": [{"type": "…"}]} — is untrusted client input and is sanitized at the
gateway boundary. The compact_20260112 compaction strategy is only valid on models that
declare support for it (sonnet-class; haiku does not), and Anthropic returns a 400
if it is sent to an unsupported model.
- Accepted: a
context_managementwhoseeditsis a JSON array. Acompact_20260112edit is forwarded only when the resolved (post-routing) model supports it. Other strategies (e.g.clear_tool_uses_20250919) are separate features and pass through untouched. - Rejected / stripped (fail closed): when the resolved model does not support
compact_20260112, the gateway strips that edit fromcontext_management.edits(and removes thecontext_managementfield entirely if no edits remain) and removes thecompact-2026-01-12anthropic-betatoken — so the gateway never forwards a request Anthropic would 400. The routed model, not the client'smodel, is authoritative (the gateway may route a compaction-capable request to an incapable model), so this adapts rather than erroring. Acontext_managementwhoseeditsis not an array (an object, a scalar, or absent) is left untouched — a malformed body is the client's own 400. /v1/messages/count_tokens: this sub-endpoint accepts nocontext_management(any model, any strategy); the gateway strips it wholesale.- Gateway-injected compaction merges, never clobbers: when the gateway itself enables
compaction (
context_compactionconfig) for a supported model, it appends itscompact_20260112edit to the client's existingcontext_management.editsarray (deduped), preserving any client edits — it does not replace the field.
Self-hosted window budgeting (our own vLLM models)
A self-hosted model enforces a hard rule: input_tokens + max_tokens must not
exceed the model's context window (max_model_len). A large max_tokens on top of
a large prompt therefore overflows the window and the request is rejected, even
though the input alone fits. For our self-hosted models only (never third-party
providers, whose windows are much larger and are not policed this way), the gateway
keeps input and output inside the window automatically:
- Pre-flight clamp. Before dispatch,
max_tokensis reduced so that the estimated prompt plus the output budget fits the window. The result is only ever smaller than the value you sent — a request that already fits is untouched. If the input is so large that no useful output budget remains (below ~512 tokens),max_tokensis left as-is rather than shrunk to a useless near-zero value. Inline data-file previews are additionally window-fitted pre-flight in a worst-case token frame with the full potential output budget reserved (see Data-file preview bounding above), so a near-window preview is shrunk before the serving stack can reject the request. - Overflow repair (one resend). If a request still overflows (the estimate
under-counts, or a multi-step tool loop grew the prompt), the model returns a
context-window error. When that error shows the input still fits the window,
the gateway resends the request once with a reduced
max_tokens; your full input is preserved and the answer is generated. The budget depends on what the error actually reports: - a measured input count (the serving stack printed its real tokenizer
count): the remaining budget minus a 128-token safety margin
(
window − measured input − 128) — the margin absorbs re-count drift between serving layers; - an "at least N input tokens" figure: this is a derived lower bound (the smallest input consistent with the rejection), not a measurement, so no remaining-budget arithmetic is trustworthy — the gateway resends with the minimum useful output budget (~512 tokens) instead, which succeeds whenever a useful rescue is possible at all. The answer may be shorter than usual on this last-resort path.
- Conservative fallback when the counts are unreadable. If the overflow error is recognisable as a context-window rejection but its token counts cannot be safely parsed (missing, malformed, or implausible numbers), the gateway still resends once, with a budget computed purely from its own prompt estimate and the model's catalog window, reduced by a further 4096-token margin and always strictly below the budget that just failed. No number from the error body is used on this path. If no useful budget remains (below ~512 tokens), the request is not resent.
- Genuine input-overflow is not repaired here. If the input alone exceeds
the window (no output budget to reclaim), the request is not resent — it surfaces
the standard
context_length_exceedederror (see Error codes).
The window-error bodies the gateway parses are untrusted model/serving output:
only the specific captured vLLM/LiteLLM phrasings are recognised, and any
implausible value (a count ≤ 0 or > 10,000,000, an input that meets or exceeds
the window, or a missing input count) is rejected — those bodies take the
conservative-fallback path above (which consumes no numeric value from the
error body) or, failing its guards, the original error is returned unchanged. In
every case the resend happens at most once — a second overflow surfaces the
standard error.
Validation
The gateway validates the request body at the trust boundary before any
inference runs. A request is rejected with 400 invalid_request when:
- the body is empty (
Empty request body) or is not valid JSON (Invalid JSON body); - the body is valid JSON but not a JSON object — a bare scalar, boolean, or
nulltop-level body (e.g.123,true,null) is rejected (Request body must be a JSON object). (A top-level JSON array is likewise not a valid request; it is rejected on the requiredmodelfield below.); messagesis present but is not an array ('messages' must be an array);messagesis an empty array[]('messages' must not be empty) — an empty array is a malformed request, not an empty conversation, and never reaches a provider;- any element of
messagesis not an object (each 'messages' entry must be an object); modelis missing, empty, or not a string (Missing or invalid 'model' field).
A JSON null messages value is treated as absent (equivalent to omitting the
field), not as an empty array.
Trailing assistant turns are normalized (Anthropic)
A well-formed chat request ends with a user turn. For Anthropic models on the
compat path, if the conversation ends with one or more trailing assistant turns
(an accidental assistant-message prefill — never intended in a chat flow), the gateway
drops the trailing assistant turn(s) so the array ends with the preceding
user/tool-result turn before the request is sent. This is required because the
Anthropic 4.x reasoning family rejects a last-assistant-turn prefill with a provider
400 ("this model does not support assistant message prefill; the conversation must
end with a user message"). The normalization is trailing-only — interior assistant
turns and tool-call/tool-result pairings are untouched — and a conversation that is
entirely assistant turns (no user turn at all) is left unchanged and rejected by
the provider. It applies only to the compat path; the anthropic native passthrough
forwards the body verbatim (where an intentional prefill is the caller's own choice).
Per-provider nuances — system-prompt extraction, extended thinking, prompt caching, model name prefixes — are covered in Providers overview and the individual provider pages.
chat_template_kwargs — the adaptive-thinking flag
Accepted shape: an optional object {"enable_thinking": true|false} (the vLLM/Qwen
chat-template knob; the Chat sends it to every non-Anthropic model). Only the strict boolean
false is read as "thinking off"; the gateway translates or drops the field per provider and
never forwards it where it would be rejected. Strict cloud providers (Mistral, Groq, OpenAI,
Cohere) never receive the field itself; where the model has a reasoning knob the OFF signal is
translated instead — Magistral and Groq's gpt-oss/qwen3 get reasoning_effort, Cohere's
command-a-reasoning-* gets thinking: {type: "disabled"}, Gemini 2.5 gets
thinkingBudget: 0, OpenRouter gets reasoning: {effort: "low"}; models without a knob
(OpenAI direct, Gemini 2.0/3.x, Mistral Large, …) get no translation. A Myra-served
Mistral-tokenizer model (mistral-small-4) receives neither the field nor a translation (see
Myra). Qwen and Gemma fleet models receive it unchanged.
Rejected/ignored: a null, non-object or otherwise malformed value is treated as absent by
every adapter (one shared null-safe test — no translation, and the field is still dropped where
the provider would 400 on it); it never causes a gateway error. A reasoning_effort you set
yourself always wins over the translation.
Retired model ids upgrade to their successor (Myra fleet)
When the Myra fleet retires a self-hosted model in favour of a newer one, the retired id is
deprecated in the catalogue (GET /model-prices shows deprecated_at) and would normally be
refused at serve time with 400 model_not_found (see Models). For the fleet's own
retired ids the gateway carries a successor map, so an integration that pinned the old id
keeps working instead of hard-failing on the day of the rollover:
Requested model |
Served as |
|---|---|
qwen3.6-27b, qwen3.6-35b-a3b |
qwen3.8-27b |
gemma-4-26b-a4b-it |
gemma-4-31b-it |
Accepted shape. The rewrite applies to the model field of an inference request — bare
(qwen3.6-27b) or provider-prefixed (myra/qwen3.6-27b), on the provider-native myra route
and on /compat/... — once the id has been normalised exactly as any other id is (provider
resolved, prefix stripped), to a routing rule's actions.model and each
fallbacks[].model when the rule remaps onto a retired id (the same resolver; an upgraded
fallback target is only logged — the response headers and the request log always describe the
model that served the turn; a rule whose model condition names a retired id simply no longer
matches, see Routing rules), and to a gateway's agentic_fetch.model
override (the inner fetch agent runs on the successor; no response signal, the inner leg does not
serve the turn). It happens before cost, capability checks and the request log read the
model, so every downstream surface names the model that was actually served: request_log.model, the cost figures and the analytics rows
carry the successor; the requested id is kept for attribution in request_log.meta under the
gateway-owned key aig_model_upgraded_from (a client cannot set it — x-aig-meta-aig_* headers
are dropped like every other reserved key).
Stored model no longer permitted. On a self-serve workspace, a run whose model comes
from stored configuration — an agent's model (/agents/{slug}/invoke, scheduled runs,
workflows, delegated sub-agents), a scheduled prompt task's model, the Copilot document-AI model —
that the plan or the EU-Gov add-on no longer permits runs on the pinned gateway's Auto choice
instead of failing (see Tenant entitlements). request_log.model and the
cost name the model that ran; the stored model is kept in request_log.meta under the gateway-owned
key aig_substituted_from. A model the caller sends is never substituted.
What the response says. Every response to a rewritten request — buffered or streamed, and
including any error raised after the rewrite (a guardrail block, a plan or gateway refusal of
the successor, a provider error) — carries the headers below. Responses produced before the
model is read — authentication, rate limiting, the tenant/token budget 429, the IP allowlist,
and an exact-match cache hit — never do.
| Header | Value |
|---|---|
X-AIG-Model-Upgraded-From |
the retired id the request — or the matching routing rule's rewrite — named (qwen3.6-27b) |
Deprecation |
RFC 9745 form @<unix-seconds> — when the retired row was deprecated. Sent only when the client's own model was retired (a routing rule's rewrite and an agent's stored model — on /agents/{slug}/invoke and scheduled runs — are the tenant's configuration, which the caller cannot act on: those carry X-AIG-Model-Upgraded-From only), and absent when the catalogue no longer has the old row at all. The RFC defines the field for the resource; here it qualifies the requested model, and it only ever appears together with X-AIG-Model-Upgraded-From — the endpoint itself is not deprecated |
X-AIG-LLM-Model (buffered) / X-AIG-Model (streamed) |
the successor that served the turn |
Both new headers are CORS-exposed to browser clients.
What is NOT rewritten (each fails closed toward the existing behaviour — the deprecation
gate answers 400 model_not_found, or the block 403 model_disabled_on_gateway):
- an id with no entry in the map — the gateway never guesses a successor from the name
(
qwen3.5-27bstays a400); - the same id on a non-Myra provider (
openrouter/qwen3.6-27bis a passthrough serve, not the fleet); - a retired id the tenant admin disabled on this gateway (Disabling a Myra model) — a rename never sidesteps an admin's block;
- a retired id whose catalogue row is live again (an operator un-deprecated it, or the fleet re-serves it) — the operator's decision wins, the request goes to the id as named;
- a successor that is itself not a live catalogue row (the map is single-hop; chains are not followed);
- platform-owned targets applied after routing: a self-serve plan's
fallback/degrade_modelnaming a retired id is still refused at the serve-time gate (400 model_not_found) — platform operators update those at a rollover; - a catalogue read that fails — the request proceeds unchanged (and the deprecation gate, which fails open on its own read, decides).
Two consequences of "the successor is what the request now names": a plan's model allowlist and a
gateway's disabled-model list are checked against the successor — a request for qwen3.6-27b
on a gateway where the admin disabled qwen3.8-27b is answered 403 model_disabled_on_gateway
naming qwen3.8-27b (with X-AIG-Model-Upgraded-From on the response), and a plan that lists
only the retired id refuses the successor by name. At a rollover, update plan allowlists and
gateway blocks to the successor id.
The exact-match response cache is keyed on the id as requested, so
qwen3.6-27b and qwen3.8-27b are separate cache entries; a cache hit on the old id serves the
stored answer without dispatch and therefore without the two upgrade headers (the entry records
the model that produced it).
Stored picks are covered too: a project default model, an agent's model or a conversation's
pinned model that names a retired id resolves to the successor at dispatch time through the
same resolver (the stored value itself is not rewritten — the conversation keeps showing the
retired id until the user picks again, while every turn it sends is served on, logged as and
billed to the successor), and /compact (its summary is recorded under the model that wrote
it), the save-time pin checks and a scheduled task's plan check use that resolver too, so they
agree with what an inference request would serve. A soft pin — a project default model without
a gateway, or a user's personal default model — that names a retired id is resolved to the
successor when a new conversation is created (a retired Gemma pin lands on the Gemma successor,
not on the Auto pick); that ladder honours a block recorded under the old name on a gateway
(a sibling gateway without the block serves it — gateways stay independent; with no such
sibling the pin degrades to the Auto pick, as an unmapped pin does) as well as the successor's
own availability. Note for the chat UI: a
conversation whose stored pin is a retired id shows the picker label Auto (the catalogue
has no row for it) while its turns are served on, and labelled with, the successor; picking a
model again stores the live id. The pre-rollover model unavailable recovery bubble does not
appear for a mapped id — the transparent upgrade replaces it.
Empty responses & inline dead-turn retry
A model can finish a turn cleanly (finish_reason: stop) yet produce no visible
answer — 0 output tokens, a few tokens of internal reasoning with nothing user-facing, or
a handful of whitespace-only tokens (the instant-EOS "dead turn" some self-hosted models
emit on certain prompts: a \n\n preamble, then stop, before any real token). A visible
answer that is empty or contains only whitespace (matched against the same whitespace set as
JavaScript's String.trim(), so a \n\n or an ideographic-space answer counts as blank)
is treated as no answer; any substantive character (e.g. a terse 4 or Ja.) is a real
answer and is delivered unchanged. The gateway substitutes a short fallback message in
content and records the event for the per-release empty-rate metric. The wording is chosen on the response's
output_tokens (the completion-token count, which for a reasoning model includes its
reasoning tokens): > 0 → an empty-answer notice; == 0 → "The model returned an
empty response. Please try again or rephrase your prompt." For the > 0 case the notice is
further split by whether the turn actually produced reasoning content: a reasoning turn
that spent its budget on hidden reasoning gets "The model produced internal reasoning but no
visible answer. Try rephrasing your prompt or turning off Reasoning…", while a thinking-off
turn that simply emitted no visible content (the instant-EOS dead turn — no reasoning at all)
gets "The model ended the turn with no visible answer. Please try again or rephrase your
prompt." — the "turn off Reasoning" advice is dropped because reasoning was never on. On a
streamed OpenAI-compat response the count is only final after the trailing usage chunk, so
the classification is made once the whole response has been read. The fallback applies
whether or not tools were offered on the turn: a first-party conversation turn (where the
file tools are offered whenever the model can call tools) that finishes with a whitespace-only
answer gets the same fallback as a tool-free request. It is the leg's terminal fallback — it
is not added to a leg the turn is about to continue or recover: a tool call the loop will run,
a length-cap continuation, the broken-promise nudge (a file was asked for and none produced —
the nudged leg answers; a second blank leg, or a nudge that never runs because its dial failed
or the loop's budget ended the turn first, gets the one fallback — after any earlier leg's text),
or a turn that
already produced a deliverable (the bare re-download card, or a file / image /
code-interpreter artifact earlier in the loop). On a streamed turn a blank leg whose read
failed after its finish chunk still gets it (nothing else covers that bubble); on a buffered
turn that cut attempt is re-dialled or reported like any other failed attempt (see the
buffered-failure rules below), and a leg that ended with a provider error event gets that
banner, not this fallback. Three
tool-related shapes get something else: a leg whose visible content is empty because an
inline tool call was stripped and recovered continues the loop into the real answer; a leg
whose only content was a tool tag the gateway could not turn into a call gets the
"unparseable tool" retry hint below; and a leg whose only content was a stripped non-tool tag
(a <memory> proposal, which was delivered as its own event) gets no wording at all — the
model did produce content, so "no visible answer" would be false (the gateway logs the strip).
A turn the client aborted is recorded as aborted, not as empty.
A third case takes priority on a max_tokens (length-capped) empty turn: when the
prompt itself ~filled the model's context window (prompt_tokens ≥ 90% of the routed
model's known context window), the turn produced no answer because there was no output
budget left. The gateway returns a distinct fallback whose wording follows the turn count of
the request — on a first turn (exactly one user message in the request, so the prompt is
the conversation) "Your input is too large for this model's context window. Shorten it, remove
attachments, or pick a model with a longer context."; on any later turn "This conversation is
too long for the model's context window. Start a new chat or shorten the conversation, then try
again." — and, because
a same-prompt retry would re-send the near-window prompt and deterministically fail again
(and re-bill the input), suppresses the retry — the inline dead-turn retry below never
fires for this class, and the chat UI hides its regenerate affordance on the bubble. When
the routed model has no known context window (the catalog lacks max_input_tokens) this
class is not selected — the turn falls back to the reasoning-only / empty wording above.
For self-hosted (inline) models a "dead turn" — an empty-visible clean finish with
~0 output tokens (an instant-EOS glitch) — is transparently retried once on a
buffered (stream: false) request before the fallback is returned. If the retry
produces a real answer you receive it normally; if the turn is still dead the fallback
is returned and a provider fault is recorded (and, for a sustained storm of dead turns,
the provider's circuit breaker opens so traffic fails over). The retry is capped at one
attempt and never applies to third-party providers. A dead turn is an instant end of
turn (finish_reason: stop): a max_tokens (length-capped) empty turn is not one — the
model was still generating when the cap cut it, typically because the caller asked for a
tiny max_tokens (a one-token health probe) or, on a model with no known context window,
because the prompt filled it. Such a turn is delivered once with the fallback above, is
never retried (the identical request would end identically), and counts as neither a
success nor a failure for the provider's circuit breaker. Live-streaming (stream: true)
turns are past the point of no return once bytes are on the wire and are not retried.
Broken-promise recovery (Workspace file writes)
On a Workspace turn where the server-side file tools are offered, a model can narrate
an action it never performs: you ask it to create or write a file, and it replies "I'll now
create the file…" and stops — a clean finish_reason: stop with prose but no write_file
call and no file produced. Because a clean stop with visible text is the tool loop's normal
successful exit, the turn would otherwise end there with nothing written. When your typed
request was to write a file and the turn produced no file and called no tool, the gateway
re-drives the turn with a short instruction to use the tool now. The write intent is read from
the text you typed — the contents and filename of an attached document are excluded, so
attaching a file whose text happens to contain a word like "generate" does not trigger the
recovery. Re-drives stop as soon as either a file is written or a nudged attempt again makes
no progress (answers with no tool call), so a turn that has genuinely finished is never
re-driven repeatedly; no answer field of a standard response is altered. This recovery applies
only to Workspace/chat turns (identified by their conversation and turn ids); a plain
OpenAI-compatible client is never re-driven and always receives exactly one finish_reason.
Streaming
Set "stream": true to receive an SSE response. The gateway emits OpenAI-shaped
chunks and terminates the stream with a data: [DONE] sentinel.
data: {"id":"...","object":"chat.completion.chunk","created":1753718400,"model":"...","choices":[{"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"...","object":"chat.completion.chunk","created":1753718400,"model":"...","choices":[],"usage":{"prompt_tokens":676,"completion_tokens":127,"total_tokens":803}}
data: [DONE]
What a plain client receives
A plain OpenAI-compatible client is one that sends neither x-aig-turn-id nor
x-aig-extensions: 1. Every data: frame it receives is either [DONE] or a JSON object
carrying choices (an array — empty on the final usage chunk) or an error object. Nothing
else. This is a hard contract: clients that validate each chunk against the OpenAI schema
(the Vercel AI SDK's @ai-sdk/openai-compatible provider and everything built on it) abort
the stream on the first non-conforming frame.
The gateway's own aig_* extension frames — aig_status (all values, including thinking),
aig_tool_call, aig_tool_input, aig_sources — are suppressed for such a client. They
are a first-party side channel for the chat UI, described throughout this page; a client opts
into them explicitly:
| Signal | Who sends it | Effect |
|---|---|---|
x-aig-turn-id: <id> |
the chat UI, every turn | turn ownership and the side channel |
x-aig-extensions: 1 |
a first-party client with no turn (agent Try-It, the playground) | the side channel only |
Accepted values for x-aig-extensions are 1, true, yes, on — matched exactly, so
TRUE and On do not count. Any other value is legal to send and is treated as absent
(fail closed): 0, an empty value, an unknown word, a non-string, or a repeated header whose
first value is not one of the four. The stream then stays pure OpenAI.
Turn persistence is tenant-fenced
A turn-owning request (x-aig-turn-id) that also names a Workspace conversation with
x-aig-conv-id has its assistant answer persisted into that conversation. Both headers are
client-supplied, so the conversation named by x-aig-conv-id is not trusted as an authorization
grant: a conversation is pinned to the tenant of the gateway that owns it, and the gateway
persists an answer into it only when that conversation belongs to the same tenant as the
gateway the request routed through (the tenant is derived server-side from the gateway, never
from the request body or a header).
- Accepted — the conversation is in the caller's tenant. This includes a conversation reached through a different gateway of the same tenant (a platform admin routing across the tenant's gateways); the answer persists normally.
- Rejected — the conversation belongs to another tenant. The turn still runs and the answer is returned to the caller (it was a legitimate request on the caller's own budget), but nothing is written into the foreign conversation: no assistant message and no streaming placeholder row. From the caller's tenant that conversation does not exist, so the gateway treats it as a ghost — no error status, the answer simply is not saved anywhere the other tenant can see. This closes a cross-tenant write (an attacker cannot inject a chosen answer, or a stalling placeholder, into someone else's thread).
Within a tenant, a second check — the same write-permission rule the Workspace UI applies — decides whether the answer is saved into the named conversation:
- Accepted — the caller is the conversation's owner, or a member of the project it is shared into who currently holds the live-driver turn (the collaborative co-editing lock). This also covers a server-produced artifact (an image or generated file) persisted under an absent or non-conforming turn id.
- Rejected — a same-tenant conversation the caller may not write: not the owner and not the current driver (for example a read-only shared member, or someone with no access at all). As with the cross-tenant case the answer is still returned but nothing is written — no assistant message and no streaming placeholder — so an attacker cannot inject an answer into, or stall the reload of, a colleague's private (or read-only) conversation.
- A service token with no user identity persists into no conversation on this path.
- The driver lock is evaluated at the start of the turn (when the placeholder is claimed), not re-evaluated at save time — so a long-running turn whose lock lapses mid-answer is still saved, never dropped. A transient database error while checking permission is surfaced as a retryable failure (the answer is reported as not-yet-saved), never silently discarded.
Consequences worth planning for as a plain client:
- Errors arrive as frames, not statuses, once the stream is open. A failure detected
after the response head is committed is delivered as
data: {"error":{"message":"…","type":"gateway_error","code":"<code>","param":null}}followed by[DONE]on an HTTP 200 — the full official OpenAIerrorshape (message,type,code, and the nullableparam; the gateway never attributes an error to a request parameter, soparamis alwaysnull). Before the head is committed a failure is still an ordinary HTTP error response. The terminal chunk of such a stream stays inside thefinish_reasonenum (stop), so treat the error frame, not the terminal, as the failure signal. - A long reasoning or tool phase carries no
data:frames. Thethinkingframes that used to fill those gaps are extension frames. To keep the connection alive the gateway writes an SSE comment line (: hb) after 15 s of silence — never while content is flowing. A comment is ignored by every conformant SSE reader (EventSource,@ai-sdk/openai-compatible,openai-python); a reader thatJSON.parses whole blocks instead of following the SSE framing must skip lines that do not start withdata:. - Exactly one
data: [DONE]. The stream is terminated once; nothing follows it. - Exactly one non-null
finish_reason— even across a server-side tool loop. When the gateway runs a tool for you (an MCP connector, web search, the code interpreter, file writes), it may call the model several times in one response. Only the final leg's terminal reaches the wire; the intermediatefinish_reason: "tool_calls"legs, where the gateway executes the tool and continues, are suppressed. So a client that treats the first non-nullfinish_reasonas the end of the turn still gets the whole answer. Server-side tool execution is the only thing that runs several legs for a plain client — the length-cap auto-continue of a long answer (below) is a first-party feature a plain client never triggers, so it never adds a secondfinish_reasonhere. This is distinct from a turn where you supplied the tools (toolsin the request body): there the gateway forwardsfinish_reason: "tool_calls"to you as the legitimate, single terminal — you run the tool and continue the conversation with a new request.
created and system_fingerprint
Every chat.completion.chunk carries created — the completion's creation time as a unix
timestamp in seconds. Every chunk of one response carries the same value (it stamps
the completion, not the frame), and the buffered chat.completion body carries it too, so a
re-serialising proxy (LiteLLM and similar) accepts either shape.
A response served from the gateway's response cache replays the body it stored, so its
created (like its id) is the original completion's — not the time of the cache hit. The
X-AIG-Cache: HIT response header tells you which case you are in.
system_fingerprint is deliberately never sent. It identifies the backend configuration
a completion ran on; a gateway routes one model id across heterogeneous providers, regions
and adapters, so there is no value it could report that would stay true. The field is
optional in the OpenAI schema — treat its absence as intentional rather than as a gap.
The final usage chunk
The last chunk before [DONE] carries token usage with an empty choices array:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":1753718400,"model":"…","choices":[],"usage":{"prompt_tokens":676,"completion_tokens":127,"total_tokens":803}}
One deliberate difference from OpenAI: OpenAI sends this chunk only when you ask for it
(stream_options.include_usage: true); the gateway sends it by default and honours
stream_options.include_usage: false as an opt-OUT.
A turn can run several upstream legs — a server-side tool loop, or an auto-continued long
answer. A plain client receives exactly one usage chunk per response, carrying the turn
total summed across every leg, as the last frame before [DONE]. You may assign it or add
it; both give the same number. Because the gateway sums it rather than relaying one leg's
block, provider-specific breakdowns that do not add up across legs
(prompt_tokens_details, reasoning_tokens) are not carried; Anthropic's
cache_creation_tokens / cache_read_tokens / cache_deletion_tokens are, when non-zero.
The upstream's prompt_tokens_details.cached_tokens is consumed, not forwarded: the
gateway nets it out of prompt_tokens and books it as cache_read_tokens instead (so the
summed chunk reports prompt_tokens = the uncached prompt plus a cache_read_tokens
bucket, the same shape an Anthropic-origin turn already reports on this envelope). Only a
finite non-negative cached_tokens no greater than prompt_tokens is accepted; a
null / non-numeric / NaN / ±inf / negative value is treated as no cache hit
(cache_read = 0), and a value exceeding prompt_tokens is clamped to it.
A client that opted into the gateway extensions (x-aig-turn-id / x-aig-extensions)
instead receives one usage chunk per leg and must SUM them for the turn total — that
cadence exposes the leg structure first-party tooling reports on. It is the only place the
two audiences see different bytes at all; finish_reason, created and every other
standard field are identical for both.
The gateway normalises each provider's stop reason to a consistent OpenAI-compatible
finish_reason:
finish_reason |
Meaning |
|---|---|
stop |
Normal completion (Anthropic end_turn/pause_turn, Gemini STOP, Cohere COMPLETE also map here), and the fallback for anything the gateway cannot classify. |
length |
Output truncated at the token limit (OpenAI length, Anthropic/Gemini/Cohere/Bedrock max_tokens/MAX_TOKENS, Mistral model_length map here). |
tool_calls |
The model requested a tool call (provider tool_use). |
content_filter |
The model refused or was blocked (Anthropic refusal, Gemini SAFETY/RECITATION/PROHIBITED_CONTENT/BLOCKLIST/SPII, Cohere ERROR_TOXIC, Bedrock guardrail_intervened/content_filtered). |
That enum is closed, and it is the same for every client and every egress — the streaming
wire and the buffered chat.completion body, whatever headers you send. An unrecognised or
future provider stop reason is mapped, never relayed verbatim, and a stream aborted
mid-flight reports stop with an accompanying error frame (below) rather than inventing a
value. function_call is accepted as the legacy enum member but the gateway never mints it.
This matters most for agents that continue automatically when a response was cut off: they
key on length. A value outside the enum does not make a strict client abort — it makes such
an agent stop silently, with a half-finished answer and no error.
Internally the gateway keeps a richer terminal vocabulary (that is what drives server-side
auto-continue, the truncation indicator, and the rule that an abnormal turn is never
persisted as a durable summary). It is deliberately not on this field: a standard OpenAI
field carries the OpenAI value, and the gateway's own reason travels on the aig_* side
channel, where extensions belong.
This normalisation now covers every provider. Gemini/Vertex, Cohere and Bedrock previously
reported no terminal reason at all, so a truncated or content-filtered turn on them was
mislabelled a clean stop: the "response was truncated" indicator never appeared, auto-continue
never armed, and a truncated compaction summary could still supersede the conversation history.
Their parsers now surface the real terminal, so these behave like OpenAI/Anthropic.
On the Anthropic-native response shape (/anthropic/v1/messages), the stop_reason field
is always one of Anthropic's own enum values (end_turn, max_tokens, stop_sequence,
tool_use, refusal) even when the request is routed to a non-Anthropic model — a foreign
provider's terminal (Gemini MAX_TOKENS, Cohere COMPLETE, Bedrock guardrail_intervened) is
translated to the matching Anthropic value, and an unrecognised foreign reason resolves to a
clean end_turn rather than being emitted verbatim. A genuine Anthropic leg's stop_reason
(including pause_turn and any newer value) is relayed unchanged.
A long answer that fills the model's context window is auto-continued across legs; if a
continuation leg cannot proceed (the input plus the answer so far already exceeds the
window), the turn ends truncated, and the gateway terminates the stream with
finish_reason: "length" — never a clean "stop" — so the client can surface that the
response is incomplete. Route context-heavy jobs to a larger-window model to avoid
truncation.
A turn truncated at the model's own output limit is reported as length on every egress —
the live SSE stream, the buffered response body, and the buffered re-emit (below).
Web-search turns are included: a search-augmented answer that hits the limit is labelled
truncated, not clean.
Buffered re-emit turns carry the same tool telemetry
A gateway whose only response detector is an inline PII token-restore masker
(pii_protector / custom_pii) streams incrementally on both the native and the
compat (OpenAI-format) wire: PII tokens are restored on the live wire (visible answer,
reasoning, tool-call args, grounded-excerpt quotes, filenames, web-search queries all
carry real values), so these turns are not internally buffered and their tool-loop
telemetry streams live. (On a wholly-local Myra/EU model leg no PII scan/restore happens —
masking is skipped; see PII Protector — Local model legs are not masked.)
A "stream": true turn is still force-buffered internally when a response detector the
live stream cannot satisfy is active — a content-safety block, or a regex/presidio
scrub (mask) — or when the turn runs buffered by nature (a stream: false tool loop).
On that path the whole answer is produced, scanned, and then re-emitted to the client as a
normal SSE stream. That re-emit replays the turn's tool-loop side-channel events —
aig_status values tool_call, tool_result, tool_deferred, tool_error, and
ephemeral_file — in their original order, before the answer content.
A PII-protected turn therefore carries the same tool telemetry as a live stream — to a client
that opted into the aig_* side channel; a plain OpenAI-compatible client receives it on
neither path. With that telemetry a client can tell a genuinely computed code_interpreter
answer from a model estimate, render tool chips, show file-write notes, and receive a
live-delivered generated file.
Bounds (the replayed frames originate from model-driven tool calls and are treated as
untrusted): at most 256 events are replayed per turn, and an unencodable frame is
dropped. An event whose encoded size exceeds 64 KiB is replayed with its args removed
and "args_truncated": true set. ephemeral_file is exempt from both the size check and the
encodability check — its payload is the user's file and every field of it is
gateway-constructed; its size is bounded at the source instead (see below). Statuses the buffered tail already reconstructs from
its own state (thinking, pii_masked, tools_skipped, aig_sources) and mid-stream
control signals (provider_error, compacted, pause_turn) are not replayed — no
event is ever delivered twice, and no stale control signal outlives the turn it steered.
Plain non-streaming requests ("stream": false) are unaffected: their single JSON
response never contains side-channel frames.
When the buffered answer is a tool call — the model returned tool_calls (for example a
client-provided tool, or a web-search direct answer that called the client's own tool) with no
visible text — the re-emitted stream carries those tool_calls as a normal OpenAI
delta.tool_calls chunk followed by a terminal finish_reason: "tool_calls", exactly as a live
stream would. A streaming client is never handed a silent empty turn on a tool-call answer. (If the
tool call's arguments were truncated by the token limit the terminal is length, not
tool_calls — the truncation signal wins. An empty or non-array tool_calls is not emitted.)
ephemeral_file — a generated file delivered live and stored nowhere
When the code interpreter produces a file in a conversation that cannot store files — a private (ghost) chat — including one opened inside a project, where the gateway refuses to write to the project's knowledge base — or any request without a conversation to write to — the file is not persisted and there is no download route to point at. Instead the gateway sends the bytes once, inline on the turn's own stream:
data: {"aig_status":"ephemeral_file","name":"report.xlsx","mime_type":"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet","size":8134,"data":"UEsDBBQ..."}
| Field | Type | Meaning |
|---|---|---|
name |
string | Bare filename, at most 100 characters. Never a path, never a leading dot, never a Windows-reserved character (: * ? " < > |) and never a control byte. |
mime_type |
string | One of the 11 accepted types: text/csv, text/plain, application/json, the .xlsx workbook type, image/png, Word (.docx), PowerPoint (.pptx), PDF (application/pdf), and OpenDocument (.odt, .ods, .odp). |
size |
integer | Size of the raw (pre-base64) bytes. |
data |
string | Base64 of the raw bytes. |
Nothing is written server-side for such a file: no conversation row, no knowledge row, no stored blob, and no id — the event is the file's only existence, and it is gone once the client discards it.
Opt-in. The event is sent only to a client that identified its turn with the
x-aig-turn-id header (the chat UI always does) or opted in with x-aig-extensions: 1. A
plain OpenAI-compatible client that sends neither never receives it, and the generated file
is reported as not shown instead. This is a delivery contract, not an authorization boundary.
Bounds. A turn delivers at most 4 MiB (4 194 304 bytes) of raw bytes this way by default — the effective per-turn cap is the tenant's code-interpreter artifact-size limit, which an admin can raise up to a 25 MiB ceiling — and at most four files plus four images (the same per-turn budgets the stored path uses). An artifact that does not fit, or that cannot be delivered (for example on a non-streaming request), is reported to the model as "could not be shown to the user" rather than silently dropped.
Client-side validation. A consumer must treat this event as untrusted, exactly like every
other stream frame. The chat UI rejects — and renders nothing for — an event whose name
contains a path separator, a control character, or a leading dot, whose mime_type is
outside the list above, whose size is not a positive integer within the budget, whose
data is not plain base64, or whose size and data length disagree (padded base64 length
is a function of the raw length, so the two must match). The filename rule is the same list the
gateway applies before emitting — a path separator, a Windows-reserved character, a control
byte, a leading dot, or a name over the byte ceiling is REJECTED, never "repaired" into a
different file, and yields no download card. Because both sides apply the same list, an
artifact the gateway delivers always renders; one whose name would fail it is never
delivered at all — the model gets the honest "could not be shown" note instead. Cosmetic
trimming (collapsing runs of whitespace, stripping trailing dots) happens only at render
time and changes how the name looks, never which bytes are written.
Tool-loop termination note on the buffered egress
When the gateway's own tool loop stops a turn early — the per-turn tool-call limit, a detected tool-call loop, an exhausted time budget (including the "answer is based on partial results" label when the gateway synthesized a partial answer over the results gathered so far), or a sub-agent deadline — it appends a short human-readable note to the answer. On a streaming turn that note is relayed live as its own content delta.
On the buffered compat egress (a "stream": false request, or a PII-force-buffered
"stream": true turn re-emitted as SSE) the note is included in the delivered response
body's message content — choices[0].message.content on a direct non-streaming response,
and the re-emitted content on the PII path — so a buffered client sees the same stop/partial
label a streaming client does, on the first response, matching the persisted turn (a reload
shows the same text). The note is appended exactly once and never onto a tool_calls body,
an error envelope, or a structured-output (response_schema) answer. A native
/anthropic/v1/messages direct non-streaming response gets the note folded into its last
text content block as well; only the native PII-force-buffered turn is excluded — it is
re-emitted through the compat-shaped tail, an existing limitation outside this scope.
Malformed upstream streams (fail-closed)
The upstream provider's SSE bytes are treated as untrusted. A single Server-Sent
Events line (the text between two newlines) has an internal size ceiling of 1 MiB.
Legitimate event lines are far smaller — token deltas of at most a few kilobytes, and
providers split even a large tool-call payload across many newline-terminated delta
events — so this ceiling never truncates a well-formed stream regardless of the total
response size (only the size of one un-terminated line is bounded). A broken or
hostile provider that streams bytes without a terminating newline past the ceiling
is cut off rather than allowed to grow gateway memory without limit: the affected leg
fails closed. On the compat (unified) path the stream terminates with a final chunk (emitted
when no terminal chunk had yet been sent) and then ends without a data: [DONE]
sentinel — the sentinel is suppressed on an errored stream. A plain client receives the
failure the way the OpenAI schema carries one: a data: {"error":{"message":…,"type":…,"code":…,"param":null}} frame before that final chunk, whose own
finish_reason is the in-enum stop (error is not an OpenAI enum member, so it is never
written to this field). The code is stream_truncated when the upstream closed early and
provider_error when the connection itself failed; when the provider announced the failure
with its own error event, the code is the gateway's classification of it (a rate-limit or
overload class) and type is the provider's own error type, so an SDK's retry policy can
read it. This frame is delivered to every client, first-party included — once the
terminal has to stay inside the enum, the error object is the only thing distinguishing a
truncated turn from a short one, so no audience may be left without it. A client that also
opted into the extensions still gets its aig_status:"provider_error" banner event.
Such a leg is not re-dialled on the same provider. Both the per-line ceiling above and
the per-leg total stream ceiling (64 MiB) are properties of the request/response pair,
not of the moment: an identical retry downloads the same oversized response and fails
identically, while the provider generates — and bills for — the output again. The gateway
therefore skips the remaining same-provider attempts and moves straight on to the next
provider in the routing chain, which may well be a model that answers. Only if the whole
chain is exhausted does the request end in all_providers_failed, and the recorded cause
names the ceiling that was hit rather than a bare stream error. Transient mid-stream
failures (a dropped connection, a read timeout) are unaffected and keep their full retry
budget. Models that return generated images inline in the stream are the common way to
hit the per-line ceiling — their base64 payload arrives as one un-terminated line — and such
a model cannot be delivered on this path at all; pick a text model, or one whose images are
returned by reference.
One case is deliberately not treated as a truncation: an Anthropic-style two-phase
terminal where the model's message_delta already carried a natural stop_reason
(end_turn/stop_sequence/pause_turn) and only the trailing message_stop was lost to a
socket close. That answer is complete, so the gateway emits neither the
stream_truncated error frame nor the visible note below — only the finish chunk (still the
in-enum stop). A genuine cut-off, where no terminal stop_reason ever arrived, is
unaffected and still gets both. (A max_tokens terminal is a real length cap, not a natural
end, so it keeps its length truncation signal.) On the native passthrough path the raw stream simply ends without a terminator
(which conforming clients treat as an error). In both cases the upstream connection is
closed rather than returned to the pool, and any pending tool call the truncated leg had
begun is discarded (never executed).
Visible truncation note (persisted), so a cut-off is never silent to a human. Because the
wire terminal reads the in-enum stop, a partial answer would otherwise render — live and on
reload — as a clean, complete one. On the compat path the gateway therefore appends a one-line
"The response was interrupted before it completed. Please try again." note to the visible
answer — as its own content delta after whatever partial text had already streamed (one blank
line between the partial and the note; when the partial already ends in a blank line the note
follows it directly, and a dangling whitespace-only last line is closed first so the note is
never indented into a code block), or as the whole bubble when nothing had — and records it into
the persisted turn, so a reload shows
the truncated turn the same way it rendered live. The note is gateway-authored English prose
(like the tool-loop termination note above); it is not localized. The chat UI additionally raises
a non-fatal response may be incomplete banner off the error frame's code. A turn that instead
ends with an actionable provider-error banner (aig_status:"provider_error") or a
context-overflow recovery banner gets neither the generic note nor the banner — those surfaces
already name the failure. This note follows the same rules as the tool-loop termination note
above (never onto a tool_calls body, an error envelope, or a structured-output answer). It is
the same on the buffered / PII-force-buffered egress: a turn the gateway buffers internally
(a PII-masking or guardrail configuration, or a stream: false tool-loop turn) whose upstream
dies after partial bytes delivers the salvaged partial with the note and the same failure
surface for its egress — the terminator error frame on the re-emitted SSE, X-AIG-Error on a
JSON body (see the Mid-stream provider failures on a buffered turn callout under
Anthropic-native passthrough) — and a conversation-owned turn
persists exactly those bytes; it used to persist the bare partial as a clean-looking complete
answer. The one verdict is shared: a leg that reached a natural end and only lost its
message_stop is complete on both egresses (no note, no frame), and a provider_error event
gets its banner and no note on both. A leg that will continue
in the tool loop gets no note — and only a leg that finished cleanly (its finish event
arrived — the finish_reason chunk on the OpenAI wire, message_stop on Anthropic — with no
read error and no provider error) with pending tool calls continues. A leg that did not finish
cleanly is dead: nothing that would make the turn continue is synthesized from it (no inline
<tool_call>, no unwrapped file tool, no file-claim steer), structurally delivered tool calls are
discarded (Anthropic tool_use blocks included — even after message_delta, if message_stop
never came), the gateway logs and traces dead_leg_no_continue, the note is appended, and the
error frame is terminal — Regenerate is the remedy. (A leg that reached a natural end and only
lost message_stop on an otherwise clean close is complete for the wire and the row — no note,
no error frame, its finish chunk emits — but equally never continues: nothing may follow that
terminal. If such a leg carries a tool call its death cost it — one the gateway would have
run had the finish event arrived: a recovered <tool_call> marker, an unwrapped <write_file> /
<read_file>, a write_file sentinel, or a structurally delivered call (an Anthropic tool_use
block, an OpenAI tool_calls delta) — it is not treated as complete: it gets the note and the
error frame, because its prose is a preamble to work that will never happen. Text the gateway
discards on a clean finish too (an anchored bare-JSON or pseudo-call naming write_file /
read_file, which the inline channel never dispatches) does not change that verdict — the answer
stays complete — but the discard is always logged, so a dropped write is never silent. A pending bare-re-download card is not such a marker —
on a streaming turn, and on a PII-force-buffered turn (whose captured card the buffered re-emit
replays), that card is re-served on any leg that neither hit a read error nor carries the
interrupted note (a leg that ended in a provider error does get it), so a re-download does not
go silently missing; a leg that carries the note re-serves no card on either egress — a card
under an "interrupted" note would contradict it, and Regenerate re-serves the file — and on a
non-streaming (stream:false) turn the card has no wire to ride, as before.) So a partial on a tool-capable turn (the common
first-party conversation turn, where file tools are always offered) is a cut-off like any other
and carries the note, live and on reload. The same holds for an API turn that supplies its own
tools: if the provider dies after tool_calls deltas but before the finish chunk, those calls
are discarded and the note is appended to the visible answer — the turn is a cut-off, not a tool
turn, and no tool_calls body is emitted for it.
The client's own connection dies mid-stream (a laptop lid, Wi-Fi gone at 80 % of a long
answer). The gateway sees the client abort, stops reading the provider at once and closes
the provider connection (the model stops generating for a turn nobody will receive), then
persists the partial answer it had already streamed once, synchronously, under the turn's id
— nothing is lost server-side and nothing is doubled; the request is accounted as a client abort
(request_log.aborted = 1, the leg status 499 / error_class cancelled / partial = 1). The chat app treats the read failure the
browser raises for a dropped connection as exactly that — one predicate accepts each engine's
wording for both moments, before a response and mid-body: Chrome Failed to fetch /
network error, Firefox NetworkError when attempting to fetch resource. /
Error in body stream, Safari Load failed / The network connection was lost. — and keeps the
partial the user was watching on screen as the assistant turn, raises the same non-fatal
response may be incomplete banner, re-enables the composer, and adopts the persisted row
(which may carry a token or two more than the client received) through the same bounded
reconcile a Stop uses — so live and reload show one partial, once. If the network is still down
when the reconcile runs, the partial simply stays as rendered — no blank, no "no response
received" row — and the next load shows the persisted partial once. (The same failure while
nothing had streamed yet is a plain error banner — there is no partial to keep.)
The provider's HTTP body framing. Before any of the above, the gateway frames the
provider's response body itself (its own coroutine-free reader over the connection — the
receive belongs to the coroutine that is cancelled on a client abort, so an abort really closes
the provider connection). Accepted: Transfer-Encoding: chunked with a hexadecimal
chunk-size line (as tonumber(line, 16) reads it — the SSE case), a Content-Length body
read to exactly that length, or — with neither header — a read-until-close body. Rejected, and
surfaced as a read error that ends the leg (never a fabricated length, never a silent end): a
chunk-size line that is not hex (a chunk extension such as 1a;name=value, a non-hex token) or
is negative; a Content-Length that is not a single non-negative integer (4.5, a duplicated
header) is ignored and the body is read until close, so a truncated read can never leave stray
bytes on a pooled connection. Sizes are honoured in bounded pieces (at most 64 KiB per read
without a caller limit), so a size the socket cannot honour ends the leg as a read error. HEAD /
1xx / 204 / 304 responses have no body and are never waited on.
Structured / malformed content values. A provider's delta.content (streaming)
and message.content (buffered) are equally untrusted. Accepted shapes are a plain
string, null, or an array of parts; the gateway flattens an array into the
visible answer text (the text parts concatenated; the same flattener, with a separator between
the parts, also joins a user message's text parts where a route needs a single string — the
reply-language detection and the Mistral web-search fold). Array normalization is two-tier: recognized typed non-text
parts (a part whose type is a string other than "text" — e.g. a thinking or
citation part) are omitted from the visible text silently; garbage shapes
(non-object/non-string elements; text-equivalent blocks — type absent, null, or
"text" — with a non-string or absent text; blocks with a non-string non-null
type; content values that are neither string, null, nor a table with a leading
element sequence) are dropped fail-closed with a gateway WARN log (deduplicated per
stream leg); an empty array yields empty text silently. The same fail-closed guards
apply to the per-part text fields of the native Anthropic, Gemini, and Cohere
parsers, and the compat streaming dispatcher additionally drops any non-string delta a
provider parser might emit (one WARN per leg) instead of letting it abort the live
stream. Before these guards, a structured content part could kill a streaming turn
mid-leg with no finish event — the user saw a silently dead chat.
Source citations (aig_sources)
When an answer draws on an uploaded document, the gateway emits a structured, deterministic
source list as a dedicated SSE chunk on the final leg of the turn — built from the turn's
actual document-retrieval calls (reading a file by name or the search_knowledge hybrid
search over the project corpus), not from model text:
data: {"aig_sources":[
{"file":"budget-report.pdf","loc":"p.1","kind":"page"},
{"file":"budget-report.pdf","loc":"p.2","section":"Section 3.1","kind":"page"}
]}
| Field | Type | Meaning |
|---|---|---|
file |
string | The document the answer drew on (required, non-empty). |
loc |
string | The locator that was retrieved, e.g. p.2 (required, non-empty). |
section |
string | Section/heading within the document, when known (optional). |
kind |
string | Locator kind — page, slide, sheet, para, line (optional). |
The list is deduplicated by file + loc + section, ordered by first retrieval, and reflects
only the portions actually delivered to the model (a document truncated to fit the context window
does not contribute citations for the pages that were cut). It is bounded — at most 100 entries,
each field at most 300 characters — so a many-page document cannot produce an unbounded list. A turn
that retrieved nothing emits no aig_sources chunk. Consumers must treat every field as untrusted
input and ignore malformed entries. Like the other aig_* gateway-extension events, aig_sources
is emitted only in streaming mode; a non-streaming ("stream": false) response returns a plain
chat.completion JSON object with no citation channel.
Only strict-verifiable retrieval is cited. Segments extracted by OCR (confidence='ocr') are
served to the model as document text but are not cited — an OCR page may be mis-located, so the
gateway never emits a deterministic page/section claim it cannot stand behind.
Citations persist. The same list is stored on the assistant message, so the citations rehydrate when a conversation is reloaded (not just for the live stream). The persisted list is re-validated through the same bounds above before it is returned, and a turn that produced no answer stores no citations (a citation attributes an answer).
fetch_url / agentic_fetch pages range (untrusted model input). For a long PDF
truncated to its first pages, the model may re-fetch the same URL with an optional pages argument
to read a specific window. Accepted shape: a string — a 1-based range "START-END" (e.g.
"110-125") or a single page "117"; surrounding whitespace is tolerated. The window is
[START, END] and its span is clamped to the 200-page extraction cap. Rejected / ignored (→
read from the first page, never an error): a non-string value (number/table/boolean), an empty or
non-numeric string, a reversed range ("10-2"), or a zero/negative page ("0-3"). ABSENT (omitted)
and every malformed shape share this identical first-page fallback — the parse is fail-closed and
never widens the read. pages applies to PDFs only (DOCX/XLSX/web pages ignore it). A range
that begins beyond the document, or hits only text-less pages, returns a served notice ("Requested
pages start beyond … the N-page document"), not an extraction failure.
💡 Note: A guardrail block in streaming mode returns
HTTP 200with a synthetic SSE error chunk (the refusal text as a normal assistantdelta.contentchunk) followed bydata: [DONE], rather than a non-200 status. This wire shape is unchanged for external API clients — integrate against the assistant message as before. See Error codes — guardrail_blocked.First-party turn-owned streams (the web app) additionally receive a
{"aig_status": "guardrail_blocked", "error_class": "<class>"}gateway-extension event so the app can render a distinct policy-block notice instead of a normal reply. Like everyaig_*event it is emitted only in streaming mode and only on a turn-owned stream; a plain API client never sees it (see What a plain client receives).error_classis one ofguardrail_blocked(content policy),guardrail_pii_mandatory,guardrail_pii_media,guardrail_role_model, orguardrail_unavailable(treat the set as open) — a stable, non-sensitive slug (it never carries the specific policy sub-category or any request content). On a turn-owned stream the refusaldelta.contentis a generic, sub-category-free message (the specific category stays in the gateway logs); external clients keep the detailed prose.🔒 File attachments on a PII-protected gateway: because PII guardrails scan content as text, a gateway with an active PII protector rejects a request that embeds inline base64 file media (an
image_url/video_urldata:URL, aninput_audio/fileblock, or an Anthropic-nativeimage/documentwith a base64source) with the block reasonpii_media_unmaskable— such binary would otherwise reach the provider unmasked. Attach files through the Chat view, which extracts each PDF/image to text server-side first so the text is masked before it leaves the gateway. Text-extracted blocks and Files-APIfile_idreferences are accepted. A request served wholly by a first-party local (Myra/EU) model is exempt from this block — there is no third-party egress to fail closed for. See Data protection — file attachments on a PII-protected gateway.♻️ Model-unavailable recovery (extensions-enabled streams): when the pinned model is no longer routable (deprecated/renamed/deleted upstream), the gateway ends a STREAMING turn from a client that opted into the side channel with
HTTP 200and a singleaig_statusrecovery event instead of a cryptic provider error (a plain OpenAI-compatible client, and ANY non-streaming request, gets the typedmodel_not_foundHTTP error instead — see the closing sentence of this note):{"aig_status": "model_unavailable", "old_model", "old_display", "fresh_model", "fresh_display", "fresh_gateway_id"}when a replacement default exists, else{"aig_status": "no_default_available", "old_model", "old_display"}. The suggested(fresh_gateway_id, fresh_model)is policy-legal by construction: it is validated against the conversation's permission tier and PII posture with the SAME predicate the conversation-routePATCHguard enforces (see Route policy validation), so the server never proposes a route the app would then be refused. If the only available default would be illegal for the conversation (a cloud model for alocal_onlyproject, or a non-scrubbing gateway under a PII mandate) — or the tier can't be proven during a transient fault — the event fails closed tono_default_availablerather than offering an illegal target. Like everyaig_*event it is emitted only to a client that opted into the side channel (x-aig-turn-id/x-aig-extensions: 1) and asked for a stream; every other caller — a plain OpenAI-compatible client, or any"stream": falserequest — receives the typed HTTP errormodel_not_found(400) instead of an SSE body it did not ask for. That error names the model id — e.g.the model '<id>' is not available on this gateway.(the exact wording varies by path) — so an integrator can see which id was wrong without reading the gateway's logs.What counts as "no longer routable": any
404from the upstream on the chat-completions call, whatever its body says. It used to depend on the provider's prose — the same fault classified three different ways depending on whether the upstream happened to name the model in its error text, so a{"message":"not found"}body was reported to the user with the generic "The request could not be completed.", the same sentence the product uses for a timeout. A 404 is now always a model failure: it is never retried against the same id (definitive, not transient), and it reaches this recovery. A 404 whose body describes a MORE specific fault (a capability gap, for instance) still classifies as that fault — the status is a floor, not an override.
Per-request header overrides
The provider pass-through supports an optional control header:
| Header | Effect |
|---|---|
x-aig-byok-alias |
Selects which stored BYOK key alias to use for the resolved provider on this request (see Provider key management). Accepted shape: a label matching an alias already stored for the provider on the gateway. Absent or empty → the default alias. An alias that does not exist is rejected with a configuration error — there is no silent fallback to default, and no fallback to a tenant-level key. Applies only to the primary provider; a fallback provider always uses its default-alias key. |
x-aig-extensions |
Opt into the gateway's aig_* SSE side channel (streaming only). Accepted: 1, true, yes, on. Anything else — including 0, an empty value, an unknown word, a non-string, or a repeated header whose first value is not accepted — is rejected and treated as absent, so the stream stays pure OpenAI. Implied by a valid x-aig-turn-id. This is a verbosity switch, never an authorization signal: it unlocks nothing about any turn but the caller's own. See What a plain client receives. |
x-aig-image-gen |
Client opt-in for the generate_image tool (see Image generation). The tool is offered to the model only when the gateway has image_generation.enabled: true (see config reference) and this header carries a plain value other than 0 / false — the web chat always sends x-aig-image-gen: 1. Accepted: any single plain-string value except 0 and false — the match is exact and case-sensitive, so an empty value, FALSE, no or off all count as opt-IN (unlike x-aig-extensions, which is an allowlist). Absent, 0, false, or a repeated header (which arrives as a list, never as a plain string) → the tool is not offered; a non-string can never arm it. On a gateway with image generation disabled the header has no effect. Saved/scheduled agents never carry it. |
x-aig-provider-* |
The x-aig-provider- prefix is stripped and the remaining header is forwarded to the upstream provider only if the stripped name is on a short allow-list — anthropic-beta, openai-organization, openai-project. Any other stripped name (a credential such as authorization, a request-framing header such as content-length / transfer-encoding / connection, or any unrecognised name) is dropped and logged at warn; it never reaches the upstream request. At most 16 overrides / 8 KB total are forwarded; the excess is dropped. |
x-aig-knowledge-ids |
Per-request knowledge-source ALLOWLIST. A comma-separated list of project/conversation knowledge-document ids the assistant may consider for THIS request. Tri-state by presence: absent → all knowledge (the default); present → exactly the listed ids; present with no valid id (e.g. a single - sentinel) → no knowledge. At most 200 ids are read; a longer list is truncated (truncating an allowlist only narrows what the assistant may use). The selection is intersected with — never widens — the caller's already-authorized project/conversation scope, and is enforced for auto-injected knowledge, the read_file tool, and the search_knowledge hybrid-search tool (the model can neither read nor search a deselected document). It does not restrict loading a project file into the code interpreter's sandbox, which is a separate, deliberate by-name action gated only by the caller's access to the file (see Code interpreter). search_knowledge itself accepts a single query string (natural language or exact terms such as a section/form number); a blank query is rejected, and the search is always confined to the caller's own project/conversation and tenant — a document from another project or tenant is never searched or returned. A file the assistant creates during this request (via write_file or the code interpreter) is always readable by that same request's tools, whatever the selection says — the selection describes the pre-existing knowledge, not the request's own output. The selection is part of the cache key, and the semantic cache is skipped when it is present. Saved/scheduled agents do not honour caller headers, so they always use all knowledge. |
x-aig-knowledge-exclude-ids |
Per-request knowledge-source DENYLIST — the complement of the header above, and what the chat UI sends when you uncheck a file. A comma-separated list of knowledge-document ids to EXCLUDE from this request; everything else in the caller's authorized scope stays in, including documents added after the client last looked. Absent → nothing excluded. Unlike the allowlist this header is honoured in full or not at all: if any token is not a document id (internal whitespace included), if the header carries no usable id at all (an empty value, or only the - sentinel), or if more than 200 ids are listed, the request is rejected with 400 invalid_request — silently dropping an id would feed the model a document the caller asked to exclude. There is no "exclude everything" value; use x-aig-knowledge-ids: - for "no knowledge". Both headers may be sent together, in which case the allowlist is applied first and the denylist removes from it; neither can ever widen the caller's authorized scope. Same enforcement points as the allowlist, and the same carve-out for a file created during the request. A request carrying this header is not cached (see Caching). Saved/scheduled agents do not honour caller headers. |
x-aig-leg |
Marks a request the chat interface made on its own behalf, rather than a message the user sent — currently followup_suggestions (the follow-up chips under an answer) or title_generation (the automatic conversation title). Both are ordinary completions and are guardrailed like any other, but they follow the user's turn and would otherwise be indistinguishable from it in the logs: an operator inspecting "the newest request" was inspecting one of these. Log-only — it never affects routing, guardrails, billing or any policy decision, and a value outside the two above is dropped (the request is logged exactly as it would be without the header). The label appears as the leg kind on request_log_legs.kind and on the log row, where the Logs list badges it. Saved/scheduled agents do not honour caller headers. |
X-Request-Id |
Your own correlation id for this request. The gateway always mints its own request id — returned in the response X-Request-Id header, and the id of the request's log entry; a value you send is never used as that id, so it can neither collide with nor suppress the logging of any request. Accepted shape: a single header whose value matches ^[A-Za-z0-9._\-:+]{1,64} anchored at the absolute end of the value — a trailing newline is rejected (the server anchors with \z, not a $ that would also match before a final \n) — it is recorded on the log entry as client_request_id (list and single-entry responses, the CSV export, the tenant export and JSON SIEM events). Rejected (dropped — not recorded, not echoed, the request is served normally and still logged): an empty value, more than 64 characters, any other character, or the header sent more than once. To the caller "sent nothing" and "sent something unusable" look the same — deliberately so, since the header unlocks nothing; the drop is logged server-side at info level (visible only with the error log at info; the shipped configs run at notice). Like the other correlation headers it is read from the first 100 request headers — placed beyond them it reads as absent. Saved/scheduled agents and delegated agent runs never carry it. |
Sending the same header twice
HTTP allows a header to appear more than once, and some SDKs do send anthropic-beta on two
lines. The gateway's rule depends on what kind of value the header carries, and it
is never "whichever copy arrives first wins by accident":
| Header kind | Repeated | Why |
|---|---|---|
List-valued — anthropic-beta, x-aig-provider-*, user-agent |
The copies are combined into one comma-separated value (a, b), exactly as RFC 7230 §3.2.2 permits and as the front-end proxy already does for the same header. |
The header's own grammar is a list, so nothing is lost. Dropping the second copy would silently discard a beta flag and surface as an opaque provider 400. |
Boolean opt-in — x-aig-web-search, x-aig-image-gen |
Treated as not set. An explicit : 0 sent twice keeps the capability off. |
The value is unusable, and an unusable value must never be more permissive than an absent one. In particular, an intermediary that appends the header cannot defeat your opt-out. (x-aig-extensions, a first-party verbosity switch, instead takes the first value on a duplicate.) |
Scoping / identity — x-project-id, x-aig-conv-id, x-aig-skill, x-aig-thinking-budget |
Treated as not asserted, and the request is not cached. | These select a permission and residency scope. A duplicated value names no single scope, so the gateway declines to guess: the request runs without that scope rather than with a guessed one, and its response is never written to (or served from) the shared cache. |
Correlation only — X-Request-Id, x-aig-turn-id, x-aig-parent-msg-id, x-aig-regen-of-turn |
Treated as not sent and logged server-side (X-Request-Id at info, the x-aig-* ids at warn); the request is served and logged normally, with the gateway's own request id. |
These name nothing that decides routing, permissions or caching — a duplicated value is simply unusable as a correlation key, and an unusable value must never be recorded as if it were the one you meant. |
A duplicated header never returns 500. Before this rule a repeated anthropic-beta produced
an internal error on the inference path that looked like a provider fault.
is forwarded to the provider as:💡 Note: To pin a specific provider, address it in the URL path (
/v1/{tenant}/{gateway}/{provider}/chat/completions) instead of thecompatpseudo-provider. The compat endpoint resolves the provider from themodelfield, not from a request header.
See Providers overview — Compat model resolution and Provider header pass-through for the full behaviour.
💡 Note: Additional
x-aig-*control headers are accepted for specific features (web search, projects, MCP tools, conversation threading). They are documented alongside the feature they configure.🔒 Scope of caller control headers. On this OpenAI-compatible endpoint, the caller-supplied
x-aig-*control headers are honored as sent — the gateway does not silently drop them here (unlike the saved-agent invoke path, which is owner-scoped and clears caller control headers). This is safe because the token is bound to a single user, and every header resolves against that identity:x-aig-conv-id,x-aig-policy-conv-idandx-project-idare workspace-scoped: they resolve only within the caller's own workspace (an id from another workspace returns nothing), and what they resolve is a policy tier, never message content — no conversation text is loaded or returned through them. MCP connector references (tools:[{type:"mcp",connector_id}]) can invoke only connectors the caller owns or is granted (using the caller's own credentials), andx-aig-web-searchmerely toggles a capability the gateway must also have enabled and provisioned. No control header can disable PII protection. The client is never the authorization boundary — these headers convey the caller's own intent within the caller's own scope.🔒 How the egress tier is resolved (residency / PII). A request's data-residency tier (
local_only— no external tools/models;pii_mandatory— masking required) governs whether the turn may egress externally, and it is resolved only from the authoritative source, never from a value the client can vary: - A request that carriesx-aig-conv-id(a conversation) resolves its tier solely from that conversation's committed project binding, server-side. A value the gateway cannot use — one that does not match the id shape, or the header sent more than once — is not treated as "no conversation was asserted": the turn fails closed, because otherwise a control could be defeated by sending the header twice. Thex-project-idheader is not consulted for a conversation-bearing request — so a stale or mismatchedx-project-idcan neither loosen a restricted conversation (no residency leak) nor spuriously restrict a plain chat. If the binding cannot be confirmed for the request, the turn fails closed (external tools/models blocked) and returns the retryableproject_tier_unresolved(503), never a false permanent "only local models" — retry resolves it. - An id that names no conversation at all is not an unconfirmed binding, and is treated exactly as if the header had been omitted: the tier comes fromx-project-idas below. Conversations are created only through the conversations API, never by an inference turn, so an id with no conversation behind it is a caller's own grouping key (it also keys the PII salt and the prompt cache). This grants nothing — omitting the header reaches the same place — and a restricted conversation is unaffected, because it has a real binding and takes the rule above. - A conversation whose project no longer exists (deleted) keeps failing closed, but the answer is permanent, not retryable:conversation_project_unresolved(403). The retryable 503 above is reserved for a binding that genuinely might resolve on a retry — promising "try again" for a project that is gone is a promise nothing can keep. - A request with no conversation (a raw/v1caller acting within a project) resolves its tier fromx-project-id, server-side and tenant-scoped — the header names a project; the gateway looks up that project's real tier. A cross-tenant or unknownx-project-idyields no tier. - A request that carriesx-aig-policy-conv-idresolves its tier from that conversation's committed project binding, exactly asx-aig-conv-idwould — but it does nothing else. It exists for the chat surface's own post-turn helper calls (auto-title, follow-up suggestions), which carry real conversation text and so must obey the same tier, but which cannot sendx-aig-conv-idbecause that header also switches on server-side system-prompt composition and would replace the helper's instruction with the conversation's persona, memories and project knowledge. The header names a conversation; the gateway looks up that conversation's real binding, tenant-scoped, and fails closed when it cannot. Accepted shape: the same id shape asx-request-id(^[A-Za-z0-9._\-:+]{1,64}, end-anchored — a trailing newline is rejected). It is ignored wheneverx-aig-conv-idresolved a conversation for this request, so it can never override a real conversation's binding. Anything it cannot use — a malformed value, an empty one, or the header sent more than once — is refused (the turn fails closed), never silently skipped: a header the caller deliberately sent must not be treated as one they never sent. - A conversation that provably has no project is a plain chat and is unrestricted; the header is ignored there too.If the database is briefly unreachable. Tier resolution always reads the database on the happy path, so a tier change takes effect on the very next request — the gateway keeps no read-through cache. It does keep a last-known-good value per gateway node, consulted only when that read fails, so a momentary database blip degrades to the last confirmed answer instead of blocking every project conversation. If nothing is cached, the request fails closed.
That fallback is an availability measure, not a guarantee, and it is worth stating what it does not promise. Moving a conversation between projects, or changing a project's access tier, clears the value on the node that handled the change — but each node keeps its own, so for up to 5 minutes another node can still answer a failed database read from what it last confirmed. Where the tier was loosened in the meantime that merely over-restricts. Where it was tightened — a conversation filed into a restricted project, or a project moved to a stricter tier — that node can treat the conversation as it was before, during a database outage, for up to that window.
Use Local only where that window is unacceptable: it constrains model routing itself, so a turn cannot reach a non-local provider regardless of what any cache answered.
Every such decision is recorded. Whenever a tier is served from the last-known-good value because the authoritative read failed, the request log carries
meta.residency_tier_from_cacheset to the value that was served — a tier, orno_projectfor a conversation last confirmed to have none. It is visible onGET /admin/v1/logs/{id}and accompanied by a[residency_tier_from_cache]warning in the gateway log. A healthy resolution is not marked, and neither is a request that failed closed: the marker means exactly "this answer may predate a change we could not read", so a compliance review can enumerate what was let through during an outage instead of inferring it.
MCP connectors from the API
Registered MCP connectors can be used from the inference API by reference — the gateway resolves the connector's tool list server-side, offers the tools to the model, and executes every tool call inside its own tool loop with the caller's stored connector credential (which never leaves the gateway):
{
"model": "…",
"messages": [{ "role": "user", "content": "What does ticket ABC-123 say?" }],
"tools": [{ "type": "mcp", "connector_id": "<connector id>" }]
}
Accepted shape per entry: exactly { "type": "mcp", "connector_id": "<id>" },
plus an optional server_label (accepted for OpenAI-shape compatibility and
ignored — gateway tool names are not label-prefixed). connector_id must
be a non-empty string of at most 64 characters. Duplicate ids are de-duplicated;
at most 8 distinct connector references per request.
Authorization. The API token's user is the credential identity: a
user-scoped connector uses that user's connected credential (connect it in the
UI first), and a private connector is visible only to its owner. A tenant-level
token (no associated user) can use shared connectors; user-scoped ones return
connector_credential_required and private ones connector_not_found. The
connector's server-side tool_policy filters the offered tools and is
re-enforced on every call; connector tool names that collide with built-in
gateway tools or the reserved agent__ prefix are silently dropped from the
offer.
Per-tool shape narrowing. Every tool definition that reaches the model —
whether discovered from a connector's tools/list or carried in the
gateway-internal x-aig-mcp-tools body field — is treated as untrusted and
narrowed fail-closed before it is offered: an entry must be a JSON object with a
function object and a usable name (non-empty string, valid UTF-8, at most
200 bytes — resolved from the entry's name or function.name), or the entry
is dropped; duplicate names keep the first entry only; at most 5000 entries per
request are considered. Within each kept tool, non-conforming fields are
repaired rather than failing the request — see
MCP connectors — tool shape narrowing
for the exact field rules.
Rejected inputs (all 400 invalid_request unless noted):
| Input | Behaviour |
|---|---|
unknown / foreign / other member's private connector_id |
404 connector_not_found (fail closed, request aborted) |
| connector requires a credential the token's user has not connected (or it expired) | 424 connector_credential_required |
| connector server unreachable / tool discovery failed / discovery deadline exceeded | 502 connector_upstream (fail-closed; but see x-aig-mcp-best-effort below — the chat surface drops the connector and proceeds instead) |
mixing {type:"mcp"} references with client function tools in one tools array |
400 — one loop owner per request |
server_url, allowed_tools, headers, authorization, require_approval, or any other key on an mcp entry |
400 naming the key — register a connector; restrict tools via its server-side tool_policy |
| more than 8 distinct references | 400 |
combining references with the gateway-internal x-aig-mcp-tools body field |
400 — one MCP source per request. (The legacy x-mcp-tools header was retired and is now ignored, never an error.) |
combining references with legacy functions / function_call |
400 |
a forcing tool_choice ("required", a named function/tool, {"type":"any"}, allowed_tools mode required, …) |
400 — the gateway drives the tool loop. Non-forcing forms pass: absent, null, "auto", "none", {"type":"auto"}, {"type":"none"} (shape-agnostic; the provider still enforces its own wire shape) |
| references on a non-chat endpoint (embeddings, …) | 400 |
references on count_tokens |
400 — connector tools are never included in a token count (a refs request cannot pre-count) |
| references with a model whose provider has no gateway tool loop (e.g. Gemini, Vertex, Bedrock) | 400 |
references on native /v1/messages with stream: true |
400 — connector references on /v1/messages require stream:false in this version |
references on native /v1/messages with a provider that relays the Anthropic response through an OpenAI-compatible reader (the self-hosted myra fleet, groq, together, …) |
400 — that provider dials /v1/messages but cannot read the Anthropic dialect, so no gateway tool loop can run there; without references the request is a transparent passthrough. A provider that dials its own endpoint (a routed cohere model) keeps the loop. |
| references in a local-only project / egress-blocked context | 403 forbidden — "MCP connectors are not available for this request (egress blocked)" |
Operational notes.
- Tool discovery (
tools/list) runs server-side with an overall deadline of ~20 seconds across all referenced connectors (plus at most one in-flight operation); results are cached for up to 5 minutes per connector + user, so a connector-side tool change can take up to 5 minutes to appear. A failed discovery (server unreachable / deadline) is briefly negative-cached (~15s) so a persistently-down connector is not re-probed on every request. - A connector that (legitimately) exposes zero offerable tools does not fail the request — the turn proceeds with the gateway's built-in tools only.
tool_choice: "none"with references still performs the (billable) tool discovery for tools the model can then never call — omit the references instead.- MCP-carrying requests are never served from or written to the response cache.
- A member-created private connector may only reach public addresses; a
private-range/internal
server_urlsurfaces asconnector_upstream.
Connector provenance notice in the model's context
When at least one connector-backed MCP tool is injected, the gateway merges
one system-level notice (sentinel [aig:mcp-connectors]) into the outgoing
request that enumerates each attached connector's display name and its tool
names, and instructs the model that these connectors are available — call
their tools instead of denying access. (Without provenance, some models deny
having an attached connector even though its tools are offered.)
Behaviour and guarantees:
- Merged at most once per request — the sentinel makes re-injection across retries and tool-loop continuation legs a no-op.
- Absent when no MCP connector tools are attached (built-in-tools-only turns, and agent-as-tool-only turns) — the notice never promises tools the model does not have.
- Untrusted input is sanitized before embedding (fail-closed): connector
display names and tool names are stripped of ASCII control characters and
structure punctuation (quotes, brackets, braces, backticks,
<>,|, backslash), whitespace-collapsed, and capped at 64 bytes (UTF-8-safe — accent and umlaut characters survive). A display name that sanitizes to nothing falls back to the connector id; if that is unprintable too, the connector is omitted from the enumeration entirely. Raw values are never embedded. - Flood-bounded: at most 20 tool names are listed per connector (then "and N more tools") and at most 16 connectors are enumerated (then "and N more connectors").
- The internal provenance fields (
connector_id,connector_name) never reach the provider'stoolsarray.
x-aig-mcp-best-effort request header. Accepted value "1" opts
the request into best-effort connector-reference resolution: a connector that
fails discovery (connector_upstream / discovery deadline) is dropped and the
turn proceeds with the remaining tools, and the drop is surfaced to the SPA via the
tools_skipped progress notice instead of failing the turn. The same
opt-in also governs the egress-policy drop: in a Tier-1 local_only project the
referenced connectors are dropped (token connectors_policy) and the turn proceeds,
where a fail-closed caller gets 403 forbidden. A run marked no-egress by its profile
(a webhook-triggered agent run) stays fail-closed regardless of the header. Any other value, a
repeated header, or its absence keeps the default fail-closed behaviour (a
discovery failure aborts the request with the typed connector_* error above). The
web chat sends x-aig-mcp-best-effort: 1 on every send — an auto-attached
connector whose server is momentarily down must not break the chat; the raw API
omits it so an explicitly referenced connector still fails loudly. Under
best-effort, any discovery failure for a connector (including
connector_not_found / connector_credential_required) drops that connector and
proceeds; per-connector authorization is still enforced, so a dropped connector's
tools are simply never offered.
The turn_committed frame (turn-owned streams)
A turn-owned stream (one sent with x-aig-turn-id) ends with one aig_status: "turn_committed"
frame — the gateway's word that the answer is durably persisted (or that there was nothing to
persist). Its fields:
| Field | Meaning |
|---|---|
turn_id |
the turn this frame closes |
message_id |
the committed assistant row's id, or null when nothing was persisted (an empty turn, or a ghost conversation — then ghost: true is also present) |
files[], images[] |
artifacts the turn produced and claimed onto the row |
usage |
input_tokens, output_tokens for the whole turn |
finish_reason |
the turn's final stop reason |
superseded_message_ids[] |
always present (an empty array when nothing was superseded): the message ids this commit replaced — the whole trailing run of the regenerated question: its prior answer(s) and the error rows among them (see x-aig-regen-of-turn below). The app removes exactly these rows in the same update that shows the new answer. |
data: {"aig_status":"turn_committed","turn_id":"…","message_id":"…","files":[],"images":[],"usage":{"input_tokens":12,"output_tokens":48},"finish_reason":"stop","superseded_message_ids":["<prior answer id>"]}
A failed turn (a provider fault, or a guardrail block on the buffered response path)
persists its failure row without a turn_committed frame; the app learns the outcome from the
error frame and the next conversation load. (A request-phase guardrail block on the streaming
path ends with a normal turn_committed — carrying the error/guardrail rows it superseded.)
How large an answer can be committed. Two ceilings, and anything that can be written can be
read back. The row's own bound: an assistant message's large text columns — content, any later
corrected_content or edit, and the sources / search-steps lists a web-search turn stores — may
total at most ~15.9 MB (the largest payload one database packet can carry back — the wire
protocol's own single-packet limit — less a reserve for the rest of the row); a write over it is
refused before anything is written, atomically with the row it would have grown. The statement's bound: one row is
written in a single statement whose packet may not exceed 16 MB. Text is escaped on the way
in, so the worst case — an answer made entirely of characters that need escaping — leaves room
for about 8 MB; ordinary prose grows by a few percent, so in practice an answer of ~12 MB still
fits (the same statement also carries the turn's sources and search steps). A multi-megabyte
answer — a long analysis, a big pasted document folded back into the reply — commits normally
and reloads byte-identical. Beyond either ceiling the commit FAILS LOUDLY rather than silently
dropping the answer: the turn takes the persist-failure path (aig_status: "persist_failed"),
and the client sees the answer it streamed but is told it was not saved. Nothing is truncated
behind your back, and no half-row is left in the conversation. On the read side the database
driver is capped at the wire's single-packet limit, so a row that somehow outgrew it (written
before this rule) fails to load with a clear error instead of loading as truncated phantom rows.
x-aig-regen-of-turn — regenerating a tool-backed answer: forced tool, and the prior answer superseded on commit
When the chat UI regenerates an assistant answer, it sends
x-aig-regen-of-turn: <original-turn-id> — the logical-turn id of the answer being
regenerated.
- Accepted shape: a string of at most 36 characters matching the request-id
pattern (the same rule as
x-aig-turn-id). Anything else — a non-string, an over-length value, or forbidden characters — is rejected (dropped, logged atWARN) and treated as absent. The value is used only as a parameterized query parameter, scoped to the caller's tenant (the leg lookup) and to the request's own conversation (the supersede), so a forged id cannot read or touch another conversation's data. - Effect 1 — tool forcing: if that original turn ran a model-driven tool loop
(its persisted request legs include a
tool_loopleg), the gateway forces the model to call a tool on the regenerate (tool_choice: "required") so it cannot decline and answer "no access". If the target provider/model does not support a forced tool choice, the gateway falls back to offering the tool (logged atWARN) rather than failing. When the header is absent, or the original turn used no tool, nothing is forced. (A forked copy of an answer carries a fresh turn id whose legs belong to the source turn — no forcing on it.) - Effect 2 — supersede on commit. The app no longer deletes the prior answer
before regenerating; it stays visible while the replacement streams. When the replacement
commits, the gateway — in the same database transaction — soft-deletes the prior answer
and any error rows trailing it, and lists them on the
turn_committedframe assuperseded_message_ids. Rules:- Who: the conversation's owner only — the same predicate the message DELETE
route's storage applies (a project member deletes nothing there, and so supersedes
nothing here — not even as the current driver of a live session; their regenerate is
appended beside the owner's answer). Decided once when the turn starts, never on the
commit path. A non-owner, an API-key turn (no user), an invisible conversation, or a
gateway-side fault ⇒ nothing is superseded, the new answer is simply appended (both
remain visible), and the refusal is logged at
WARN+ recorded as the trace stepregen_supersede_refusedwith areason(not_owner,no_user,no_turn_id,not_found,db_error:<err>; and from inside the commit:no_anchor,not_trailing,self_anchor— the header named the turn itself). Note that the turn itself is not refused on ownership grounds: the streaming path checks the live-session lease, not conversation ownership (a pre-existing gap, tracked separately) — so a member's regenerate on a shared conversation still lands its answer in that conversation; only the supersede is withheld. - What: the answer with that turn id is the anchor (resolved even if already
soft-deleted, so a stale tab regenerating twice still ends with ONE live answer —
last-commit-wins); the whole trailing run is superseded — every assistant row after
the question the anchor answers (the last user message at or before the anchor; with no
such message, the anchor itself and everything after it) — but only if no live user
message follows the anchor (a stale tab regenerating an older turn never deletes
another question's answer: refused, logged at
WARN, traced asnot_trailing). So an earlier answer stacked under the same question (an appended non-owner regenerate, a rolling-deploy gateway) goes with the swap too. Livestreamingclaim rows of a sibling turn and the new row itself are never touched; the rows soft-deleted are exactly the live subset of the ids on the frame (an already-superseded row is listed so a stale tab removes it too, but is never deleted twice). - A failed regenerate never removes a real answer. A provider failure or a guardrail block persists a failure row that supersedes only the error/guardrail rows of that run — so the prior answer stays, the failure lands beneath it (and repeated failures never stack), and a later successful regenerate replaces both.
- A stopped regenerate (the client aborts after tokens streamed) commits its partial and supersedes the prior answer like a full answer would.
- Idempotent and atomic: a retried commit of the same turn re-runs the rules with no effect on already-deleted rows; the supersede and the new row commit together or not at all.
- Rolling deploys: a gateway container still on the previous image supersedes nothing
(both rows stay visible until the fleet has rolled); its turn-owned answers carry a turn
id like any other, but rows it writes through the legacy paths (
POST /messages, a bridge pair, a fork copy) have none and cannot be superseded later. Every assistant row written by a current gateway carries a turn id on every path (migration 0283 backfilled the historical ones).
- Who: the conversation's owner only — the same predicate the message DELETE
route's storage applies (a project member deletes nothing there, and so supersedes
nothing here — not even as the current driver of a live session; their regenerate is
appended beside the owner's answer). Decided once when the turn starts, never on the
commit path. A non-owner, an API-key turn (no user), an invisible conversation, or a
gateway-side fault ⇒ nothing is superseded, the new answer is simply appended (both
remain visible), and the refusal is logged at
Download-request tool forcing (server-side, no header)
When a chat turn's latest user message asks for a downloadable file — e.g. "gib mir die Datei zum Download", "erstelle mir eine Excel-Datei zum Download", "give me the file to download", "export this as a PDF" — the gateway forces a file tool on that turn so the model produces a real file instead of narrating a phantom "here is the file to download: X.xlsx" with no tool call (the prod class: no file, no download card). This needs no header; it is inferred from the user's message.
- When it fires: the latest real user message is classified as a download request
(a deliver-verb — gib mir / give me / erstelle / export / … — plus a file/format token, OR an
explicit "zum download" / "as a download" phrase) and at least one file tool
(
write_fileorcode_interpreter) is resolved for the turn and the provider honors a forcedtool_choiceand the request carries no client/gatewaytool_choicealready. - What it does: for that first leg only, the offered tool set is narrowed to the resolved
file tools
{write_file, code_interpreter}andtool_choiceis set to"required", so the model must call one of them (it still chooses write_file for text/HTML/office vs code_interpreter for computed/formatted output). Continuation legs are unforced (they answer from the tool result). A provider that cannot guided-decode"required"is not forced (logged atWARN); a rare 400 on the forced leg strips the force and retries once. - Precision (fail toward NOT forcing): requests about a file — "gib mir eine Zusammenfassung / Analyse / einen Überblick der Datei", "give me a summary of the file" — and how-to / question / inability phrasings do not force (the user wants an inline answer, not a file). A genuine download ask this misses is caught by the honesty backstop below.
- Backstop: if a turn still ends with the model claiming a download it did not produce, the claim-steer (
detect_file_claim) re-legs it into a real write. The download card itself always renders from the durablechat_message_filerelation, never from model prose.
Bare re-download → server re-serve of the existing file
A bare re-download — the user asks for the file they already have back with no change:
"lade die Datei herunter", "download it", "gib mir die Datei nochmal" — is handled by
re-serving the file that already exists, not by forcing a fresh creation. Forcing a new
creation here is both wasteful and unsafe on a weaker model: at 1000-case scale gemma4's forced
regeneration produced accepted=0 and then simply stopped, leaving the user with no download
card even though a perfectly good file was sitting in the conversation.
- When it fires: the turn is classified as a download request (as above) and carries
no modifier — no format/conversion token (als PDF / as pdf / convert), no version verb
(aktualisier / updated / upgedatet / Ändere / Überarbeite), no create verb (erstell /
mach mir / create), no daraus / draus, and no redirect (email / slack /
@). Any modifier → this does not fire and the normal download-force creation path runs, so "als PDF" still produces a real new PDF. (The modifier list dual-cases sentence-initial umlauts — Ä/Ü do not ASCII-lower to ä/ü — so "Ändere die Datei und lade sie herunter" is correctly treated as a modify, not a bare re-download.) - What it does (two coordinated steps, model-INDEPENDENT): at turn setup the gateway looks up the
conversation's most recent generated file (
chat_project_knowledge WHERE conversation_id = ? AND chat_generated = 1 ORDER BY (filename = ?) DESC, created_at DESC— a file the user named wins over the newest) and, if one exists, skips forcing the model to regenerate and caches the row. Then at the streaming terminal it re-emits that file's download card deterministically — whatever the (now-unforced) model says or doesn't say, including an empty turn or an honest decline. It is the same card the originalwrite_file/code_interpreterturn produced, reconciled intochat_message_fileso it survives a reload. No new model creation leg runs, so a weak model can neither botch nor corrupt it, and no duplicate file is minted. - Neutral re-download card, not "Updated": the re-served link reuses the original
file's
file_idin a new message, so without a marker the client would read it as a same-name rewrite and label the card "Updated" (and mark the original "superseded") — implying a change that never happened. The re-serve therefore persistschat_message_file.reserved = 1(migration 0257; fromctx.reserved_file_ids), and the conversation-messagefiles[]relation carries areserved: trueboolean on that entry. A reserved link renders as a plain re-download: the client ignores it when inferring Created/Updated, so the original card stays Created and the re-download card carries no revision label. Every genuinewrite_filelink isreserved: false. - Trust boundary (owner-gated, fail-closed): the lookup is keyed on the authenticated
user_idviaget_conversation_route_scalars(the owner gate,WHERE user_id) — never on the untrustedx-aig-conv-idheader or any history marker alone. A nil/empty principal, a foreign/shared-reader caller, or a conversation with no owned generated file each yield no card (no metadata leak, no cross-tenant IDOR). A re-served id is tagged inctx.reserved_file_idsso the artifact-drop guard does not misfire on it, and set inctx.written_file_idsso the branch-e file-claim and promise steers do not double-fire. - Precision: if no owned generated file exists (e.g. the first turn is "download the file"), nothing is re-served and the normal download-force creation path runs — a bare re-download only short-circuits when there is genuinely a validated file to hand back.
x-aig-datafile-auto — permit a per-turn model upgrade for data-file turns
The web chat sends x-aig-datafile-auto: 1 when a conversation routed by Auto
(no explicit model pick) sends a turn that carries an inline data-file block (an
attached spreadsheet/CSV preview). The header is a permission hint, never an
authorization: it only allows the gateway to consider upgrading this single
turn to a more capable model; every decision input is validated server-side.
- Accepted shape: exactly the string
"1". Any other value — absent, empty, another spelling, an over-long value, or a repeated header — is treated as absent (no upgrade consideration; the request proceeds unchanged). - Effect: when the request also actually carries an inline data-file block (verified server-side — a hint without a file is inert), the resolved model is not reliably tool-capable per the gateway's curated capability registry (provider/catalog self-claims do not count) with a context window covering the request, staging the raw file is possible on this gateway, and the tenant is not in a reduced-model (over-cap) window, the gateway rewrites this turn only to the best eligible model: providers the gateway can actually dispatch to (configured or managed fleet, intersected with data-residency and provider-allowlist enforcement) ∩ the plan's model allowlist ∩ usable chat models with registry-proven function-tool support and an adequate context window — ranked sovereign-first, then by price. No eligible model → the request proceeds unchanged. The upgrade is never persisted to the conversation; a later turn without a data file routes by the conversation's own model again.
- Disclosure: an upgraded turn reports the actually served model in the
X-AIG-Modelresponse header on streamed responses (X-AIG-LLM-Modelon non-streamed ones), and the assistant message row is persisted with the served model — the web chat renders a per-turn "Answered by …" notice from it. - Security: because the target set is confined to models the caller could already select explicitly on the same gateway, a spoofed hint cannot cross a residency, plan, or authorization boundary — at most it moves the caller's own turn to another model the caller is already entitled to use.
x-aig-auto-complexity — permit a per-turn Auto complexity upgrade
The web chat sends x-aig-auto-complexity: 1 on every turn of a conversation routed
by Auto (no explicit model pick), to mark that the turn used the conversation's
sticky Auto model rather than an explicit choice the gateway must never override. Like
x-aig-datafile-auto, it is a permission hint, never an authorization: it only
allows the gateway to consider upgrading this single turn; every decision input is
validated server-side.
- Accepted shape: exactly the string
"1". Any other value — absent, empty, another spelling, an over-long value, or a repeated header — is treated as absent (no upgrade consideration; the request proceeds unchanged). - Effect: only on a gateway with per-gateway complexity routing enabled
(
gateway_config.laya_routing=shadoworlive), the hint lets the gateway run the self-hosted Laya classifier on the turn's user text. Inlivemode a turn whose hard-probability mass clears the floor is rewritten this turn only to the best eligible frontier model — intersected with data-residency and provider-allowlist enforcement ∩ the plan's model allowlist ∩ registry-proven tool-capable chat models with an adequate context window, ranked quality-first. No eligible model, a below-floor signal, or an easy/medium tier → the request proceeds unchanged. Inshadowmode the verdict is only logged; nothing is rewritten. The upgrade is never persisted to the conversation; a later turn routes by the conversation's own model again. - Disclosure: an upgraded turn reports the actually served model in the
X-AIG-Modelresponse header (X-AIG-LLM-Modelon non-streamed responses), and the assistant message row is persisted with the served model. - Security: the target set is confined to models the caller could already select explicitly on the same gateway, so a spoofed hint cannot cross a residency, plan, or authorization boundary — at most it moves the caller's own turn to another model the caller is already entitled to use. A best-effort classifier failure (Laya unreachable) fails open: routing is unchanged.
x-aig-web-search-freshness — restrict web search to a recency window
When web search is enabled for the request (via x-aig-web-search: 1, or a gateway
configured with mode: "always"), this optional header narrows the underlying search
to a time window so the model grounds its answer in recent results.
- Accepted shape: one of the discrete windows
pd(past day),pw(past week),pm(past month),py(past year), or an explicit inclusive date range in the formYYYY-MM-DDtoYYYY-MM-DD(e.g.2022-04-01to2022-07-30). Anything else — an unknown token, a partial range, wrong separators, extra query fragments, or a non-string — is rejected: the header is ignored (fail closed) and the search runs unfiltered rather than forwarding an unvalidated value to the search provider. - Effect: the validated window is passed to the web-search provider so only results
within that period are returned. When the header is absent or rejected, web search
behaves exactly as before (no recency filter). The filter applies to every
cross-provider path that runs the gateway's own web-search backbone (buffered
OpenAI-format providers and the streaming
vllm/myratool loop); providers that use their own native search (Anthropic, OpenRouter, Gemini, Mistral Agents) ignore it. Recency is a Brave-adapter feature: the EUlinkupadapter (see below) currently ignores the window and searches unfiltered. - In the app: the chat composer's + menu exposes a Recency dropdown (Any time / Past 24 hours / Past week / Past month / Past year) whenever web search is on; picking a window sets this header for the next send. "Any time" sends no header (unfiltered).
tool_choice on a web-search turn is honored one-shot
When x-aig-web-search: 1 injects the gateway's web_search tool and the request carries
a forcing tool_choice ("required", or a named function/tool), the gateway forces
the model to call the tool on the first leg only, then frees it: the injected search
runs, and the model is unforced on the continuation leg so it can produce a final answer.
Without the one-shot drop, a forcing choice would be re-sent on every leg and the tool loop
could never terminate. Non-forcing values (absent, null, "auto", "none") are
untouched. (This differs deliberately from an MCP connector-references loop, which
rejects a forcing tool_choice with 400 — there the tools are resolved server-side and
invisible to the client, so forcing them is meaningless; a header-injected web_search, by
contrast, is a tool the client explicitly opted into. Both paths share one fail-closed
forcing definition.) A provider/model that does not support a forced tool choice is
unaffected by web search; a forcing choice sent to such a provider follows that provider's
own wire behaviour.
Web-search direct answer relays tool_calls and labels the terminal reason
When x-aig-web-search is on but the model answers directly on leg 1 (it does not call
web_search), the gateway returns that leg as the response. On the compat (OpenAI) egress:
- if the model made its own tool call on that leg (a client-provided tool), those
tool_callsare relayed onchoices[0].message.tool_calls(withfinish_reasontool_calls) — the turn is not collapsed to an empty completion; and - otherwise the terminal
finish_reasonis routed through the gateway's single stop-reason authority: a length-capped answer surfaces aslength(so a truncated web-search answer is labelled incomplete, matching the non-web-search paths) rather than a cleanstop; every other terminal reads asstop.
(An empty tool_calls array or a non-array value is not relayed. On a PII gateway the relayed
tool-call arguments are token-restored with the correct depth-2 escaping.)
If the web-search direct answer instead comes back with zero visible content (no text and no
tool call — e.g. a reasoning-only turn), it is classified through the same shared empty-response
path as the non-web-search buffered egress: the turn is counted as an EmptyResponse in the
gateway's telemetry, and the compat egress delivers a short fallback bubble (a "produced internal
reasoning but no visible answer" or "returned an empty response" note) instead of a blank message —
so a web-search direct answer is never a silently empty turn, matching the other egress paths.
Web-search provider selection and EU residency
The gateway's own web-search backbone is provider-pluggable, selected per gateway by
web_search.provider in the gateway configuration:
brave(default) — Brave Search (US). Used for every gateway that does not set the field.linkup— linkup.so (EU, France; GDPR Art. 28 DPA, zero data retention). The EU-sovereign backend for tenants that must keep the model-generated search query in the EU.
The per-gateway search API key stays in the existing web_search.api_key field (it holds the
key for whichever provider is selected). Web search is turned off per gateway with
web_search.enabled: false.
Platform Linkup key injection at read
Every new production gateway — the pay-first default gateway, both trial gateways
(internal and external), and a gateway created by an administrator via
POST /admin/v1/tenants/{id}/gateways — is provisioned with web search on via Linkup —
web_search: { enabled: true, provider: "linkup", max_results: 5 } — so a fresh organisation can
search immediately with no manual setup (AGF-2874; before it, only trial gateways were seeded).
The block is seeded only when the gateway has no web_search block at all: an explicit
web_search.enabled: false, a block with its own api_key (BYOK), another provider, or any
existing value is a deliberate choice and is never overridden. Gateways created with
purpose test, benchmark or archived (E2E fixtures, rigs) are not seeded. Existing trial
gateways were backfilled the same way (migration 0302); existing non-trial gateways that have no
web_search block are not backfilled (a separate decision).
The Linkup api_key for these gateways is not stored in the gateway config. It is a shared
Myra platform key held in the database settings row trial_linkup_api_key (the historical
name — it applies to every plan) — a super-admin sets it live from the admin console
(Feature Flags), taking effect within about a minute. It is injected into the effective config
at read time, for any tenant plan, gated only on the block being enabled, the provider
resolving to linkup, and the block carrying no api_key of its own. It is a low-value
search key, not a credential: a plain setting — the settings GET returns it verbatim and
it is admin-editable live (see Global settings API). Consequences:
- The stored gateway config (and therefore the admin
GET/list responses and the tenant config export) contains theweb_searchblock but never the platformapi_key— the shared key is injected only at read time and is not written into per-gateway config or exports. This holds for every plan, including callers holdingGATEWAYS_MANAGE. - If
trial_linkup_api_keyis unset (or empty), new gateways are still seeded with the keyless block (so a later key set lights them up with no backfill), but they have no web search until then: they reportweb_search_setup_pending: true, the/easycomposer shows a visible "not set up" notice, the Health dashboard flagweb_search_platform_key_missingistrue, the gateway logsERRevery 10 minutes and a[platform_alert]fires (at most every 6 h). Signup and gateway creation never fail because of it. Set it to a dedicated, rotatable Linkup key from the admin console. - A trial → paid conversion changes nothing for web search: the injection is plan-agnostic,
so the converted organisation keeps searching. A customer who wants their own Linkup or
Brave account sets
web_search.api_key(andprovider) on the gateway — an explicit key always takes precedence over the platform key. max_resultsis honored by the Brave adapter but ignored by Linkup (kept for shape parity).
Regardless of plan, the server-computed web_search_configured boolean the gateways
endpoint returns (and the composer web-search globe reads) reflects the effective config —
so a keyless Linkup gateway reports web search as configured via the injected key, without the
key ever reaching the client.
- Fail-closed EU gate. When a gateway (or its tenant) enforces EU residency
(
eu_region_routing, including the inheritedtenant_eu_region_routingfloor), the model-generated search query may egress only to an EU-vouched search provider. A non-EU provider — including thebravedefault a sovereign gateway forgot to switch — is blocked: the query is never sent, and the model is told the search was withheld. A sovereign tenant can therefore turn web search off, but can never silently fall back to Brave (US). This also neutralizes the buffered path'swttr.inweather side-channel (which would otherwise embed the raw query in a URL to a US host). The query PII scrub (seeguardrails.md) runs on top of this on every provider path. - The linkup response is untrusted third-party content and is validated fail-closed
(principle 11): a non-2xx status, a non-JSON / non-object body, or a body with no
resultsarray yields a classified failure and zero results (never a partial or fabricated answer); every resulturlmust be an absolutehttp(s)URL or it is dropped (ajavascript:/data:/scheme-relative/empty URL can never become a citation or link). linkup'sresults[].{name,url,content}are normalized to the internal{title,url,snippet}shape at the single dispatch point, so downstream formatting/citation code is provider-agnostic.
Inline numbered citations
When the gateway's own web-search backbone runs (the cross-provider vllm/myra path — not
a provider's native search), each search returns a numbered source list to the model, and
the model is instructed to mark the claims it grounds with inline [1], [2] … references.
The answer ends with a numbered Quellen: (sources) footer listing those URLs in the same
order, so [N] in the prose maps to source N in the footer. In the chat UI this footer is
rendered as a collapsed-by-default disclosure ("Quellen (N)") so a long source list does
not flood the answer; the inline [N] markers remain direct links to the source, so a source
is still one click away without expanding the list.
Citations inside a written file. The [N] markers belong to the chat answer
only — a file produced with write_file has no footer to resolve them, so a document that
carried them showed bare [1], [3]. The model is therefore instructed (in the write_file
tool description, the file-tool rules of the composed system prompt and the file-first system
notice — one shared rule, emitted only on turns where write_file is actually offered) not to
use inline [N] markers inside a file and to put the sources in a final Sources /
Quellen section (one line per source with title and URL), or to omit them when the user asks
for a clean document. The gateway does not rewrite file content: what the model writes is what
is stored and delivered.
- Source URLs are untrusted third-party content. They are validated at two boundaries and
fail closed: non-
http(s)result URLs (e.g.javascript:/data:) are dropped when the provider (Brave or linkup) response is parsed and again when the footer is rendered (so Anthropic-native citations, which don't pass through the search path, are covered). Source titles are stripped of markdown-link metacharacters. The SPA re-validates every URL scheme before rendering a link. - Best-effort per claim. Inline
[N]anchoring depends on the model following the instruction (as with Perplexity/ChatGPT Search). The numbered source footer is the guaranteed floor; models that don't emit[N]still get the footer, and an out-of-range or hostile-scheme[N]is left as plain text. Providers with native search cite in their own style and are unaffected.
Auto-disable of injected tools on a non-tool-calling model
Some routed models (for example the Perplexity Sonar family) expose no upstream endpoint
that accepts a function-tools array; injecting the gateway's own function tools (web search,
URL fetch, file access, knowledge, code, connectors) would make the provider reject the whole
request (OpenRouter returns 404 "No endpoints found that support tool use").
For a model the gateway knows cannot do function calling (recorded in its capability registry), the gateway now proactively strips its injected function tools before the upstream call and lets the run proceed as a plain answer — instead of relaying the opaque 404.
The notice has a second, reactive trigger. The proactive strip only
fires for a model the gateway has explicitly curated as non-tool, which cannot keep up with a
catalog that grows (an OpenRouter image model, openai/o1-pro, whatever is added next). So when
a model the gateway did not recognise as non-tool refuses the injected tools at the
upstream — OpenRouter answers 404 "No endpoints found that support tool use" — the gateway
now strips its own injected set, rebuilds the system prompt so it no longer offers tools the
model cannot call, and re-sends the turn once as a plain answer. Same notice, same category
tokens; the only difference is that the strip was learned from the refusal rather than known in
advance. At most one extra upstream call per turn, and only for tools the gateway injected. A caller
that named its own tools gets its provider's error passed through verbatim — both a plain
tools array (the client is driving its own agentic loop, and silently removing its tools would
leave it waiting for tool_calls that never come) and an array of {"type":"mcp"} connector
references (those were explicitly requested, and the typed error names the actual lever:
remove the references or pick a tool-capable model).
Scope, stated plainly: the retry is triggered by the upstream saying it cannot accept tools, so
it covers the refusal phrasings the gateway recognises — OpenRouter's "No endpoints found that
support tool use", and the "does not support tools / tool use is not supported / function calling
is not supported" family. A backend that refuses tools in wording the gateway does not recognise
still surfaces the typed error. That is a shorter and slower-moving list than one entry per model
(openai/o1-pro is not an image model, which is why per-model curation could not keep up), but
it is a list.
The client is told which tool categories were skipped:
- Streaming (chat / playground): an additive SSE status event
{"aig_status": "tools_skipped", "tools": ["web_search", "url_fetch", ...]}just before[DONE]. The SPA renders an informational notice on the assistant message.toolsis a deduped list of coarse category tokens (web_search,url_fetch,files,knowledge,image,code,connectors,connectors_policy, or a generictools).connectorsmeans a connector was unavailable (transient — a retry may help);connectors_policymeans the project forbids external egress, so the connectors could never have run (permanent for that conversation, and nothing is broken). When a dropped connector can be identified, the same event also carries an optionalconnectorsarray —[{"id","name","reason"}]— naming each dropped connector so the SPA can name it (instead of the bareconnectorsword) and, for a credential-classreason(credential_required/credential_expired), offer a Connect/Reconnect deep link to/connectors?focus=<id>. The coarsetoolscategories still cover nameless drops (policy / discovery-deadline). The gateway also injects a turn-scoped system note telling the model the named connector exists but is unavailable this turn (so it asks the user to reconnect rather than claiming the connector does not exist). - Buffered agent
/invoke: atools_skippedtrace step (visible via the run'sX-AIG-Trace-Id/ run-detail) records the skipped categories, the model, and the provider.
Provider-native web search (OpenRouter plugins, Gemini grounding, Anthropic's native
search) needs no function-calling capability and is injected in the provider's request build,
so it intentionally survives this strip and is not listed as skipped — and that is true of the reactive strip too: the notice enumerates only what the gateway
actually removed, so a turn whose OpenRouter web plugin still runs is never told it has no web
search. A tool-incapable model the gateway does not recognize up front is not stripped
here (this section describes the proactive strip, which fires only for a curated
non-tool model) — it is stripped reactively instead, once the upstream refuses the
injected tools, as described above. model_capability_mismatch (see
error codes) remains for the case where that stripped retry ALSO fails, or
where the refusal names a capability stripping cannot fix (an image, a document).
EU-residency exception. Auto-disable + proceed applies only on a non-EU gateway. On an
EU-residency-enforced gateway (eu_region_routing), a definitively-non-tool online model
(e.g. a Perplexity Sonar model) is not auto-disabled — it stays fail-closed. Such a
model's intrinsic web search egresses outside the EU and cannot be stripped by the gateway, so
proceeding would breach the tenant's data residency; rejecting is the safe outcome. In practice
today's such models (Perplexity Sonar via OpenRouter, a non-EU provider) are refused by the
residency dispatch gate with data_residency_blocked before any upstream call; skipping
the auto-disable here simply avoids emitting a misleading tools_skipped notice ahead of that
block. (A future definitively-non-tool model on an EU-vouched provider whose intrinsic search
still egressed would instead be rejected with model_capability_mismatch.)
Web search enabled but not used / no results / failed
When web search is enabled for a turn the gateway keeps tool_choice: "auto" — it never
forces the model to search (that would burn a turn on greetings, rewrites, and
context-answerable follow-ups). A consequence: a weaker model may answer from its own training,
or ask "should I search?", without running a live search — and a search that the model
does run can still fail or return nothing. In all three cases the answer used to look normal
with no signal that no trustworthy live search happened.
The gateway now emits an additive side-channel status event so the SPA can render a distinct chip on the assistant bubble:
reason |
Meaning | Chip |
|---|---|---|
not_used |
Web search was enabled but the model answered without calling it (replied from training / asked whether to search). | "Answered without web search" |
no_results |
The search ran but returned zero hits. | "Web search returned no results" |
failed |
The search provider errored (quota / auth / network). | "Web search failed" |
reason is the only field; a value outside the three above is ignored by the SPA (the boundary
rejects unknown values rather than rendering a raw slug). The event is additive — an older
client ignores the unknown aig_status (no schema bump) — and is delivered only to a client
that opted into the aig_* side channel. It is emitted on the two-leg web-search engine
(not_used at the buffered leg-1 direct answer, re-emitted on the
buffered→SSE tail; no_results/failed inline on the streaming leg-2 wire) and on the
tool-loop web-search path. It is a notice, not an error — the
turn proceeds and the answer is still delivered; the chip only tells the user the answer is not
backed by a fresh search. Live-session only — like the tools_skipped and pii_masked
notices, it is not persisted, so a reloaded conversation carries no chip.
Declining all tools with tool_choice: "none"
A request that carries tool_choice: "none" (the string, or {"type": "none"}) is a
declaration that the model will call no tool this turn. On such a request the gateway
injects no built-in tools — not web search, URL fetch, file access, knowledge, image
generation, nor code interpreter — and strips the now tools-less tool_choice from the
upstream body (a tool_choice with no tools array is rejected by several providers). This
is honored uniformly across every provider, including the pass-through providers that have no
gateway tool loop (Gemini, Vertex, Bedrock).
- Only an explicit
"none"opts out. Absent,null, and"auto"are not a decline — they receive the normal offered tool set. A forcing value ("required", a named tool,{"type":"any"}, …) is not a decline either (and is separately rejected on the MCP path — see the connector table above). - MCP connector references are the one exception: a request that ALSO carries
tools:[{"type":"mcp", …}]still resolves those connectors and runs the loop withtool_choice:"none"(the model declines to call), exactly as the connector table above documents — a client that named connectors asked for that loop. The built-in-tool suppression above applies to the plain (no-refs) request. - The client is never the authorization boundary:
tool_choice:"none"can only reduce the offered tools, never widen them. - The decline is snapshotted from the client's request before any gateway middleware mutates
the body, so a gateway-internal
tool_choice:"none"set on a later leg (e.g. the web-search two-leg synthesis leg) is never mistaken for a client decline.
This is how the /easy surface's post-turn utility completions (auto-title, follow-up
suggestions) stay pure text: they send tool_choice:"none" so a code-interpreter-enabled
gateway does not run the sandbox for a completion the user never sees.
Inline tool-call recovery (self-hosted vllm/myra models)
A self-hosted model (our qwen/vLLM fleet) occasionally emits a tool call — most commonly
web_search — as inline text in the assistant content instead of a structured
tool_calls field, when the serving-side tool parser fails to extract it. The gateway
recovers these so the tool still runs. Recognised dialects:
- the wrapped XML form
<tool_call><function=web_search><parameter=query>…</parameter></function></tool_call>; - a whole-content JSON object
{"name":"web_search","arguments":{"query":"…"}}(also accepted inside a<tool_call>…</tool_call>block, and whenargumentsis a JSON-encoded string); - a whole-content pseudo-call
web_search(query="…")/{web_search(query="…")}.
The recovered call is untrusted model output and is validated fail-closed at this boundary — a call is recovered only when ALL hold, otherwise the text is left as-is (no tool runs):
- Accepted: the JSON/pseudo call is the whole trimmed message OR the sole content of a
<tool_call>…</tool_call>block (the wrapper is itself an anchor, so a wrapped call is honoured even with surrounding prose — consistent with the existing<function=…>behaviour); itsnameis a string on the request's resolved/allow-listed tool set; and itsargumentsis a JSON object with at least one string key. - Rejected (rendered as plain text, no dispatch): an un-wrapped JSON/pseudo tool call
embedded in surrounding prose (a documentation example, a pasted snippet, "what does
this JSON look like"); an unknown / non-allow-listed / hallucinated tool name; missing /
empty / non-object
arguments(incl. a JSON array); malformed or truncated JSON.write_file/read_fileare never synthesized from this ambiguous channel (a mutating write is only dispatched from an unambiguous marker). When awrite_file/read_fileIS emitted as a bare-JSON/pseudo call and therefore not dispatched, the gateway does not lose it silently — it records a fail-loudsilent_write_file_drop/silent_tool_call_droptrace step and anERRlog (so a new serving dialect is caught and patched), rather than letting the reply falsely claim success. - Recognised tool tag the visible stream removed but nothing synthesized: when the model's
entire visible output is an inline tool tag the gateway strips from the bubble (a
<tool_call>/<invoke>/<function_calls>wrapper, or a<write_file>with no parseable filename) but no dialect/allow-listed call is recovered, the assistant turn used to render as a silent empty bubble — the user had to send a bare?for the model to produce the answer on the next turn. It now surfaces a short retry hint ("I tried to call a tool but the call wasn't in a format the gateway could parse. Please ask me to continue…") instead, and records anEmptySuppressedNoTool(gateway_fault) signal so the unrecognised dialect is caught and patched. (A<memory>tag stripped from an otherwise-empty turn is not a tool call and gets neither this hint nor the empty-response fallback — see Empty responses above; a whitespace-only turn does get the fallback, tools offered or not.) - Unwrapped
<write_file>/<read_file>block bounding: in the gemma/qwen inline-XML channel the next tool opener is a hard wall for every block: a block ends at its own close tag when that lies before the wall, otherwise at the earliest of a Hermes terminator (</function>/</tool_call>), the wall, or the end of the message. A close tag beyond the wall belongs to that later block and never ends the earlier one, so two<write_file>blocks where the first was left unclosed yield TWO calls — the first marked torn (the dead-leg diagnosis reads it as such), the second whole — instead of one call for the first file with the second file's markup embedded and the second file silently never written. An open tag whose>lies past the next opener is unparseable and skipped — no call is yielded for it (the next opener is parsed on its own). Bounding is dialect-free (the block's own close tag before the wall ends any block; otherwise the earliest of a Hermes terminator, the wall, or the end of the message); whether a block counts as closed follows its dialect: the attribute form (<write_file filename="…">) closes only with its own</write_file>— a</function>/</tool_call>inside it is quoted content, which bounds the block (the two are indistinguishable on the wire) but leaves it torn; the parameter forms (<write_file>/<function=write_file>with<parameter=…>children) close with the Hermes terminators or their own</write_file>. A body in the attribute form that quotes a Hermes terminator keeps it when its own close tag arrives before the wall. The scan is linear in the message on every axis (blocks, body bytes, repeated<parameter=…>openers): every literal lookup is a plain find, memoised. The wall's inherent trade-off: a body that quotes another tool opener (a write documenting a tool call) is split at that opener — a nested quote and an interleaved torn block are indistinguishable without a depth model — so the quoting block comes out torn and the quoted call is dispatched. web_searchquery typing: on the buffered path a non-stringqueryis not recovered (left as plain text); on the streaming path the call is dispatched and the tool returns an"Error: query required (must be a string)"result (it never crashes).
Degraded write_file sentinel recovery + file-claim honesty guard (streaming)
A weaker model sometimes announces a file instead of calling write_file: it prints a
pseudo-sentinel like DATEI_ERSTELLT:report.html, a bare write_file token, or narrates
"the file is saved — click the file card" while the turn produced zero write_file
artifacts. Previously nothing recovered, stripped, or corrected these — the sentinel
leaked into the visible answer, no download card existed, and the user was stranded
retrying. On streaming tool-loop legs the gateway now:
- Sentinel recovery (whole-content anchored, like the JSON/pseudo dialects above).
Accepted: the entire trimmed message is a single
IDENT:<filename>line (IDENT all-caps, ≥4 chars; the filename a single token with a 1–8-char alphanumeric extension; at most one space after the colon — e.g.DATEI_ERSTELLT:Report_Q2.html,FILE_CREATED: notes.md), a barewrite_file/write_file()token, or a contentless pseudo-callwrite_file(filename="…"). The blob is suppressed from the visible answer and ONE contentlesswrite_filecall is dispatched instead — the tool rejects it deterministically (content required …, before any persistence or cap tick, so the dispatch is provably side-effect-free) and the error text steers the model into a realwrite_filecall on the next leg. Rejected (left as plain text / loud-guarded only): a sentinel embedded in prose or spanning multiple lines; lowercase or short identifiers; a post-colon token without an extension (WICHTIG: Bitte …); a pseudo-call with acontentkey (a mutating write is never synthesized from this ambiguous channel — it stays with the fail-loud guards); any sentinel whenwrite_fileis not among the resolved tools (fail-closed); a sentinel co-occurring with a structured tool call (silent_write_file_dropfires with file provenance instead of dispatching). - File-claim honesty guard (once per turn). When the leg's visible text claims a
produced file while the turn delivered no downloadable artifact and nothing is
dispatched, the gateway dispatches the same contentless steering
write_file(hard one-shot per turn; a claim that merely references an earlier turn's file costs one steering leg and the error text tells the model to answer honestly). A claim is one of: gateway card vocabulary (Datei-Karte/file card) plus a named document file; a sentinel-shaped line mid-prose; a gateway delivery-card term (downloadable card/download card/herunterladbare Karte) on a line that is not a modal offer, a how-to/imperative, or 2nd-person attribution — scanned per line (a modal on another line cannot excuse a phantom line) and, uniquely, without a filename, since the prod repro (opus-5: "the CSV is attached above as a downloadable card") carried none; a verbatim echo of the gateway's own history CREATED_FILE frame ([In this reply you created these downloadable file(s)…: X]), a frame that is gateway scaffolding and never legitimate model prose (so, like the sandbox surface below, it needs no exclusion filter); a download-offer term (zum Download/to download/for download/as a download) on a non-excluded line plus a document filename anywhere (the offer and filename may be on different lines); or a hallucinated markdown download link with thesandbox:scheme ([name.xlsx](sandbox:/mnt/data/name.xlsx)) — the OpenAI/ChatGPT training-data artifact path that resolves to nothing here (no producer or resolver exists in the backend or SPA), which weak models (e.g.gemma-4-26b-a4b-it) emit instead of callingwrite_file, so the client renders a dead, non-clickable link and no download card. Untrusted-input shape (principle 11): the sandbox surface is anchored on the presented-link opener](sandbox:— a rendered clickable link — so it is rejected/steered; a baresandbox:prose mention or quoted path (no](opener) is accepted (not steered), as is a link to any non-sandbox:scheme. Because](sandbox:has no legitimate producer it is not run through the modal/how-to/2nd-person exclusion filter (a dead link is a phantom even when framed 2nd-person, e.g. "Du kannst x öffnen"); the accepted residual is a user pasting ChatGPT output that the model then reproduces (pasted or illustrative) on a no-artifact turn → one bounded steer leg, since the honesty clause only forbids claiming a saved file; a folder claim —export folder/Export-Ordner/Export Ordner/Exportordner, per line, without a filename, since the prod repro ("Fertig — bitte prüfen Sie den Export-Ordner") carried none: excluded by a how-to, a filesystem path or menu crumb ("unter Einstellungen > Berichte") anywhere on the line, or by a first-person offer ("I can put it in the export folder"), a negation ("there is no export folder here") or a question mark in the sentence holding the term (so a courteous closer — "…Export-Ordner. Brauchen Sie noch etwas?" — never excuses the claim) — not by 2nd-person framing, because no export folder exists in this product (the same structural rule as thesandbox:link: the natural prod phrasings are 2nd-person); or a delivery claim without a card term —saved/gespeichertin the same sentence as an office-format token (PDF,DOCX,XLSX,PPTX,ODT,ODP,ODS— with or without a filename, so "in eine PDF-Datei umgewandelt und gespeichert" counts; the filename, when named, is the last office filename of the sentence, the delivery target — a heuristic: a source named after the target still wins, harmlessly, since the steering call is contentless), excluded by how-to words, a filesystem path or a menu crumb anywhere on the line, or on one of the three preceding non-blank lines when that line was instructional in shape — a numbered step, a bullet, a table row, a heading, a "…:" lead-in, or an explicit how to / Anleitung phrase ("Öffnen Sie Word, klicken Sie auf Datei > Exportieren. Die Datei ist dann als PDF gespeichert." is a how-to answer, and so is the same answer rendered as steps, bullets or a table, where the instruction words sit above the concluding step; a merely courteous line — "Use the button below." — does not silence what follows it, and a fenced code block consumes the carry like any other line; the same rule applies to the folder branch), and — in the claim sentence — by a present-passive description ("can be saved", "wird … gespeichert", "…, dass X gespeichert wird"), 2nd-person attribution of the save ("the report.pdf you saved", "Sie haben … gespeichert", and with words in between: "Die Datei, die Sie hochgeladen haben, wurde als angebot.pdf gespeichert"), a negation (an honest decline never steers — "wasn't saved", "nichts gespeichert"), a date/year (provenance: "wurde 2023 im SharePoint gespeichert"), or a final question mark ("Wurde report.pdf gespeichert?"), or awie … herunterladeninstruction idiom; or a readiness / delivery-availability idiom — a German<verb> bereit(liegt/liegen/steht/stehen, then whitespace, thenbereit) in the same sentence as a document filename ("Outline_Template.md liegt bereit" — the confirmedclaude-opus-5prod escape, conv 87c8d619: a completion claim carrying no card term, no download-offer term and nogespeichert/savedparticiple, so every other branch misses it and the turn ended with neither a card nor an honest notice). It is German-only by design: the English "is ready" and German "ist fertig"/"ist bereit" are generic completion adjectives that routinely describe an operation over the user's OWN upload ("die Auswertung von umsatz.xlsx ist fertig") or advice ("make sure report.md is ready before you deploy"), so they are deliberately excluded (precision-first). The filename must be the subject of the idiom — it directly precedes<verb> bereit(markdown emphasis and whitespace allowed between), with a document extension from the defaultCLAIM_DOC_EXT(so a.mdcounts — unlike the delivery branch's office-only set). Subject-adjacency is what keeps it precise: it does not fire when a genitive/prepositional OBJECT marker sits before the filename — either directly or with one intervening noun — from a closed class (von/vom, the genitive/dative determinersdes/der/dem/den/dieser…, the genitive-inflected possessivesIhrer/des/seiner…, and the object-prepositionszu/zur/zum/aus/mit/bei): "die Auswertung von umsatz.xlsx liegt bereit", "die Auswertung der Datei umsatz.xlsx liegt bereit", "die Zusammenfassung des Dokuments bericht.pdf steht bereit" — the analysis is ready, not a produced file. The NOMINATIVE determinersdie/dasare deliberately NOT markers, so apposition to a real subject still fires ("die Dateien report.md und notes.csv liegen bereit"). It also does not fire on a filename that sits after the idiom ("Folgende Vorlagen liegen bereit: vertrag.docx" — a RAG list of files that already exist), nor when excluded by a same-line 2nd-person / modal / how-to token or — in the claim sentence — by a negation (the contiguous idiom also self-excludes "liegt nicht bereit"), a first-person offer, a 2nd-person save-attribution ("Die von Ihnen bereitgestellte report.pdf liegt bereit"), or a final question mark. Named residuals (each costs one bounded steer leg): a bare possessive with no attribution verb ("Ihre umsatz.xlsx liegt bereit" — nominativeIhre, the file is the subject) FIRES; a genitive object reached by TWO+ intervening nouns ("die Auswertung der neuen Datei umsatz.xlsx liegt bereit") FIRES (a full fix needs German case parsing, which no branch here does); a non-contiguous idiom ("liegt zum Abruf bereit") and the German verb-BRACKET where the subject follows the verb ("steht die datei.pdf bereit") are MISSED. This is a BACKEND-only steer — unlike the SPA phantom-notice detector, it has no frontend twin: the lead-less idiom cannot be carried precisely by that detector's lead-anchored, exclusion-free design, so the authoritative backend re-drive is the fix. Fenced code blocks (or ~~~) are not scanned by any of the (g) sub-branches** — a quoted script or log line (*"`print('report.pdf saved')`"*) is not a delivery claim. A fence is three or more markers and closes only on the same character, so an inline-code span never opens one and a-fence is not closed by a~~~line inside it; an UNCLOSED fence suppresses every later claim (fail-safe). 4-space-indented code, inline single-backtick code, quoted log prose and a blockquoted fence are not fences (accepted residuals). URLs are removed before the term, format-token and filename scans of all three (g) sub-branches, so neither a vendor help link nor a foreign path can become a claim or the recovered name. All three (g) sub-branches run before branch (a) (the card-vocabulary + filename branch), so when a turn matches both, the recovered filename is the delivery sentence's own — or, when that sentence names none, no filename, and the steer lands inwrite_file's filename-required branch (all carry the same honesty clause; the readiness sub-branch always names one, since it requires a filename). The correctivewrite_fileerror text denies the invented location without denying the product's real delivery paths ("There is no export folder the user can browse: a file reaches them only as a download card from awrite_fileorcode_interpreterrun that succeeded (a file in another program is unaffected)." — stated about the product, never about this turn's deliverables: the same filename-validation branch is reachable after a successful write) — the model must not be told that a Word or Lexware export folder does not exist, nor that the sandbox's own/outputpath (which thecode_interpretertool description tells it to write to) is unavailable. Accepted residuals of the folder/delivery branches (each costs one bounded steer leg): a third-party-software answer naming an export folder without a path or how-to marker; an undated statement of where an existing office document is stored ("ist im SharePoint gespeichert"); third-party attribution ("Word hat die Datei als PDF gespeichert"); a memory/settings sentence carrying a bare format token; indented or inline-backtick code, quoted log prose and a blockquoted fence (only fenced blocks are skipped); a how-to answer under a heading that carries no how-to vocabulary ("## Wie Sie ein PDF erzeugen"), which therefore cannot arm the carry; drafted correspondence or a translation quoting a save; and a possessive 2nd-person provenance line ("Ihre Datei wurde als angebot.pdf gespeichert"). In the other direction (missed phantoms, equally deliberate): a claim within three lines of a how-to token, a claim sentence that happens to pair a 2nd-person pronoun with an attribution verb, andzwischengespeichert. Rejected (no steer): any turn that actually produced a downloadable artifact — awrite_filecard, a persisted written-file id, acode_interpreteroutput data file, a live ephemeral artifact, or a generated image (the airtightturn_produced_deliverablegate, tested by emptiness* so a failed image/CI attempt that only lazily-inits its registry still steers a phantom); a same-line modal / how-to / 2nd-person token; a bare "attached above" without card vocabulary (ambiguous with a user-upload reference). Detection runs on the visible text only — model thinking* about file cards never fires it. The delivery-card vocabulary and thesavedparticiple are shared, test-tied (test_system_prompt_composerF6), with the system prompt's document-delivery wording and thewrite_filenudge so the three cannot drift; the folder vocabulary is detector-only (the no-file-tool prompt already denies an export folder). The full fire / no-fire / accepted-residual corpus is executable:test_inline_tool_callsC12–C17. - Observability:
file_sentinel_recovery/file_claim_no_artifactWARN logs + trace steps (filenames gated behindinclude_bodies, as for the drop guards).
Errors
Inference errors use a structured JSON envelope and set the X-AIG-Error response
header to the error code:
{
"error": {
"code": "invalid_request",
"message": "Human-readable detail about what went wrong"
}
}
Branch logic on the stable code, never on the human-readable message. The full
list of codes (HTTP status, cause, and the rate_limited / quota_exceeded /
guardrail_blocked specifics) is in Error codes.
An error raised after the stream has already opened
On a gateway with a PII detector, a streaming turn's response is committed early — the
gateway emits a {"aig_status":"scanning_pii"} progress event before it starts scanning, so the
client sees HTTP 200 and Content-Type: text/event-stream immediately.
If the request is refused after that point but before the model is called, a fresh HTTP status is no longer possible. The gateway therefore delivers the error over the open stream and ends it normally:
data: {"aig_status":"scanning_pii"}
data: {"aig_status":"provider_error","error_class":"forbidden","message":"…","user_message":{"body":"…"}}
data: [DONE]
So on a PII gateway a refusal that would otherwise be 403/413 arrives as HTTP 200 plus a
terminal provider_error frame. Clients must treat a provider_error event as an error
regardless of the HTTP status; error_class carries the same stable code the JSON envelope would
have used.
The response status is not the whole story for such a turn, so the gateway records the logical
outcome separately: the request log's status column carries the real typed status (e.g. 403),
not the 200 seen on the wire, and meta.aig_stream_error holds the error code. Previously nothing was recorded at all and these turns were indistinguishable from successful ones.
See also
- API quickstart — first request in five minutes
- Providers overview
- Authentication
- Error codes
- Playground