Gateway configuration reference
Every gateway has a config JSON object that controls authentication, caching, timeouts, security, routing, and provider settings. Update it with PATCH /admin/v1/gateways/{id}.
The config is merged at the top level on each PATCH — only the fields you include are changed. Nested objects (rate_limit) are replaced in full when provided.
🔑 Reading the config back — credential fields are role-scoped.
GET /admin/v1/gateways/{id}and the tenant gateways listing are open to any member of the tenant, so the third-party credential fields below —web_search.api_key,semantic_cache.embedding_api_key,webhooks.secret, the wholesiemblock, andtracing.headers— are returned only to a caller holdingGATEWAYS_MANAGE; a lower-privileged reader gets them withheld. See the read-scoping note in the gateways API reference.
Full default config
{
"auth_required": true,
"budget_usd": null,
"budget_period": "monthly",
"tenant_budget_usd": null,
"tenant_budget_period": "monthly",
"cache_ttl": 0,
"retry_count": 2,
"timeout_ms": 60000,
"log_payloads": true,
"rate_limit": null,
"ip_allowlist": [],
"guardrails": [],
"circuit_breaker": null,
"webhooks": null,
"siem": null,
"azure_endpoint": null,
"azure_deployment": null,
"azure_api_version": "2024-02-01",
"azure_resource": null,
"bedrock_region": "us-east-1",
"vertex_project": null,
"vertex_region": "us-central1",
"cf_account_id": null,
"hf_endpoint": null,
"provider_base_urls": {},
"provider_allowlist_enforced": false,
"provider_allowlist": [],
"tracing": null,
"web_search": null,
"semantic_cache": null,
"agentic_fetch": null,
"url_fetch": null,
"av_scan": null,
"prompt_caching": { "enabled": true, "ttl": "1h" },
"context_compaction": {
"enabled": true,
"threshold_tokens": 200000,
"keep_last_turns": 10
}
}
Core fields
| Field | Type | Default | Description |
|---|---|---|---|
auth_required |
boolean | true |
Require a valid x-aig-token, Authorization: Bearer, or x-api-key header on all inference requests. Set to false only for development. |
budget_usd |
number | null | null |
Gateway-level spend cap in USD for the current budget period. Blocks all requests once exhausted. null = no cap. |
budget_period |
string | "monthly" |
Period over which gateway spend is accumulated. One of: "daily" (resets each day at midnight in the server's local time zone), "monthly" (resets on the first of each calendar month, local time), "total" (lifetime, never resets). |
test_headers_allowed |
boolean | absent | Reserved for Myra's test tooling; never set on a customer gateway. Platform-admin-only: on a config write by any other actor (a tenant admin, or a custom role holding GATEWAYS_MANAGE) the submitted value is ignored and the stored value kept — on PATCH and on a POST to an existing slug alike, so no one below a platform admin can arm or disarm a gateway — and the 2xx response then carries "ignored_fields": ["test_headers_allowed"] when the dropped value differed from the stored one (omitting the field, or re-sending the stored value, is a plain no-op). A platform admin's value is applied as sent; a platform admin's non-boolean value that differs from the stored one is rejected 400 ("test_headers_allowed must be a boolean"). |
tenant_budget_usd |
number | null | null |
Tenant-level spend cap in USD. Applies across all gateways belonging to the tenant. null = no cap. |
tenant_budget_period |
string | "monthly" |
Period for the tenant-level budget. One of: "daily", "monthly", "total". |
budget_alert_pct |
number | null | 0.8 |
Soft-alert threshold as a fraction of the budget cap: when current-period spend crosses this fraction of any configured budget (token / tenant / gateway) without being over it, a proactive budget_threshold signal fires once — a budget_threshold webhook (if webhooks is configured) plus an ops notification (log + operator alert e-mail + alert webhook + Mattermost). This is advisory only and never blocks the request; the hard stop still happens at 100%. Must be a number in the open interval (0, 1); 0 disables the soft alert; null, absent, or any out-of-range/non-number value falls back to 0.8 (fail-safe read). The alert fires at most once per (scope, entity, period, cap); a new period or a changed budget re-arms it. Note: the SPA's "Budget Warnings" badge is fixed at 80% and is independent of this knob (it changes only the webhook/notification timing, not the badge). |
pii_masking_enforced |
boolean | false |
Tenant-level. When true, PII masking is mandatory for every user of the tenant: the chat privacy toggle is locked, the client hint_pii_preference opt-out is ignored, routing is forced through a PII-masking gateway (or the request is rejected fail-closed), manual unmask is denied, and every third-party tool egress (web search, fetch_url, MCP arguments) is scrubbed or blocked. A wholly-local (Myra/EU) model leg — the resolved primary and every failover fallback first-party — is never masked on the model input (masking a first-party leg protects nothing and corrupts the user's own data); its tool egress is still guarded, and agent-invoke output (delivered externally) always masks. If any provider in the egress set is external, the leg is masked fail-closed. Set on the tenant, not per-gateway. See PII Protector → tenant-enforced masking. |
cache_ttl |
integer | 0 |
Response cache TTL in seconds. 0 disables the cache. Cached responses are keyed on SHA-256(provider:model:canonical_body). |
retry_count |
integer | 2 |
Maximum number of retry attempts against the primary provider on a retryable status (500, 502, 503, 504, or 429 — the last honouring the Retry-After / Retry-After-Ms header) before the fallback chain is walked. |
timeout_ms |
integer | 60000 |
Per-upstream-request timeout in milliseconds. Applies to each attempt individually, not the total request time. Bounds the connection, the request send, and a buffered (non-streamed) read; a streamed response's read phase is bounded by the built-in stall budgets (20 s to the first byte, 300 s per inter-chunk gap) instead — to bound a streamed read explicitly, set timeout_ms on a routing rule (see the routing rules API). The admin UI pre-fills 120000 when a gateway is created; the 60000 default applies only when the key is absent. |
log_payloads |
boolean | true |
Store request and response bodies in the log table. Disable for sensitive workloads where prompt/response content must not be persisted. |
max_parallel_tools |
integer | — | Maximum number of tool calls the gateway runs in parallel within one assistant turn. Raises the built-in limit. A value of 0 or below defers every tool call to a separate turn. |
Rate limiting
| Field | Type | Default | Description |
|---|---|---|---|
rate_limit |
object | null | null (disabled) |
Gateway-level sliding-window rate limit applied to all callers. Default: null (disabled). Example: {"requests": 100, "window_sec": 60}. Per-token limits are checked independently — a request can be blocked by either limit. Accepted shape (validated on POST /gateways and PATCH /gateways/{id}): a JSON object whose only keys are requests and window_sec, or null to clear. Rejected with 400: a string, number, boolean or array; an unknown key (a typo such as reqs used to be stored and silently mean 100 per 60 s); a requests / window_sec that is not a whole number ≥ 1 (a string "100", 0, a negative or fractional value, nan/inf) or above the ceilings (requests ≤ 1,000,000,000; window_sec ≤ 31,536,000 — one year). A key sent as null is treated as omitted and is not stored. At request time, a stored value the gateway cannot use — not an object (a string, number, boolean or array), or a present requests / window_sec that is not a whole number in range — makes the gateway refuse every request with 500 configuration_error and log the reason once per worker (until the value is corrected) — never "no limit" and never a crash per request. A row written before this validation existed is otherwise served as-is: extra keys are ignored at read, a null key takes the default. |
rate_limit.requests |
integer | 100 |
Maximum requests allowed in the window. Used when rate_limit is set but this field is omitted or null. |
rate_limit.window_sec |
integer | 60 |
Window duration in seconds. Used when rate_limit is set but this field is omitted or null. |
⭐ Example:
IP allowlist
| Field | Type | Default | Description |
|---|---|---|---|
ip_allowlist |
array of strings | [] |
CIDR blocks permitted to call this gateway. An empty array allows all source IPs. Requests from IPs outside the list return 403 forbidden. |
⭐ Example:
Guardrails
| Field | Type | Default | Description |
|---|---|---|---|
guardrails |
array | [] |
Ordered list of guardrail configs. Evaluated in array order within each tier. First block verdict short-circuits the pipeline. |
💡 Note: The legacy key
detectorsis still accepted as a fallback, but only whenguardrailsis absent (the key is missing) or the booleanfalse. A present-but-nullor otherwise non-arrayguardrailsvalue takes precedence and resolves to zero detectors (it does not fall back todetectors). New configurations should useguardrails.
Each guardrail object has a common set of fields plus type-specific fields:
| Field | Type | Description |
|---|---|---|
type |
string | Guardrail type: regex, keyword, jailbreak, json_schema, contains_code, gibberish, language, custom_pii, presidio, prompt_guard, pii_protector. |
name |
string | Human-readable name used in block messages and logs. |
action |
string | One of block, scrub, or flag. |
target |
string | One of request, response, or both. |
See the Guardrail Pipeline page for full per-type field documentation.
Provider-specific fields
Azure OpenAI
| Field | Type | Default | Description |
|---|---|---|---|
azure_endpoint |
string | null | null |
Azure OpenAI resource URL, e.g. https://myresource.openai.azure.com. Required for Azure provider. Routed through the same SSRF guard as provider_base_urls: it is resolved once and the dial IP is pinned; a value resolving to a private/internal address is rejected at dispatch. |
azure_resource |
string | null | null |
Azure resource name. An alternative to azure_endpoint: the gateway builds the URL https://<resource>.openai.azure.com/… from the resource name and azure_deployment. Accepted: a bare vendor label (1–63 chars, alphanumeric plus -/_, starting alphanumeric). Rejected: any value with a dot, slash, colon, @, or whitespace — these would break out of the host authority and are refused (the request fails with a configuration error). |
azure_deployment |
string | null | null |
Azure deployment name. Replaces the model name in the request URL path. Accepted: a single path segment (may contain dots/-/_). Rejected: any value with a slash, ?, #, :, @, or whitespace (path/query injection). |
azure_api_version |
string | path-dependent | Azure OpenAI API version appended as the ?api-version= query param. Defaults to "2024-02-01" on the OpenAI-provider path (a gateway with azure_endpoint set) and to "2024-10-21" on the dedicated azure provider path. |
AWS Bedrock
| Field | Type | Default | Description |
|---|---|---|---|
bedrock_region |
string | "us-east-1" |
AWS region for Bedrock API calls. Used in SigV4 request signing. |
Google Vertex AI
| Field | Type | Default | Description |
|---|---|---|---|
vertex_project |
string | null | null |
Google Cloud project ID. Required for Vertex AI; if unset, falls back to a deployment-wide default project set by your operator. If neither is set the request fails closed with a configuration_error (no empty-project URL is dialed). |
vertex_region |
string | "us-central1" |
Google Cloud region for Vertex AI API calls. Falls back to a deployment-wide default region when unset. |
Data residency (EU model hosting)
For deployments where model data must stay in the EU (a hard requirement in
regulated / public-sector tenders), a gateway can be put into EU-region
routing mode. When enabled the gateway fails closed: any request whose
resolved provider/region is not EU-hosted is rejected with HTTP 403
(data_residency_blocked) — it is never sent to a US endpoint.
| Field | Type | Default | Description |
|---|---|---|---|
eu_region_routing |
boolean | false |
When true, only EU-hosted providers may serve this gateway. Coerced to a strict boolean at config load — a non-boolean value is treated as false (disabled) and logged. |
azure_region |
string | null | null |
Operator-asserted EU region of the Azure OpenAI resource (e.g. swedencentral). The Azure host does not encode a region, so this declaration is what the residency guard validates. |
What counts as EU-hosted (fail-closed allowlist):
- AWS Bedrock —
bedrock_regionis an EU-member region:eu-central-1,eu-west-1,eu-west-3,eu-north-1,eu-south-1,eu-south-2. Only a bare single-region on-demand model id (e.g.anthropic.claude-3-haiku-20240307-v1:0) is vouched under enforcement — it pins processing to that one region. AWS cross-region inference profiles (eu.anthropic.*,us.anthropic.*,global.anthropic.*) are refused: a geographic profile has a multi-region destination set that is not provably EU-27 — AWS's own "EU" geography for Claude routes toeu-central-2(Zurich, Switzerland), which is outside strict EU-27. (Admittingeu.profiles would require widening residency scope to EU + adequacy incl. CH — a deliberate policy change to the EU-region allowlist, not enabled by default.) - Google Vertex —
vertex_regionis an EU-member region:europe-west1,europe-west3,europe-west4,europe-west8,europe-west9,europe-west10,europe-west12,europe-north1,europe-north2,europe-central2,europe-southwest1. - Azure OpenAI —
azure_regionis an EU-member region:westeurope,northeurope,germanywestcentral,germanynorth,swedencentral,francecentral,francesouth,polandcentral,italynorth,spaincentral. - Myra-hosted models (
myra/vllm) — run on the on-prem EU estate. - Mistral AI (
mistral) — the gateway calls only Mistral's EU-hosted endpoint (api.mistral.ai, EU by default); it never enables Mistral's opt-in US endpoint. A fixed EU-direct provider with no region to configure, so it is vouched like Myra. (Requires a Mistral API key configured for the gateway.)
Everything else — the native anthropic / openai / gemini adapters (US
endpoints), any US SaaS provider, custom providers, and UK /
Switzerland / Norway regions (not EU-member — EEA is not in scope) — is
rejected. A per-provider provider_base_urls override is also rejected under
enforcement (the gateway cannot vouch for an arbitrary host).
The guarantee holds on every model dispatch path, not just the primary
request. The check sits at the single upstream dispatch chokepoint, so the same
fail-closed refusal applies to the primary attempt and every failover
attempt, each tool-loop leg, the conversation summarize call, and the
agentic-fetch inner call — all of which funnel through that one chokepoint. It
also covers the web-search query egress (a non-EU search provider such as
Brave is blocked; only an EU provider such as Linkup may receive the
model-generated query, and provider-native search is forced through the EU
two-leg), so an armed gateway can turn web search off but never silently fall back
to a US search backend. Embeddings and image generation are served solely
by the on-prem Myra EU fleet (MYRA_BASE_URL) and never reach a commercial US
provider. Net effect on an armed gateway: no user request — on any path — reaches
a US region; a US route is refused before any network call rather than
egressing.
Routing the big three via EU adapters. Point the model at the EU cloud adapter (routing rule / catalog target) and set the EU region:
# Anthropic → AWS Bedrock (Frankfurt); OpenAI → Azure OpenAI (Sweden);
# Google → Vertex (Netherlands).
curl -X PATCH "https://<host>/admin/v1/gateways/<ID>" -H "Content-Type: application/json" -d '{
"config": {
"eu_region_routing": true,
"bedrock_region": "eu-central-1",
"vertex_region": "europe-west4", "vertex_project": "<proj>",
"azure_resource": "<eu-resource>", "azure_deployment": "<deployment>", "azure_region": "swedencentral"
}
}'
A deployment-wide default is available as an operator-only setting, applied on the sovereign/EU stack. It is not applied on a shared stack, because it would flip every gateway to enforced and 403 existing US-routed gateways.
EU-hosted model names & regions
Three vendor-independent EU-27 model routes are available:
| Route | Provider | EU-27 models / regions |
|---|---|---|
| Myra-Eigenbetrieb | myra / vllm |
On-prem EU inference fleet (chat, embeddings, image generation). |
| Mistral AI | mistral |
mistral-large-latest, mistral-small-latest, codestral-latest, … — EU host api.mistral.ai. |
| Google Gemini via Vertex | vertex |
Gemini 2.5 Pro / 2.5 Flash / 2.0 Flash / 1.5 Pro / 1.5 Flash on a europe-west* regional endpoint (europe-west4 Netherlands, europe-west1 Belgium, europe-west3 Frankfurt, …). A regional endpoint pins processing to that single EU-27 region — the usable big-vendor EU-27 route. |
| Anthropic Claude via Bedrock | bedrock |
Bare on-demand Claude ids (e.g. anthropic.claude-3-haiku-20240307-v1:0) in an EU-27 bedrock_region. Cross-region inference profiles (eu.anthropic.*) are refused under strict enforcement (destinations include Zurich/CH) — see the AWS Bedrock note above. |
Newer Claude models are offered by AWS only via cross-region inference profiles, so under strict EU-27 they are not available on Bedrock; Vertex Gemini is the delivered big-vendor EU-27 route. The full A-5 answer (incl. the EU-vs-adequacy policy decision) is maintained by Myra.
Region resolution — the precedence chain
The effective region for a region-bearing provider (Bedrock / Vertex) is resolved from an ordered candidate list; the first candidate that is a valid region token wins:
- the gateway's own region (
bedrock_region/vertex_regionin the gateway config); - the tenant's region (
bedrock_region/vertex_regionon the tenant — set in the tenant settings, see below), inherited by every gateway under that tenant that does not set its own; - the deployment-wide default region set by your operator — now a fallback tier, not the primary source;
- the provider's hard-coded US default (
us-east-1/us-central1) — logged with a WARN so a gateway with no region configured anywhere is visible.
The fall-through is explicit, not OR-coalescing: a non-empty but invalid earlier candidate (a whitespace / uppercase / malformed paste) is treated as "not set" and falls through to the next tier — it can never short-circuit past a valid tenant or env region to the US default.
Per-tenant region SELECT. A tenant sets its Bedrock/Vertex region in the tenant
settings (a dropdown of EU regions, sourced from GET /admin/v1/cloud-regions). The
value is validated server-side as an EU-member region (region.is_eu, fail
closed) — a non-EU or malformed value is rejected with HTTP 400. Clearing it (empty
selection / null) reverts to the deployment or hard-coded default. This is a tenant-column
change, folded into every gateway's cached config, so it propagates within the
gateway-config cache TTL (≤ 30 s) — unlike a per-gateway admin write, which is
invalidated immediately. Allow for the TTL when verifying a tenant region change.
Region value vs the EU guarantee — two distinct controls. The precedence above decides the region value; the EU guarantee is the separate
eu_region_routingenforcement floor. A per-gateway region overrides the tenant region — including to a non-EU value — and when enforcement is OFF that value is used as-is (the region token is still sanitized against header/URL injection, but not EU-checked). To guarantee EU-only routing, enableeu_region_routing(per gateway or the tenant floor): any resolved non-EU region is then blocked at dispatch (403data_residency_blocked), fail closed. Setting a tenant EU region alone is a default, not a floor.
Provider allowlist (no silent fallback)
Independent of data residency, a gateway can pin an explicit set of approved
providers so it can never silently route — primary or fallback — to a provider
the operator did not approve. Enforced at the same single
dispatch chokepoint as residency: a request whose resolved provider is not on the list
is rejected with HTTP 403 (provider_not_allowed) before any upstream call.
| Field | Type | Default | Description |
|---|---|---|---|
provider_allowlist_enforced |
boolean | false |
When true, only providers named in provider_allowlist may serve this gateway. Coerced to a strict boolean at config load — a non-boolean value is treated as false (disabled) and logged. |
provider_allowlist |
array of strings | [] |
Approved provider names (e.g. ["myra", "mistral"]). Matched case-insensitively after trimming; vllm folds to myra. Under an armed tenant default this per-gateway list may only narrow the tenant-approved baseline (it is intersected with it) — a gateway can never add a provider outside the tenant list. |
Fail-closed semantics. When enforcement is on, a missing, empty ([]), or
malformed (non-array) allowlist denies every provider — a forgotten or mistyped
list can never silently fall back to allow-all. Non-string array elements (numbers,
null, nested objects) are ignored and can never widen the set; a near-miss or
substring of an approved name never matches. When enforcement is off the list is not
consulted (zero behaviour change for existing gateways).
Scope. The allowlist governs routed chat/inference dispatch. Fixed EU-infrastructure services that call the Myra on-prem fleet directly — embeddings (RAG), image generation/editing, transcription, and vision image-analysis — always use that sovereign fleet and are not routable providers, so they are not gated by this list.
A deployment-wide backstop is available as an operator-only setting, applied on the sovereign stack: with it applied, every gateway must
carry a valid allowlist or its inference fails closed. It is not applied on a shared
stack, because it would 403 every gateway that has not configured an allowlist. An
unrecognised value is treated as disabled and logged at startup.
Per-tenant sub-processor objection (Art. 28 Abs. 6 deactivation)
Orthogonal to residency and the allowlist, a per-tenant sub-processor objection narrows the
effective provider set for one tenant without touching any gateway config or any other tenant.
When Myra Ops records an objection (see the
objection API),
the provider(s) mapped to that sub-processor are deactivated for that tenant: they are dropped
from the tenant's model-offer set and refused at the single dispatch chokepoint (primary, fallback,
failover, and the observed LiteLLM-overflow leg) with 403 provider_objected — the same posture as
the allowlist deny, but scoped by tenant_id and driven by the objection store rather than a config
flag. The deactivated set is derived at config load (a {provider: true} set on the resolved
gateway config, never stored in gateway.config, so a tenant admin cannot PATCH away their own
objection), and it fails closed: a hard read error on the objection store fails the whole config
load rather than serving the objected provider. It is enforcement-flag-free — presence of an
objection is itself the enforcement, so it works on an ordinary non-sovereign gateway too.
Tenant-level residency default (default ON for a sovereign tenant)
For a sovereign tenant on a shared stack (where the deployment-wide backstops
above are not an option — they would flip every tenant), the three residency controls
above (eu_region_routing, provider_allowlist_enforced, provider_allowlist) can be
set as a tenant-level default. Every gateway under that tenant then inherits the
default — including a newly created gateway with an empty config, which is the
behaviour a public-sector "commercial models exclusively via EU regions" obligation
requires: enforcement must be default ON, never fail-open on a forgotten per-gateway
flag.
- Where it lives: tenant columns
tenant.eu_region_routing,tenant.provider_allowlist_enforced,tenant.provider_allowlist.eu_region_routingis settable in the product: Settings → Organisation → Routing → Organisation-wide EU routing (platform admins only — lowering an organisation's residency floor is not self-service), and onPATCH /admin/v1/tenants/{id}aseu_region_routing. It is tri-state:1armed,0explicitly not armed,null(the empty option) = no organisation default at all, each gateway governs itself.provider_allowlist_enforcedandprovider_allowlistremain operator/DB-managed only (a migration or adb.shUPDATE) — there is no admin-UI / PATCH field for those two. - Precedence — a FLOOR (per enforcement control): when a tenant enables a residency
control (
eu_region_routing/provider_allowlist_enforced), that is a floor — a gateway can only raise the control, never lower it. A gateway that setseu_region_routingtofalse/0/"off"under a EU-only tenant is still enforced (it cannot opt a sovereign tenant out of EU-only routing — that would be fail-open on the "commercial models exclusively via EU regions" exclusion criterion). The floor covers the two enforcement booleans and theprovider_allowlistlist contents: when the tenant armsprovider_allowlist_enforcedwith a baseline list, a gateway that supplies its ownprovider_allowlistmay only narrow it — its entries are intersected with the tenant-approved list, so a gateway can never approve a provider the tenant did not (a gateway list of["openai"]under a tenant baseline of["myra","mistral"]resolves to[]— deny all — becauseopenaiis not tenant-approved). A gateway with no own list inherits the baseline. For a fully-sovereign posture arm botheu_region_routingandprovider_allowlist_enforced(as the bid tenant does) — the residency guard is then the hard EU floor regardless of any gateway's allowlist. A gateway's own enforcement-flag value governs only when the tenant default is a recognized tenant default is a recognized OFF (0/false) — the ordinary per-gateway opt-in for non-sovereign tenants. An unrecognized/garbage tenant value (e.g.2) is a mis-provisioned compliance flag and fails closed to ENFORCED (with a warning), never a silent US-leak. For a residency control an absent ornullper-gateway value resolves to the tenant default (the safe ON under a floor). The deployment-wide operator backstop still forces enforcement on where set. - Fail-closed: a tenant that enforces the allowlist but whose
tenant.provider_allowlistis missing/malformed denies every provider on its inheriting gateways (same fail-closed rule as the per-gateway list). - Accepted shape — enforced by the database:
tenant.provider_allowlistmust be SQLNULLor a valid JSON array (e.g.["myra","mistral"]). A DBCHECKconstraint (migration 0190) rejects any other write — invalid JSON (a bare-token[myra,mistral]from a mis-quoted provisioning statement), a JSON object, a scalar, or a literalnull— withCONSTRAINT chk_tenant_provider_allowlist_json failed, so a malformed value can no longer be stored at all. Should a pre-constraint/foreign row still carry a non-array value, the gateway treats it as absent (fail-closed: deny-all while enforced) and logs it at ERROR level naming the tenant slug — never silently. - Admin view caveat: the admin gateway GET/PATCH shows the raw per-gateway config, so an inheriting gateway shows these keys absent while enforcement is active at runtime. Read the tenant default to know the effective policy.
- Cache: the per-gateway config is cached briefly (≤ 30 s). A per-gateway change through the admin API invalidates the cache entry immediately; flipping a tenant default (a tenant-column change) still takes effect on a deploy (workers restart) or within the cache TTL, since it is not a per-gateway write.
- Widening for a 3rd EU cloud provider: when a further EU route (e.g. Claude via
Bedrock-EU or Gemini via Vertex-EU) is provisioned, add it to the tenant
provider_allowlist— otherwise it is residency-OK yet allowlist-blocked.
Cloudflare Workers AI
| Field | Type | Default | Description |
|---|---|---|---|
cf_account_id |
string | null | null |
Cloudflare account ID. Required for the Cloudflare Workers AI provider; the gateway builds the request URL from it. Rejected: a value with a slash, ?, #, :, @, or whitespace (path injection); the request then fails with a configuration error. |
HuggingFace
| Field | Type | Default | Description |
|---|---|---|---|
hf_endpoint |
string | null | null |
URL of a HuggingFace dedicated inference endpoint. When set, requests go to this endpoint instead of the serverless API. Routed through the same SSRF guard as provider_base_urls: resolved once with the dial IP pinned; a value resolving to a private/internal address is rejected at dispatch. |
Provider base URL overrides
| Field | Type | Default | Description |
|---|---|---|---|
provider_base_urls |
object | {} |
Map of provider_name → base URL. Overrides the hardcoded default endpoint of the gateway for any provider. Useful for internal proxies, self-hosted endpoints, and staging environments. A base only — the provider's endpoint path is appended, and a query string / fragment on the value is discarded. Validated at save time: the field must be an object mapping supported provider slugs to non-empty http:// or https:// URLs with no userinfo (user@host) and no whitespace / control characters; a value that fails this shape — or a key that is not a provider the gateway can route to (an unknown slug, or a legacy alias such as vllm, which is rejected in favour of its canonical id myra) — is rejected with 400 on POST/PATCH — it can no longer be saved and silently no-op at runtime. Setting the field to null (or omitting every entry) clears the override. The full SSRF guard remains at request time: the override host is resolved once and the dial IP pinned, so an entry whose value is not a non-empty URL string (including false or null) or which resolves to a private/internal address is rejected at dispatch with configuration_error — it never silently falls back to the default endpoint. Use the Test Connection endpoint to check reachability before saving. |
⭐ Example:
Circuit breaker
| Field | Type | Default | Description |
|---|---|---|---|
circuit_breaker |
object | null | null |
Set to enable the per-provider circuit breaker. null disables it entirely. |
circuit_breaker.enabled |
boolean | false |
Must be true to activate the breaker. |
circuit_breaker.failure_threshold |
integer | 5 |
Number of failures within window_sec before the breaker opens. |
circuit_breaker.window_sec |
integer | 60 |
Window in seconds over which failures are counted. The counter's lifetime is window_sec x 2 and is not extended by later failures, so the window is tumbling, not sliding. |
circuit_breaker.cooldown_ms |
integer | 30000 |
Milliseconds to wait in the Open state before allowing a probe request. |
circuit_breaker.failure_status_codes |
array | [500,502,503,504] |
HTTP status codes that count as failures. Connection and timeout errors always count. Must be an array of whole numbers ([] = never trip on status); a malformed value is rejected with 400 on write and fails closed to the default at runtime. |
See Circuit Breaker for state machine details and examples.
Webhooks
Webhooks deliver structured event payloads to an external HTTP endpoint for integration with alerting, ITSM, and automation systems.
| Field | Type | Default | Description |
|---|---|---|---|
webhooks |
object | null | null |
Set to enable outgoing webhooks. null disables them. |
webhooks.url |
string | — | HTTPS endpoint that receives POST requests for each event. |
webhooks.secret |
string | — | Optional shared secret. When set, each delivery carries an X-AIG-Signature: sha256=<hex> header — the HMAC-SHA256 (RFC 2104) of the exact request body, keyed with this secret. Recompute it over the raw body in your receiver to verify origin and integrity. |
webhooks.events |
array | all events | Event types to deliver. Supported values: blocked, budget_exceeded, budget_threshold, circuit_open. When this key is absent (or not an array), every supported event is delivered. Providing an array restricts delivery to the listed events; note that an existing filter that does not list budget_threshold will not receive the new soft-alert event until you add it. |
⭐ Example:
"webhooks": {
"url": "https://hooks.slack.com/services/...",
"secret": "mysecret",
"events": ["blocked", "budget_exceeded", "budget_threshold"]
}
Webhook payload
Every delivery is a POST with a JSON body of the shape:
ts is the Unix epoch (seconds) at send time. The data object depends on the event:
| Event | data fields |
|---|---|
budget_exceeded |
scope ("token" | "tenant" | "gateway" — which budget was hit), budget_usd (the cap), spent_usd (current-period spend), period (the period key, e.g. "2026-07", "2026-07-15", or "total"). Fired when a request is blocked because spend reached the cap — except for self-serve tenants with cap soft-degrade configured, where reaching the base allowance fires it once per billing period with degraded: true and stage: "primary" while requests continue on the efficient models; the true stop at the secondary cap then fires per-request with degraded: false and stage: "backstop". Payloads without degraded/stage are the unchanged manual-tenant / token / gateway hard stops. |
budget_threshold |
Same fields as budget_exceeded, plus threshold_pct (the configured budget_alert_pct, e.g. 0.8). Fired once when spend crosses the soft threshold without being over the cap; the request is not blocked. |
SIEM (gateway-level override)
A gateway-level siem key overrides the tenant-level SIEM (Security Information and Event Management) config for that specific gateway. All fields are identical to the tenant-level config.
| Field | Type | Default | Description |
|---|---|---|---|
siem |
object | null | null |
SIEM backend config for this gateway. Overrides the tenant default. Set null to remove the override and fall back to the tenant config. |
siem.type |
string | — | Backend: splunk_hec, elasticsearch, vector, syslog. |
siem.events |
array | ["blocked"] |
Event filter. Values: blocked, guardrail, scrubbed, all. |
See SIEM Integration for the full field reference and per-backend examples.
Tracing
| Field | Type | Default | Description |
|---|---|---|---|
tracing |
object | null | null |
Set to enable request tracing. null disables tracing entirely. |
tracing.enabled |
boolean | false |
Activate internal pipeline tracing (Traces API + Playground traces). |
tracing.include_bodies |
boolean | false |
Store raw content in the request and web-search trace-step fields (request message arrays, leg-1 response previews, and web-search queries/URLs/page previews). When off, those fields are omitted and only metadata (counts, sizes, status) is stored. Enable only for debugging. |
tracing.otlp_endpoint |
string | — | Base URL of an OpenTelemetry collector (e.g. http://otel-collector:4318). Setting this enables OTLP span export. |
tracing.service_name |
string | "ai-gateway" |
service.name resource attribute on all emitted OTLP spans. |
tracing.headers |
object | {} |
Extra HTTP headers to include in the OTLP request (e.g. auth tokens for managed collectors). |
tracing.sample_rate |
number | 1.0 |
Fraction of requests to export via OTLP (0.0 = never, 1.0 = always). |
💡 Note: Trace retention is not a per-gateway field. Stored traces are purged automatically by a deployment-wide background job; the window is a deployment setting configured by your operator (default 48 hours). See Request Tracing — Trace retention.
💡 Note: Conversation (chat content) retention is separate. It is disabled by default and enabled per deployment by your operator (default off), then per tenant via
conversation_retention_days. See Conversation retention and deletion.💡 Note: Request-log retention (deletion of
request_logprompt/response rows) is separate again, disabled by default and enabled per deployment by your operator (default off), then per tenant viarequest_log_retention_days. Therequest_log_legscost ledger is never deleted. See Data retention — Request-log retention.💡 Note: Replay-record capture is separate again and default OFF. When a tenant sets
replay_capture_enabled, the gateway persists a raw per-turn replay record — the client messages, the intra-turn tool calls/results, the available tool set, and the sampling params — i.e. the full context needed to faithfully re-run a turn through the gateway with a different model (for offline model evaluation). It is bounded byreplay_capture_ttl_hours(default 48, reaped hourly) andreplay_capture_cap(default 200 per tenant+gateway), residency-tagged, deleted on Art. 17 user erasure (by both the turn owner and the run-as driver) and on tenant purge, and is not captured on guardrail-blocked turns. It is a raw (unmasked) store — enable it only where that retention is acceptable.
See Request Tracing for the full pipeline step reference and OTLP integration guide.
Image generation
| Field | Type | Default | Description |
|---|---|---|---|
image_generation |
object | null | — (absent) | Controls whether the EU image-generation tool (generate_image) is offered on this gateway. Absent by default — an admin- or API-created gateway has no key, and the tool is withheld until { "enabled": true } is set; self-serve signup is the one path that seeds it on. Set { "enabled": false } to withhold the tool (the admin UI's toggle persists null, which withholds it too). The tool is offered only when this is on and the request carries the x-aig-image-gen header (see the inference API reference). |
image_generation.enabled |
boolean | false |
Offer the generate_image tool to the model. |
image_generation.model |
string | — | Optional override of the EU-hosted image model the tool calls (see Image generation); absent → the deployment's default image model. |
Web search
| Field | Type | Default | Description |
|---|---|---|---|
web_search |
object | null | null |
Set to enable web-search augmentation. null disables the feature. |
web_search.enabled |
boolean | false |
Activate web search on this gateway. |
web_search.api_key |
string | — | API key for the configured web_search.provider (Brave Search or linkup.so). Required for all providers except Google Gemini. |
web_search.max_results |
integer | 5 |
Maximum number of search results to retrieve per query. |
web_search.mode |
string | "opt-in" |
"opt-in" — only triggered when the client sends the x-aig-web-search: 1 request header. "always" — attempted on every request. |
web_search.max_searches_per_turn |
integer | 8 |
Per-turn cap on distinct web searches on the model-driven tool-loop path (near-duplicates are de-duplicated). API-managed; read fail-safe (invalid/zero/non-finite → default). See Web Search settings. |
web_search.provider |
string | "brave" |
Search backend: "brave" or "linkup" (EU-sovereign). When the gateway sets no provider of its own, the tenant web_search_provider default fills it (and, under an armed EU residency floor, a non-EU per-gateway provider is clamped up to an EU tenant default). See Web Search — Supported providers and tenant web_search_provider. |
💡 Note: Web-search per-query cost is a DB catalog row, not a deployment setting. Migration
0226seeds amodel_pricerow forprovider='linkup', model='web_search'withinput_per_1k = 10.0— the per-1,000-queries rate (10.0=$0.01per query), so a web-search leg reportspricing_source=dband is visible and editable on the Provider Costs page (Settings › Costs) like every other model. In the UI that row is flagged per query (its Input field is USD per 1,000 queries, not per token) so an edit is not a 1000× mis-price. Only linkup is metered (Brave is the historical unmetered default and has no cost row). A built-in default rate (0.01) is now only the last-resort fallback used when the catalog row is absent (a fresh DB before the migration, or an operator DELETE). Propagation: the resolved price is memoised per worker for up to 1 hour (uniform across all model prices; the cache is node-local and not invalidated on edit), so a live UI edit takes effect within 1 hour or on the next worker restart; the migration restarts workers, so the seed is immediate. Overshoot bound: the budget cap is checked before a request starts, but search cost is recorded during the turn, so a single turn can overshoot its cap by at most 100 linkup queries (25 tool-loop rounds × 4 parallel calls, ≈$1.00 at the default rate) before the next request is blocked. This is bounded and accepted, not a bug — see Web Search for the full provider notes.
See Web Search for provider support details and usage examples.
Semantic cache
The semantic cache serves a stored response when a new request is close in meaning to an earlier one, even when the request text is not identical. The semantic cache is separate from the exact-match cache that cache_ttl controls and requires an external embedding service.
| Field | Type | Default | Description |
|---|---|---|---|
semantic_cache |
object | null | null |
Set to enable the semantic cache. null disables it. |
semantic_cache.enabled |
boolean | false |
Activate the semantic cache for this gateway. |
semantic_cache.embedding_url |
string | — | URL of the embedding service that produces the request vectors. Required; the semantic cache does nothing when this is unset. |
semantic_cache.embedding_api_key |
string | — | Optional bearer token for the embedding service. |
semantic_cache.embedding_model |
string | "text-embedding-3-small" |
Model the embedding service uses to produce request vectors. |
semantic_cache.max_candidates |
integer | 100 |
Maximum number of stored vectors compared against an incoming request. |
semantic_cache.threshold |
number | 0.95 |
Minimum cosine similarity for a cached entry to be served. |
semantic_cache.ttl |
integer | 86400 |
Lifetime of a semantic cache entry in seconds. |
Outbound PII egress guard
The egress guard scans a request's outbound tool traffic (web search, URL fetch, MCP tool arguments) for personal data and secrets before it leaves the gateway. It is on by default; any internal error in the guard fails closed (blocks). See Guardrails — threat model.
| Field | Type | Default | Description |
|---|---|---|---|
egress_guard |
object | null | on | Configures the outbound PII egress guard. Omit to keep the default-on behaviour; set enabled: false to disable it. |
egress_guard.enabled |
boolean | true |
false disables the guard for this gateway. |
egress_guard.mode |
string | "block" |
"block" refuses a request whose outbound tool traffic carries flagged data; "flag" logs the near-miss and allows it (for an open-web-browsing gateway). |
egress_guard.pii_sets |
array | high-signal set | The identifier/secret pattern sets to scan for. The default set omits loose numeric identifiers (SSN, phone, IP) that false-positive on ordinary URL IDs — opt those in explicitly. |
Agentic fetch
Agentic fetch lets the model fetch a URL and read its content during a conversation.
| Field | Type | Default | Description |
|---|---|---|---|
agentic_fetch |
object | null | null |
Set to enable agentic fetch. null disables it. |
agentic_fetch.enabled |
boolean | false |
Activate agentic fetch for this gateway. |
agentic_fetch.model |
string | — | Optional model for the fetch sub-request. When unset, the gateway uses the Myra fleet's canonical local chat model (qwen3.8-27b). A value naming a retired Myra-fleet id with a registered successor is upgraded to the successor at dispatch (logged; no response header — the inner leg does not serve the turn). See Retired model ids. Self-serve workspaces: the override must be in the plan's model list, and a Myra-hosted override needs the EU-Gov add-on — a PATCH setting a non-permitted value is refused (400 plan_model_not_allowed / eu_gov_model_not_allowed), and a stored non-permitted value runs on the default instead. The default itself is platform infrastructure and exempt from the EU-Gov rule. |
URL-fetch per-turn budgets
The fetch_url/agentic_fetch tools are bounded per turn. Both budgets are API-managed (no UI control in the Edit modal, which preserves a stored section untouched) and read fail-safe: an invalid, zero, negative, or non-finite value falls back to the default — a config typo can never uncap the tool or brick it.
| Field | Type | Default | Description |
|---|---|---|---|
url_fetch |
object | null | null |
Optional per-turn fetch budgets. null/absent means the defaults apply. |
url_fetch.max_fetches_per_turn |
integer | 24 |
Total URL fetches (web pages and documents) the model may make in ONE turn. On reaching it the model gets a notice naming the limit and answers from what it has. Raise it for research-heavy workloads. |
url_fetch.max_doc_fetches_per_turn |
integer | 8 |
Sub-cap on document extractions (PDF/DOCX/XLSX — the expensive, possibly-OCR unit) within that total. |
Upload malware scanning (ClamAV)
Scans uploaded file bytes for malware before they are processed or persisted, using a self-hosted ClamAV daemon (clamd). It is opt-in per gateway and OFF by default — malware scanning is never a global default. When a gateway opts in, its upload byte-ingress paths (chat file attachments, conversation attachments, anonymous form uploads) scan fail-closed: an infected file is rejected (422, code: malware_detected) and an unreachable/erroring scanner rejects rather than accepting unscanned bytes (503, code: scan_unavailable, with Retry-After). Tenant-scoped uploads that have no single gateway (project knowledge, personal My-PII lists, app-feedback evidence images — the last scanned on the reporter's tenant) are scanned when the tenant has scanning enabled on any of its gateways. That enablement read is itself fail-closed: if it cannot be determined whether scanning applies (the tenant's gateway list is unreadable, or a gateway config will not decode) the upload is refused with 503, code: scan_unavailable, not accepted unscanned. An account with no tenant at all (a platform admin's normal shape, or a tenantless user) owns no gateway, so no gateway can have opted in — such a caller's My-PII list / feedback image is knowably not scanned; a tenant id that is present but empty is a malformed row and refuses (503). The ClamAV endpoint is deployment infrastructure configured by your operator (a unix socket, or a host and port), not a config field; the daemon's StreamMaxLength must be at least the largest upload cap (100 MiB).
| Field | Type | Default | Description |
|---|---|---|---|
av_scan |
object | null | null |
Per-gateway upload malware scanning. null / absent = OFF. |
av_scan.enabled |
boolean | false |
When true, scan this gateway's upload byte-ingress fail-closed. Default OFF. |
Code interpreter
The code interpreter lets the model run a short program in a sandboxed, network-isolated interpreter and read back its output — for calculation, data analysis, and transforming data during a conversation. It works on the open / EU model stack (the sandbox runs on a self-hostable runner service, independent of any single provider).
The feature is ON by default on every gateway; the per-gateway flag is an opt-OUT. It is available when the deployment-wide runner is configured and the gateway has not explicitly disabled it:
- Deployment-wide runner (required) — your operator must point the gateway at a running sandbox control plane over HTTPS (a non-
https://URL is treated as a misconfiguration and keeps the feature dormant — code egress must be encrypted). When no runner is configured, the code interpreter is fully dormant: the tool is never offered, the system prompt is unchanged, and no request is ever made. - Per-gateway opt-OUT —
code_interpreter.enableddefaults totrue; set it tofalseon a gateway to disable the code interpreter there. An absent /nullvalue leaves it on (the default). Only an exact booleanfalsedisables it; a malformed value (non-booleanenabled, or a non-objectcode_interpreter) is rejected at the write boundary with400, so a fumbled opt-out can never silently leave it enabled.
Runner authentication (mTLS or bearer). The gateway authenticates to the sandbox with a client certificate (mTLS) or a bearer token; the sandbox may additionally enforce an IP allowlist. Configure at most one client-auth method:
| Deployment setting | Purpose |
|---|---|
| Runner URL | HTTPS base URL of the sandbox control plane. Required; a non-https URL keeps the feature dormant. |
| Client certificate (mTLS) | The PEM client certificate. |
| Client private key (mTLS) | The PEM client private key. |
| Bearer token | Optional bearer token, sent as Authorization: Bearer …. |
| Session signing key | Optional. Gateway-only secret for the persistent kernel (see Persistent kernel below). Unset ⇒ stateful sessions unavailable; stateless unaffected. |
- mTLS is all-or-nothing: set both cert and key, or neither. Setting only one, or a PEM that fails to parse, is a fail-closed misconfiguration — the tool is not advertised and no request is ever made (never an unauthenticated fall-through).
- The PEM files are read and parsed once per worker at first use (loaded at start-up); a parse failure is sticky until the service restarts.
- The gateway always verifies the sandbox's server certificate (TLS verification is never disabled). If the sandbox presents a private-CA certificate, add that CA to the gateway's trusted-certificate bundle; for an
IP:portURL the server certificate must carry a matching IP SAN.
In addition, the tool is never offered when the active provider does not support native tool-calling, when the client supplied its own tools, or for a Tier-1 local_only project (the runner is an external egress and model-generated code could carry local data).
| Field | Type | Default | Description |
|---|---|---|---|
code_interpreter |
object | null | null (→ on) |
Container for the code-interpreter opt-out. Absent / null leaves the feature on (the default). A present, non-null, non-object value is rejected with 400. |
code_interpreter.enabled |
boolean | true |
On by default. Set to false to disable the code interpreter on this gateway (opt-out). Absent / null → on. A non-boolean value is rejected with 400. Still requires a runner configured for the deployment. |
code_interpreter.stateful |
boolean | false |
Opt into the persistent kernel (variables survive across turns), layered under enabled. Requires the deployment's session signing key. Default off ⇒ stateless one-shot. See Persistent kernel below. |
Enabling it for a specific tenant/gateway
The code interpreter is on by default, so nothing has to be written to enable it — it is
available on every gateway once the deployment-wide runner is configured. To disable it on a
specific gateway, a tenant administrator PATCHes the gateway config —
PATCH /admin/v1/gateways/{id} with { "config": { "code_interpreter": { "enabled": false } } }.
The caller must be a tenant admin and the gateway must belong to their tenant (a plain
user gets 403; another tenant's gateway gets 403). The merge is shallow at the top
level: it sets code_interpreter and leaves every other config key intact — but it
replaces the whole code_interpreter object, so always send the full sub-object. A malformed
code_interpreter (a non-object, or a non-boolean enabled) is rejected with 400.
Two operational facts to plan around:
- Dormant until the runner env is set. The flag alone does nothing: unless the deployment-wide runner points at a running HTTPS sandbox control plane, the capability is fully inert (tool never offered, no request made). Enabling the flag on a deployment whose runner is not configured is a no-op until the runner is provisioned.
- Config-cache propagation. Gateway config is cached in a shared dictionary
(across all workers of the instance) for
config_cache_ttl(default 30 s). An admin PATCH of the gateway config invalidates that cache entry immediately, so enabling — and, more importantly, disabling — the flag on a gateway that is already serving traffic takes effect on the next request, not after the TTL. (A change written directly to the DB, bypassing the admin API, still propagates only within one cache TTL.)
PII / Datenschutz gateways — operator judgement, not a runtime block. Enabling the code
interpreter on a PII-active gateway is not blocked at runtime (the only runtime egress
block is a Tier-1 local_only project or a per-run no-egress profile; the PII axis is
independent and is not consulted for the code-interpreter offer). Model-generated code egresses
to the runner and can carry data the model saw in context — including PII it transcribes into
code literals. The input_files raw-attachment path is separately PII-gated fail-closed (see
below), but the code itself is not. So on a Datenschutz gateway this is an operator decision:
enable it only when the runner is EU-hosted / in-house (there is no code-enforced residency
tie between the runner and a gateway's eu_region_routing — the runner's
location is an operator obligation) and the residual code-literal egress is acceptable for that
gateway's data class.
Runner contract (trust boundary)
The runner is treated as an untrusted upstream. The gateway posts to POST <runner-url>/v1/execute. The request body is data-minimized — it carries only the program, optional stdin, and resource limits; never conversation history, PII, secrets, or tenant identity (the sandbox authenticates the gateway, not the end user):
{ "code": "<program>", "stdin": "<optional>",
"limits": { "wall_seconds": 30, "memory_mb": 512, "output_bytes": 65536 } }
The only request headers are Content-Type, an opaque x-aig-request-id correlation id, and the client-auth header (Authorization, when a bearer token is set). The encoded job is capped at 128 KiB; an over-cap program is rejected before egress with a clear "too large" error.
Data-file analysis (input_files)
The code_interpreter tool accepts an optional input_files argument — the filenames of data files (spreadsheets / CSV) the user attached to the conversation. When the model passes one, the gateway looks up that attachment's raw bytes (scoped to the conversation and its owner — no cross-conversation or cross-user access), validates it, and stages it into the sandbox at /input/<filename> so the program can read the real file directly (e.g. pandas.read_excel("/input/data.xlsx")). This lets any model compute exact figures over an uploaded spreadsheet instead of estimating from extracted text, and the raw file is analysed inside the self-hosted sandbox — it is never uploaded to the model provider.
- Accepted files: spreadsheet / tabular data only, by extension —
xlsx,xlsm,ods,csv,tsv(an.xlsxuploaded asapplication/octet-streamis accepted; OOXML/ODS files are magic-verified as zip containers). Legacy.xls, images, PDFs and other types are not staged. - Which version is staged (filename resolution): a name may match a chat attachment the user uploaded, a file in the conversation's project knowledge, or a file the model itself generated on an earlier turn (so a follow-up like "add a total row to that spreadsheet you made" can re-open and edit it). When more than one of these shares the same filename, the gateway stages the most-recently-created version — a later generated edit supersedes the upload it was derived from, and a genuine re-upload supersedes an older generated file. If that newest version cannot be loaded (mid-ingest, corrupt, or over a cap) the gateway falls back to the next-available same-named version and notes the substitution to the model rather than failing the run. Resolution is always scoped to the caller's own conversation/project — no cross-conversation, cross-user, or cross-tenant access.
- Bounds (fail-soft): ≤ 8 files, ≤ 8 MiB each, ≤ 16 MiB total. A file that is missing, unowned, the wrong type, corrupt, or over a cap is skipped with a note to the model — the rest of the run proceeds.
- PII gating (fail-closed): raw bytes may enter the in-house sandbox, but the analysis result returned to the model can contain PII (names, salaries) and, on an external-provider gateway, would then leave to the provider. So on a PII-active gateway, data-file staging is allowed only when the gateway has a fail-closed general-PII detector (
pii_protector) whose tool-result masking reliably masks that result before it egresses (the result passes the standard tool-result PII scrub, which reversibly tokenises PII for the provider and restores it in the final answer to the user). A PII-active gateway that masks only withcustom_pii(keyword-only) orpresidio(fails open on an outage outside a PII mandate) does not stage the raw file — the model still runs code, just without the attachment. A non-PII gateway stages freely. - Runner dependency: the gateway emits
input_filesin the/v1/executerequest as[{ "name": "<bare filename>", "b64": "<standard-base64 raw bytes>" }]; the sandbox runner must stage each into the guest/input/<name>(see the internal sandbox design,ops/sandbox/DESIGN.md§1.1). Data-minimization holds: only the explicitly-requested, owned attachment(s) are sent — never conversation history, other attachments, or other-tenant data.
The gateway accepts a JSON object response:
{ "stdout": "string", "stderr": "string", "exit_code": 0, "timed_out": false,
"truncated": false, "error": null,
"artifacts": [ { "mime_type": "image/png", "filename": "chart.png", "b64": "<standard-base64 PNG>" },
{ "mime_type": "text/csv", "filename": "results.csv", "b64": "<standard-base64 CSV>" } ] }
Accepted / rejected shape (all response fields optional, but type-strict when present; the gateway fails closed on anything malformed):
stdout/stderr— must be strings when present; a non-string is rejected. Each stream is scrubbed of invalid UTF-8 and capped (64 KB) before being returned to the model.exit_code— must be an integer when present (NaN, Infinity and non-integers are rejected). Absent → reported as “unknown”.timed_out— boolean; anything else is treated asfalse.truncated— boolean; when the sandbox trims a stream tooutput_bytes, the gateway propagates the truncation notice to the model.error— a broker-level rejection reason (e.g.no_isolation_available,busy,job_too_large). Any present, non-nullerrorfails the whole call closed — it is never rendered as an empty success, regardless ofexit_codeor HTTP status (a sandbox that could not isolate the job must not look like a job that ran and produced nothing).artifacts— optional array of files the program produced (a chart, or a data file such as a CSV or.xlsxof results). Each element carriesb64(standard-alphabet base64; base64url is rejected) plus advisorymime_type/filename. Every artifact is validated fail-closed and any that fails is dropped (its bytes never reach the user), not rendered. Three kinds are accepted; the runner's declaredmime_typeis never trusted for the decision:- PNG chart — accepted by magic bytes (filename-independent; a non-PNG labelled
image/pngis rejected), then stored and rendered inline in the conversation, downloadable, and included in chat export via the existing generated-image path. Also gated on pixel dimensions ≤ 25 MP (a small PNG declaring gigapixel dimensions — a browser decompression bomb — is rejected; the same ceiling the document exporter enforces). - Text data file — a
.csv/.txt(must be valid UTF-8 with no NUL byte) or.json(must parse as JSON). Thefilenamemust be a bare filename (a path-bearing or control-character name is rejected); its extension selects the kind and the gateway-pinned content type. Accepted files are delivered as a downloadable file card in the conversation (the same path thewrite_filetool uses); a duplicate name is auto-suffixed (results-2.csv) so it never overwrites an earlier file. Up to 4 data files are persisted per turn. .xlsxworkbook (binary) — a genuine Excel workbook the program produced (e.g.wb.save("report.xlsx")). Accepted by shape, not by the runner's declared type: the bytes must be a ZIP package (PK\x03\x04magic) carrying the[Content_Types].xmlpart and anxl/part (a.docx/.pptxOOXML zip, or a.xlsxname over non-ZIP bytes, is rejected). Thefilenamemust be a bare filename. The gateway stores and serves the bytes verbatim and never parses them (no server-side XXE / zip-bomb surface); it is delivered as a downloadable binary file card, served by an authz'd conversation download route. Shares the same per-artifact / total / count caps below.- Everything else is rejected: an unknown or absent extension, a non-PNG binary that is not a shape-valid
.xlsx, a.docx/.pptx(or other non-spreadsheet OOXML) mislabeled.xlsx, unparseable JSON, a.txt/.csvcontaining binary, an oversize file, or more than 4 files. - Shared caps across all artifacts: per-artifact size ≤ 4 MB, total ≤ 8 MB, count ≤ 4 (extras rejected).
- Degrades gracefully: a runner that returns no
artifactsfield (or only PNGs) is fully supported — the run still returns its text output; there is simply nothing extra to deliver. - The whole response body is size-bounded (12 MB; no OOM), and a non-
200status (including400/401/413/429/503/500), a connection error, or a timeout yields a friendly “temporarily unavailable” tool result — never a crash or a leak of the sandbox's internal status to the model.
The model-supplied code and stdin are size-bounded (64 KB each, and ≤ 128 KB combined job) before egress and are never executed by the gateway — they run only inside the runner sandbox. The rendered runner output is fenced as untrusted content, so stdout cannot forge the gateway's content markers or override instructions.
The gateway integration is default-dormant and fail-closed; deploying the sandbox runner service itself is a separate operational step.
Persistent kernel (stateful sessions)
By default every code-interpreter run is stateless — a fresh sandbox per call, with no variables carried between turns. A gateway may opt into a persistent kernel: a live per-user Python kernel whose namespace survives across turns (load a spreadsheet once, then reference the dataframe in a later turn without re-loading), evicted after ~2 hours idle. This is a phase-2 capability enabled for demo / internal gateways; stateless remains the default everywhere else and is unaffected.
Enabling it requires two things beyond the stateless prerequisites:
| Config / setting | Purpose |
|---|---|
code_interpreter.stateful (gateway config, boolean, default false) |
Per-gateway opt-in, layered under code_interpreter.enabled — a stateful-on gateway that isn't offering the tool at all can never mint a kernel. |
| Session signing key (deployment setting) | Gateway-only secret from which the opaque session token is derived. Unset → stateful sessions are silently unavailable (the gateway falls back to the stateless one-shot path); stateless is unaffected either way. |
When both are set and the request has an authenticated tenant, conversation, and user, the gateway mints a per-(tenant, conversation, user) session token and threads it to the runner; a missing principal falls back to stateless (never a mis-scoped kernel).
Isolation (per user, not per conversation). The token is "v1:" + HMAC-SHA256(key, "sess:v1:" + tenant + \x1f + conversation + \x1f + user). Because user is in the pre-image, a different participant in the same live-shared conversation derives a different token → a different kernel → they structurally cannot read each other's variables. The token is an opaque, unguessable map key — the runner never receives the tenant/conversation/user and does not (and cannot) verify the HMAC; the trust boundary remains mTLS + the client-CN pin exactly as for the stateless path. A tampered or guessed token simply misses the runner's session map and gets a fresh kernel — never someone else's.
Request additions (session mode). Alongside the stateless {code, stdin, limits} body, a session request carries session (the token) and, on a kernel-creating run, session_tags (three opaque HMAC correlators — conv_tag, user_tag, tenant_tag) the runner records so an erasure can destroy kernels by conversation / user / tenant. These are HMACs, never raw ids — the only widening of the data-minimized wire, EU-resident and mTLS-only. A request with no session is byte-for-byte the stateless body.
Response addition. The runner may return kernel_reset: true when it spawned a fresh kernel for this run (first use, idle-evicted, or a wall-clock timeout wiped the namespace — a timeout SIGKILLs the kernel, so any timeout loses all persistent state). The gateway surfaces a one-line honesty note to the model ("variables from earlier turns are not available; re-create what you need"). A stateless or older runner omits the field → treated as no reset.
GDPR teardown (Art. 17). Kernels are ephemeral (idle-evicted, never persisted to disk beyond the VM's discarded overlay), but the gateway also actively destroys them on the three erasure events by calling POST <runner>/v1/session/destroy with { "scope": "conversation" | "user" | "tenant", "tag": "<the matching HMAC tag>" } (same mTLS + CN pin). The tag is computed from the erasure event's tenant/conversation/user via the same derivation, so no gateway-side session tracking is needed and a cross-tenant destroy is impossible (every tag embeds the tenant in its pre-image). The call is best-effort / fail-soft: a runner-unreachable teardown is logged loudly for the retention audit but never blocks the database erasure.
Accepted / rejected (session inputs, at the runner boundary): the runner validates the token shape only (v1: + 64 lowercase hex) fail-closed — a malformed token is a 400; a well-formed but unknown token misses the map and spawns a fresh kernel. The destroy body is validated fail-closed: scope must be one of session/conversation/user/tenant, and tag must be a token (session scope) or a 64-hex tag (the others); anything else is a 400. Capacity is bounded (a small ceiling of live kernels sized to the runner's memory, plus per-tenant and per-user caps); exhaustion returns a clean "temporarily unavailable", never a queue or crash.
The side-channel residual of co-resident per-user micro-VMs on a shared host is a documented, accepted risk for the demo / internal gate (the feature is not enabled for production PII tenants without dedicated-host / SMT-off mitigation) — see ops/sandbox/DESIGN.md and docs/internal/code-exec-sandbox.md.
Context compaction
Context compaction automatically summarises older conversation turns when the estimated input token count exceeds a configured threshold. This keeps long conversations within model context limits while reducing the cost of each subsequent turn.
Context compaction is available for Anthropic models that support the native compact_20260112 strategy only. For other providers — and for Anthropic models without that strategy (for example claude-haiku-4-5, claude-sonnet-4-5, claude-opus-4-5) — the setting is present in the config but injects nothing.
| Field | Type | Default | Description |
|---|---|---|---|
context_compaction |
object | {enabled: true, threshold_tokens: 200000, keep_last_turns: 10} |
Context compaction settings. When the key is absent (or is not an object), the default applies. Set enabled: false — or the key to JSON null — to disable. |
context_compaction.enabled |
boolean | true |
Activate automatic context compaction for this gateway. Strictly boolean: only true enables; any other value (1, "true", null) leaves compaction off. |
context_compaction.threshold_tokens |
integer | 200000 |
Trigger compaction when the estimated input token count reaches this value. Minimum: 50000 (Anthropic API requirement). |
context_compaction.keep_last_turns |
integer | 10 |
Retained for compatibility. Compaction is performed by the native context management of the Anthropic API, which controls how many recent turns are kept. The gateway does not enforce this value separately. |
See Context compaction on the Anthropic provider page for background and behavioural details.
Compaction retry threshold
compact_error_threshold (top-level integer, default unset) configures the consecutive provider-error count after which the gateway forces a context-compaction probe. Set to a positive integer to enable; leave unset or set to null to disable.
Prompt caching
Anthropic prompt caching writes frequently-used parts of the prompt (system instructions, large context blocks) to a cache and bills cache reads at a steep discount. It is on by default; configure per gateway:
| Field | Type | Default | Description |
|---|---|---|---|
prompt_caching |
object | {enabled: true, ttl: "1h"} |
Prompt-caching settings. When the key is absent (or is not an object), prompt caching is enabled by default. Set the key to JSON null to disable it. |
prompt_caching.enabled |
boolean | true |
Activate prompt caching for Anthropic requests. Send a real boolean: only false (or an omitted field) disables; any other value — including the string "false" — is read as enabled. |
prompt_caching.ttl |
string | "1h" |
Cache TTL for written entries. One of "5m" or "1h". The "1h" default applies when the whole prompt_caching key is absent; an explicit prompt_caching object that omits ttl (or carries an unrecognised value) caches with Anthropic's "5m" default. |
Per-request header overrides
These headers can be sent on individual inference requests to override gateway config for that request only.
| Header | Type | Description |
|---|---|---|
x-aig-byok-alias |
string | Use a non-default BYOK provider key alias for this request. Must match an alias stored for the resolved provider. |
x-aig-meta-{key} |
string | Attach a custom key-value pair to the request log entry and make it available in routing rule conditions as meta:{key} (colon notation) and load-balancer sticky fields as meta.{key}. See the caps below. |
x-aig-collect-log |
"0", "false", or "1" |
"0" or "false" = skip writing this request to the log table entirely. "1" = log (default). |
x-aig-collect-log-payload |
"0", "false", or "1" |
"0" or "false" = log request metadata but omit the prompt and response body. "1" = log body (default). Does not affect the gateway-level log_payloads setting. |
x-aig-provider-{field} |
string | Strip the x-aig-provider- prefix and forward the header to the upstream provider only if the stripped name is allow-listed (anthropic-beta, openai-organization, openai-project). Any other stripped name — a credential, a request-framing / hop-by-hop header, or an unrecognised name — is dropped and logged at warn, so a client can neither inject an upstream credential nor tamper with request framing. At most 16 overrides / 8 KB total per request. Useful for provider-specific beta flags. |
💡 Note: An allow-listed
x-aig-provider-*header is forwarded to whatever provider handles the request, not just the one the name belongs to. Sendingopenai-organizationto an Anthropic request is harmless (the provider ignores the unknown header). Only the three allow-listed names are ever forwarded.🔒
x-aig-meta-*accepted shape and caps. Client metadata is sanitized at the trust boundary before it becomes a routing input or is persisted to the request log: - Value — a single string. If the samex-aig-meta-{key}header is sent more than once, only the first value is used; a value longer than 256 characters is truncated on a UTF-8 character boundary (invalid byte sequences are scrubbed). - Key — the{key}suffix must be at most 64 characters; a longer key is rejected (dropped, not truncated, so it can never collide with another key or a routing rule). - Count — at most 16 client metadata keys are kept per request. When more are sent, a deterministic subset (the keys sorted lexicographically, first 16) is kept so the same request always routes identically. - Reserved namespace — keys beginningaig_/aig-are gateway-owned and are dropped if sent by a client.Over-cap or malformed metadata is dropped or clamped (never a 4xx), and the request log row is always written — an oversized
metaobject is bounded so it can never overflow the column and lose the row.
Knowledge-search reranking
Project knowledge search (hybrid dense + keyword retrieval) can optionally run a final cross-encoder rerank stage that re-scores the fused candidate passages for relevance to the query and keeps the most relevant few. It is a single deployment-wide setting configured by your operator — there is no per-gateway field:
| Deployment setting | Default | Effect |
|---|---|---|
| Reranker model | (unset) | When unset (the default), the rerank stage is dormant and knowledge search returns exactly the hybrid-fused order — no rerank call, no behaviour change. When set to a reranker model id available on the EU model fleet (e.g. bge-reranker-v2-m3), retrieved passages are reranked and the top few kept. |
The reranker response is treated as untrusted: if the reranker is unavailable, times out, or returns a malformed/incomplete response, retrieval fails closed to the pre-rerank hybrid order — it never blanks or drops results. Enabling it requires the reranker model to be deployed on the fleet first (an operations dependency).
💡 Note: The reranker model is a deployment setting configured by your operator. Leaving it unset keeps reranking off.
Config and BYOK cache TTLs
These deployment-wide settings tune how long the gateway caches resolved config and decrypted BYOK keys. Both apply per instance, not per gateway.
| Deployment setting | Default | Effect |
|---|---|---|
| Config cache TTL | 30 (seconds) |
How long a gateway's resolved config is cached in the shared dictionary across all workers of the instance. An admin PATCH of the config invalidates the entry immediately; a change written directly to the database propagates only within this TTL. |
| BYOK cache TTL | 60 (seconds) |
How long a decrypted BYOK (provider key vault) key is held in the in-memory cache after decryption, so a burst of requests does not re-decrypt the stored key on every call. |
Applying config changes
# Enable caching and set a budget
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
-H "Content-Type: application/json" \
-d '{
"config": {
"cache_ttl": 300,
"budget_usd": 200.00
}
}'
# Disable authentication (development only)
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
-H "Content-Type: application/json" \
-d '{"config": {"auth_required": false}}'
# Add an IP allowlist
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
-H "Content-Type: application/json" \
-d '{"config": {"ip_allowlist": ["10.0.0.0/8"]}}'