Skip to content

Gateway configuration reference

Every gateway has a config JSON object that controls authentication, caching, timeouts, security, routing, and provider settings. Update it with PATCH /admin/v1/gateways/{id}.

The config is merged at the top level on each PATCH — only the fields you include are changed. Nested objects (rate_limit) are replaced in full when provided.

🔑 Reading the config back — credential fields are role-scoped. GET /admin/v1/gateways/{id} and the tenant gateways listing are open to any member of the tenant, so the third-party credential fields below — web_search.api_key, semantic_cache.embedding_api_key, webhooks.secret, the whole siem block, and tracing.headers — are returned only to a caller holding GATEWAYS_MANAGE; a lower-privileged reader gets them withheld. See the read-scoping note in the gateways API reference.


Full default config

{
  "auth_required": true,
  "budget_usd": null,
  "budget_period": "monthly",
  "tenant_budget_usd": null,
  "tenant_budget_period": "monthly",
  "cache_ttl": 0,
  "retry_count": 2,
  "timeout_ms": 60000,
  "log_payloads": true,
  "rate_limit": null,
  "ip_allowlist": [],
  "guardrails": [],
  "circuit_breaker": null,
  "webhooks": null,
  "siem": null,
  "azure_endpoint": null,
  "azure_deployment": null,
  "azure_api_version": "2024-02-01",
  "azure_resource": null,
  "bedrock_region": "us-east-1",
  "vertex_project": null,
  "vertex_region": "us-central1",
  "cf_account_id": null,
  "hf_endpoint": null,
  "provider_base_urls": {},
  "provider_allowlist_enforced": false,
  "provider_allowlist": [],
  "tracing": null,
  "web_search": null,
  "semantic_cache": null,
  "agentic_fetch": null,
  "url_fetch": null,
  "av_scan": null,
  "prompt_caching": { "enabled": true, "ttl": "1h" },
  "context_compaction": {
    "enabled": true,
    "threshold_tokens": 200000,
    "keep_last_turns": 10
  }
}

Core fields

Field Type Default Description
auth_required boolean true Require a valid x-aig-token, Authorization: Bearer, or x-api-key header on all inference requests. Set to false only for development.
budget_usd number | null null Gateway-level spend cap in USD for the current budget period. Blocks all requests once exhausted. null = no cap.
budget_period string "monthly" Period over which gateway spend is accumulated. One of: "daily" (resets each day at midnight in the server's local time zone), "monthly" (resets on the first of each calendar month, local time), "total" (lifetime, never resets).
test_headers_allowed boolean absent Reserved for Myra's test tooling; never set on a customer gateway. Platform-admin-only: on a config write by any other actor (a tenant admin, or a custom role holding GATEWAYS_MANAGE) the submitted value is ignored and the stored value kept — on PATCH and on a POST to an existing slug alike, so no one below a platform admin can arm or disarm a gateway — and the 2xx response then carries "ignored_fields": ["test_headers_allowed"] when the dropped value differed from the stored one (omitting the field, or re-sending the stored value, is a plain no-op). A platform admin's value is applied as sent; a platform admin's non-boolean value that differs from the stored one is rejected 400 ("test_headers_allowed must be a boolean").
tenant_budget_usd number | null null Tenant-level spend cap in USD. Applies across all gateways belonging to the tenant. null = no cap.
tenant_budget_period string "monthly" Period for the tenant-level budget. One of: "daily", "monthly", "total".
budget_alert_pct number | null 0.8 Soft-alert threshold as a fraction of the budget cap: when current-period spend crosses this fraction of any configured budget (token / tenant / gateway) without being over it, a proactive budget_threshold signal fires once — a budget_threshold webhook (if webhooks is configured) plus an ops notification (log + operator alert e-mail + alert webhook + Mattermost). This is advisory only and never blocks the request; the hard stop still happens at 100%. Must be a number in the open interval (0, 1); 0 disables the soft alert; null, absent, or any out-of-range/non-number value falls back to 0.8 (fail-safe read). The alert fires at most once per (scope, entity, period, cap); a new period or a changed budget re-arms it. Note: the SPA's "Budget Warnings" badge is fixed at 80% and is independent of this knob (it changes only the webhook/notification timing, not the badge).
pii_masking_enforced boolean false Tenant-level. When true, PII masking is mandatory for every user of the tenant: the chat privacy toggle is locked, the client hint_pii_preference opt-out is ignored, routing is forced through a PII-masking gateway (or the request is rejected fail-closed), manual unmask is denied, and every third-party tool egress (web search, fetch_url, MCP arguments) is scrubbed or blocked. A wholly-local (Myra/EU) model leg — the resolved primary and every failover fallback first-party — is never masked on the model input (masking a first-party leg protects nothing and corrupts the user's own data); its tool egress is still guarded, and agent-invoke output (delivered externally) always masks. If any provider in the egress set is external, the leg is masked fail-closed. Set on the tenant, not per-gateway. See PII Protector → tenant-enforced masking.
cache_ttl integer 0 Response cache TTL in seconds. 0 disables the cache. Cached responses are keyed on SHA-256(provider:model:canonical_body).
retry_count integer 2 Maximum number of retry attempts against the primary provider on a retryable status (500, 502, 503, 504, or 429 — the last honouring the Retry-After / Retry-After-Ms header) before the fallback chain is walked.
timeout_ms integer 60000 Per-upstream-request timeout in milliseconds. Applies to each attempt individually, not the total request time. Bounds the connection, the request send, and a buffered (non-streamed) read; a streamed response's read phase is bounded by the built-in stall budgets (20 s to the first byte, 300 s per inter-chunk gap) instead — to bound a streamed read explicitly, set timeout_ms on a routing rule (see the routing rules API). The admin UI pre-fills 120000 when a gateway is created; the 60000 default applies only when the key is absent.
log_payloads boolean true Store request and response bodies in the log table. Disable for sensitive workloads where prompt/response content must not be persisted.
max_parallel_tools integer — Maximum number of tool calls the gateway runs in parallel within one assistant turn. Raises the built-in limit. A value of 0 or below defers every tool call to a separate turn.

Rate limiting

Field Type Default Description
rate_limit object | null null (disabled) Gateway-level sliding-window rate limit applied to all callers. Default: null (disabled). Example: {"requests": 100, "window_sec": 60}. Per-token limits are checked independently — a request can be blocked by either limit. Accepted shape (validated on POST /gateways and PATCH /gateways/{id}): a JSON object whose only keys are requests and window_sec, or null to clear. Rejected with 400: a string, number, boolean or array; an unknown key (a typo such as reqs used to be stored and silently mean 100 per 60 s); a requests / window_sec that is not a whole number ≥ 1 (a string "100", 0, a negative or fractional value, nan/inf) or above the ceilings (requests ≤ 1,000,000,000; window_sec ≤ 31,536,000 — one year). A key sent as null is treated as omitted and is not stored. At request time, a stored value the gateway cannot use — not an object (a string, number, boolean or array), or a present requests / window_sec that is not a whole number in range — makes the gateway refuse every request with 500 configuration_error and log the reason once per worker (until the value is corrected) — never "no limit" and never a crash per request. A row written before this validation existed is otherwise served as-is: extra keys are ignored at read, a null key takes the default.
rate_limit.requests integer 100 Maximum requests allowed in the window. Used when rate_limit is set but this field is omitted or null.
rate_limit.window_sec integer 60 Window duration in seconds. Used when rate_limit is set but this field is omitted or null.

⭐ Example:

"rate_limit": {"requests": 100, "window_sec": 60}

IP allowlist

Field Type Default Description
ip_allowlist array of strings [] CIDR blocks permitted to call this gateway. An empty array allows all source IPs. Requests from IPs outside the list return 403 forbidden.

⭐ Example:

"ip_allowlist": ["10.0.0.0/8", "192.168.1.0/24"]

Guardrails

Field Type Default Description
guardrails array [] Ordered list of guardrail configs. Evaluated in array order within each tier. First block verdict short-circuits the pipeline.

💡 Note: The legacy key detectors is still accepted as a fallback, but only when guardrails is absent (the key is missing) or the boolean false. A present-but-null or otherwise non-array guardrails value takes precedence and resolves to zero detectors (it does not fall back to detectors). New configurations should use guardrails.

Each guardrail object has a common set of fields plus type-specific fields:

Field Type Description
type string Guardrail type: regex, keyword, jailbreak, json_schema, contains_code, gibberish, language, custom_pii, presidio, prompt_guard, pii_protector.
name string Human-readable name used in block messages and logs.
action string One of block, scrub, or flag.
target string One of request, response, or both.

See the Guardrail Pipeline page for full per-type field documentation.


Provider-specific fields

Azure OpenAI

Field Type Default Description
azure_endpoint string | null null Azure OpenAI resource URL, e.g. https://myresource.openai.azure.com. Required for Azure provider. Routed through the same SSRF guard as provider_base_urls: it is resolved once and the dial IP is pinned; a value resolving to a private/internal address is rejected at dispatch.
azure_resource string | null null Azure resource name. An alternative to azure_endpoint: the gateway builds the URL https://<resource>.openai.azure.com/… from the resource name and azure_deployment. Accepted: a bare vendor label (1–63 chars, alphanumeric plus -/_, starting alphanumeric). Rejected: any value with a dot, slash, colon, @, or whitespace — these would break out of the host authority and are refused (the request fails with a configuration error).
azure_deployment string | null null Azure deployment name. Replaces the model name in the request URL path. Accepted: a single path segment (may contain dots/-/_). Rejected: any value with a slash, ?, #, :, @, or whitespace (path/query injection).
azure_api_version string path-dependent Azure OpenAI API version appended as the ?api-version= query param. Defaults to "2024-02-01" on the OpenAI-provider path (a gateway with azure_endpoint set) and to "2024-10-21" on the dedicated azure provider path.

AWS Bedrock

Field Type Default Description
bedrock_region string "us-east-1" AWS region for Bedrock API calls. Used in SigV4 request signing.

Google Vertex AI

Field Type Default Description
vertex_project string | null null Google Cloud project ID. Required for Vertex AI; if unset, falls back to a deployment-wide default project set by your operator. If neither is set the request fails closed with a configuration_error (no empty-project URL is dialed).
vertex_region string "us-central1" Google Cloud region for Vertex AI API calls. Falls back to a deployment-wide default region when unset.

Data residency (EU model hosting)

For deployments where model data must stay in the EU (a hard requirement in regulated / public-sector tenders), a gateway can be put into EU-region routing mode. When enabled the gateway fails closed: any request whose resolved provider/region is not EU-hosted is rejected with HTTP 403 (data_residency_blocked) — it is never sent to a US endpoint.

Field Type Default Description
eu_region_routing boolean false When true, only EU-hosted providers may serve this gateway. Coerced to a strict boolean at config load — a non-boolean value is treated as false (disabled) and logged.
azure_region string | null null Operator-asserted EU region of the Azure OpenAI resource (e.g. swedencentral). The Azure host does not encode a region, so this declaration is what the residency guard validates.

What counts as EU-hosted (fail-closed allowlist):

  • AWS Bedrock — bedrock_region is an EU-member region: eu-central-1, eu-west-1, eu-west-3, eu-north-1, eu-south-1, eu-south-2. Only a bare single-region on-demand model id (e.g. anthropic.claude-3-haiku-20240307-v1:0) is vouched under enforcement — it pins processing to that one region. AWS cross-region inference profiles (eu.anthropic.*, us.anthropic.*, global.anthropic.*) are refused: a geographic profile has a multi-region destination set that is not provably EU-27 — AWS's own "EU" geography for Claude routes to eu-central-2 (Zurich, Switzerland), which is outside strict EU-27. (Admitting eu. profiles would require widening residency scope to EU + adequacy incl. CH — a deliberate policy change to the EU-region allowlist, not enabled by default.)
  • Google Vertex — vertex_region is an EU-member region: europe-west1, europe-west3, europe-west4, europe-west8, europe-west9, europe-west10, europe-west12, europe-north1, europe-north2, europe-central2, europe-southwest1.
  • Azure OpenAI — azure_region is an EU-member region: westeurope, northeurope, germanywestcentral, germanynorth, swedencentral, francecentral, francesouth, polandcentral, italynorth, spaincentral.
  • Myra-hosted models (myra / vllm) — run on the on-prem EU estate.
  • Mistral AI (mistral) — the gateway calls only Mistral's EU-hosted endpoint (api.mistral.ai, EU by default); it never enables Mistral's opt-in US endpoint. A fixed EU-direct provider with no region to configure, so it is vouched like Myra. (Requires a Mistral API key configured for the gateway.)

Everything else — the native anthropic / openai / gemini adapters (US endpoints), any US SaaS provider, custom providers, and UK / Switzerland / Norway regions (not EU-member — EEA is not in scope) — is rejected. A per-provider provider_base_urls override is also rejected under enforcement (the gateway cannot vouch for an arbitrary host).

The guarantee holds on every model dispatch path, not just the primary request. The check sits at the single upstream dispatch chokepoint, so the same fail-closed refusal applies to the primary attempt and every failover attempt, each tool-loop leg, the conversation summarize call, and the agentic-fetch inner call — all of which funnel through that one chokepoint. It also covers the web-search query egress (a non-EU search provider such as Brave is blocked; only an EU provider such as Linkup may receive the model-generated query, and provider-native search is forced through the EU two-leg), so an armed gateway can turn web search off but never silently fall back to a US search backend. Embeddings and image generation are served solely by the on-prem Myra EU fleet (MYRA_BASE_URL) and never reach a commercial US provider. Net effect on an armed gateway: no user request — on any path — reaches a US region; a US route is refused before any network call rather than egressing.

Routing the big three via EU adapters. Point the model at the EU cloud adapter (routing rule / catalog target) and set the EU region:

# Anthropic → AWS Bedrock (Frankfurt); OpenAI → Azure OpenAI (Sweden);
# Google → Vertex (Netherlands).
curl -X PATCH "https://<host>/admin/v1/gateways/<ID>" -H "Content-Type: application/json" -d '{
  "config": {
    "eu_region_routing": true,
    "bedrock_region": "eu-central-1",
    "vertex_region": "europe-west4", "vertex_project": "<proj>",
    "azure_resource": "<eu-resource>", "azure_deployment": "<deployment>", "azure_region": "swedencentral"
  }
}'

A deployment-wide default is available as an operator-only setting, applied on the sovereign/EU stack. It is not applied on a shared stack, because it would flip every gateway to enforced and 403 existing US-routed gateways.

EU-hosted model names & regions

Three vendor-independent EU-27 model routes are available:

Route Provider EU-27 models / regions
Myra-Eigenbetrieb myra / vllm On-prem EU inference fleet (chat, embeddings, image generation).
Mistral AI mistral mistral-large-latest, mistral-small-latest, codestral-latest, … — EU host api.mistral.ai.
Google Gemini via Vertex vertex Gemini 2.5 Pro / 2.5 Flash / 2.0 Flash / 1.5 Pro / 1.5 Flash on a europe-west* regional endpoint (europe-west4 Netherlands, europe-west1 Belgium, europe-west3 Frankfurt, …). A regional endpoint pins processing to that single EU-27 region — the usable big-vendor EU-27 route.
Anthropic Claude via Bedrock bedrock Bare on-demand Claude ids (e.g. anthropic.claude-3-haiku-20240307-v1:0) in an EU-27 bedrock_region. Cross-region inference profiles (eu.anthropic.*) are refused under strict enforcement (destinations include Zurich/CH) — see the AWS Bedrock note above.

Newer Claude models are offered by AWS only via cross-region inference profiles, so under strict EU-27 they are not available on Bedrock; Vertex Gemini is the delivered big-vendor EU-27 route. The full A-5 answer (incl. the EU-vs-adequacy policy decision) is maintained by Myra.

Region resolution — the precedence chain

The effective region for a region-bearing provider (Bedrock / Vertex) is resolved from an ordered candidate list; the first candidate that is a valid region token wins:

  1. the gateway's own region (bedrock_region / vertex_region in the gateway config);
  2. the tenant's region (bedrock_region / vertex_region on the tenant — set in the tenant settings, see below), inherited by every gateway under that tenant that does not set its own;
  3. the deployment-wide default region set by your operator — now a fallback tier, not the primary source;
  4. the provider's hard-coded US default (us-east-1 / us-central1) — logged with a WARN so a gateway with no region configured anywhere is visible.

The fall-through is explicit, not OR-coalescing: a non-empty but invalid earlier candidate (a whitespace / uppercase / malformed paste) is treated as "not set" and falls through to the next tier — it can never short-circuit past a valid tenant or env region to the US default.

Per-tenant region SELECT. A tenant sets its Bedrock/Vertex region in the tenant settings (a dropdown of EU regions, sourced from GET /admin/v1/cloud-regions). The value is validated server-side as an EU-member region (region.is_eu, fail closed) — a non-EU or malformed value is rejected with HTTP 400. Clearing it (empty selection / null) reverts to the deployment or hard-coded default. This is a tenant-column change, folded into every gateway's cached config, so it propagates within the gateway-config cache TTL (≤ 30 s) — unlike a per-gateway admin write, which is invalidated immediately. Allow for the TTL when verifying a tenant region change.

Region value vs the EU guarantee — two distinct controls. The precedence above decides the region value; the EU guarantee is the separate eu_region_routing enforcement floor. A per-gateway region overrides the tenant region — including to a non-EU value — and when enforcement is OFF that value is used as-is (the region token is still sanitized against header/URL injection, but not EU-checked). To guarantee EU-only routing, enable eu_region_routing (per gateway or the tenant floor): any resolved non-EU region is then blocked at dispatch (403 data_residency_blocked), fail closed. Setting a tenant EU region alone is a default, not a floor.

Provider allowlist (no silent fallback)

Independent of data residency, a gateway can pin an explicit set of approved providers so it can never silently route — primary or fallback — to a provider the operator did not approve. Enforced at the same single dispatch chokepoint as residency: a request whose resolved provider is not on the list is rejected with HTTP 403 (provider_not_allowed) before any upstream call.

Field Type Default Description
provider_allowlist_enforced boolean false When true, only providers named in provider_allowlist may serve this gateway. Coerced to a strict boolean at config load — a non-boolean value is treated as false (disabled) and logged.
provider_allowlist array of strings [] Approved provider names (e.g. ["myra", "mistral"]). Matched case-insensitively after trimming; vllm folds to myra. Under an armed tenant default this per-gateway list may only narrow the tenant-approved baseline (it is intersected with it) — a gateway can never add a provider outside the tenant list.

Fail-closed semantics. When enforcement is on, a missing, empty ([]), or malformed (non-array) allowlist denies every provider — a forgotten or mistyped list can never silently fall back to allow-all. Non-string array elements (numbers, null, nested objects) are ignored and can never widen the set; a near-miss or substring of an approved name never matches. When enforcement is off the list is not consulted (zero behaviour change for existing gateways).

Scope. The allowlist governs routed chat/inference dispatch. Fixed EU-infrastructure services that call the Myra on-prem fleet directly — embeddings (RAG), image generation/editing, transcription, and vision image-analysis — always use that sovereign fleet and are not routable providers, so they are not gated by this list.

A deployment-wide backstop is available as an operator-only setting, applied on the sovereign stack: with it applied, every gateway must carry a valid allowlist or its inference fails closed. It is not applied on a shared stack, because it would 403 every gateway that has not configured an allowlist. An unrecognised value is treated as disabled and logged at startup.

Per-tenant sub-processor objection (Art. 28 Abs. 6 deactivation)

Orthogonal to residency and the allowlist, a per-tenant sub-processor objection narrows the effective provider set for one tenant without touching any gateway config or any other tenant. When Myra Ops records an objection (see the objection API), the provider(s) mapped to that sub-processor are deactivated for that tenant: they are dropped from the tenant's model-offer set and refused at the single dispatch chokepoint (primary, fallback, failover, and the observed LiteLLM-overflow leg) with 403 provider_objected — the same posture as the allowlist deny, but scoped by tenant_id and driven by the objection store rather than a config flag. The deactivated set is derived at config load (a {provider: true} set on the resolved gateway config, never stored in gateway.config, so a tenant admin cannot PATCH away their own objection), and it fails closed: a hard read error on the objection store fails the whole config load rather than serving the objected provider. It is enforcement-flag-free — presence of an objection is itself the enforcement, so it works on an ordinary non-sovereign gateway too.

Tenant-level residency default (default ON for a sovereign tenant)

For a sovereign tenant on a shared stack (where the deployment-wide backstops above are not an option — they would flip every tenant), the three residency controls above (eu_region_routing, provider_allowlist_enforced, provider_allowlist) can be set as a tenant-level default. Every gateway under that tenant then inherits the default — including a newly created gateway with an empty config, which is the behaviour a public-sector "commercial models exclusively via EU regions" obligation requires: enforcement must be default ON, never fail-open on a forgotten per-gateway flag.

  • Where it lives: tenant columns tenant.eu_region_routing, tenant.provider_allowlist_enforced, tenant.provider_allowlist. eu_region_routing is settable in the product: Settings → Organisation → Routing → Organisation-wide EU routing (platform admins only — lowering an organisation's residency floor is not self-service), and on PATCH /admin/v1/tenants/{id} as eu_region_routing. It is tri-state: 1 armed, 0 explicitly not armed, null (the empty option) = no organisation default at all, each gateway governs itself. provider_allowlist_enforced and provider_allowlist remain operator/DB-managed only (a migration or a db.sh UPDATE) — there is no admin-UI / PATCH field for those two.
  • Precedence — a FLOOR (per enforcement control): when a tenant enables a residency control (eu_region_routing / provider_allowlist_enforced), that is a floor — a gateway can only raise the control, never lower it. A gateway that sets eu_region_routing to false/0/"off" under a EU-only tenant is still enforced (it cannot opt a sovereign tenant out of EU-only routing — that would be fail-open on the "commercial models exclusively via EU regions" exclusion criterion). The floor covers the two enforcement booleans and the provider_allowlist list contents: when the tenant arms provider_allowlist_enforced with a baseline list, a gateway that supplies its own provider_allowlist may only narrow it — its entries are intersected with the tenant-approved list, so a gateway can never approve a provider the tenant did not (a gateway list of ["openai"] under a tenant baseline of ["myra","mistral"] resolves to [] — deny all — because openai is not tenant-approved). A gateway with no own list inherits the baseline. For a fully-sovereign posture arm both eu_region_routing and provider_allowlist_enforced (as the bid tenant does) — the residency guard is then the hard EU floor regardless of any gateway's allowlist. A gateway's own enforcement-flag value governs only when the tenant default is a recognized tenant default is a recognized OFF (0/false) — the ordinary per-gateway opt-in for non-sovereign tenants. An unrecognized/garbage tenant value (e.g. 2) is a mis-provisioned compliance flag and fails closed to ENFORCED (with a warning), never a silent US-leak. For a residency control an absent or null per-gateway value resolves to the tenant default (the safe ON under a floor). The deployment-wide operator backstop still forces enforcement on where set.
  • Fail-closed: a tenant that enforces the allowlist but whose tenant.provider_allowlist is missing/malformed denies every provider on its inheriting gateways (same fail-closed rule as the per-gateway list).
  • Accepted shape — enforced by the database: tenant.provider_allowlist must be SQL NULL or a valid JSON array (e.g. ["myra","mistral"]). A DB CHECK constraint (migration 0190) rejects any other write — invalid JSON (a bare-token [myra,mistral] from a mis-quoted provisioning statement), a JSON object, a scalar, or a literal null — with CONSTRAINT chk_tenant_provider_allowlist_json failed, so a malformed value can no longer be stored at all. Should a pre-constraint/foreign row still carry a non-array value, the gateway treats it as absent (fail-closed: deny-all while enforced) and logs it at ERROR level naming the tenant slug — never silently.
  • Admin view caveat: the admin gateway GET/PATCH shows the raw per-gateway config, so an inheriting gateway shows these keys absent while enforcement is active at runtime. Read the tenant default to know the effective policy.
  • Cache: the per-gateway config is cached briefly (≤ 30 s). A per-gateway change through the admin API invalidates the cache entry immediately; flipping a tenant default (a tenant-column change) still takes effect on a deploy (workers restart) or within the cache TTL, since it is not a per-gateway write.
  • Widening for a 3rd EU cloud provider: when a further EU route (e.g. Claude via Bedrock-EU or Gemini via Vertex-EU) is provisioned, add it to the tenant provider_allowlist — otherwise it is residency-OK yet allowlist-blocked.

Cloudflare Workers AI

Field Type Default Description
cf_account_id string | null null Cloudflare account ID. Required for the Cloudflare Workers AI provider; the gateway builds the request URL from it. Rejected: a value with a slash, ?, #, :, @, or whitespace (path injection); the request then fails with a configuration error.

HuggingFace

Field Type Default Description
hf_endpoint string | null null URL of a HuggingFace dedicated inference endpoint. When set, requests go to this endpoint instead of the serverless API. Routed through the same SSRF guard as provider_base_urls: resolved once with the dial IP pinned; a value resolving to a private/internal address is rejected at dispatch.

Provider base URL overrides

Field Type Default Description
provider_base_urls object {} Map of provider_name → base URL. Overrides the hardcoded default endpoint of the gateway for any provider. Useful for internal proxies, self-hosted endpoints, and staging environments. A base only — the provider's endpoint path is appended, and a query string / fragment on the value is discarded. Validated at save time: the field must be an object mapping supported provider slugs to non-empty http:// or https:// URLs with no userinfo (user@host) and no whitespace / control characters; a value that fails this shape — or a key that is not a provider the gateway can route to (an unknown slug, or a legacy alias such as vllm, which is rejected in favour of its canonical id myra) — is rejected with 400 on POST/PATCH — it can no longer be saved and silently no-op at runtime. Setting the field to null (or omitting every entry) clears the override. The full SSRF guard remains at request time: the override host is resolved once and the dial IP pinned, so an entry whose value is not a non-empty URL string (including false or null) or which resolves to a private/internal address is rejected at dispatch with configuration_error — it never silently falls back to the default endpoint. Use the Test Connection endpoint to check reachability before saving.

⭐ Example:

"provider_base_urls": {
  "openai": "https://my-openai-proxy.internal"
}

Circuit breaker

Field Type Default Description
circuit_breaker object | null null Set to enable the per-provider circuit breaker. null disables it entirely.
circuit_breaker.enabled boolean false Must be true to activate the breaker.
circuit_breaker.failure_threshold integer 5 Number of failures within window_sec before the breaker opens.
circuit_breaker.window_sec integer 60 Window in seconds over which failures are counted. The counter's lifetime is window_sec x 2 and is not extended by later failures, so the window is tumbling, not sliding.
circuit_breaker.cooldown_ms integer 30000 Milliseconds to wait in the Open state before allowing a probe request.
circuit_breaker.failure_status_codes array [500,502,503,504] HTTP status codes that count as failures. Connection and timeout errors always count. Must be an array of whole numbers ([] = never trip on status); a malformed value is rejected with 400 on write and fails closed to the default at runtime.

See Circuit Breaker for state machine details and examples.


Webhooks

Webhooks deliver structured event payloads to an external HTTP endpoint for integration with alerting, ITSM, and automation systems.

Field Type Default Description
webhooks object | null null Set to enable outgoing webhooks. null disables them.
webhooks.url string — HTTPS endpoint that receives POST requests for each event.
webhooks.secret string — Optional shared secret. When set, each delivery carries an X-AIG-Signature: sha256=<hex> header — the HMAC-SHA256 (RFC 2104) of the exact request body, keyed with this secret. Recompute it over the raw body in your receiver to verify origin and integrity.
webhooks.events array all events Event types to deliver. Supported values: blocked, budget_exceeded, budget_threshold, circuit_open. When this key is absent (or not an array), every supported event is delivered. Providing an array restricts delivery to the listed events; note that an existing filter that does not list budget_threshold will not receive the new soft-alert event until you add it.

⭐ Example:

"webhooks": {
  "url": "https://hooks.slack.com/services/...",
  "secret": "mysecret",
  "events": ["blocked", "budget_exceeded", "budget_threshold"]
}

Webhook payload

Every delivery is a POST with a JSON body of the shape:

{
  "event": "budget_threshold",
  "gateway_id": "…",
  "tenant_id": "…",
  "ts": 1752300000,
  "data": { }
}

ts is the Unix epoch (seconds) at send time. The data object depends on the event:

Event data fields
budget_exceeded scope ("token" | "tenant" | "gateway" — which budget was hit), budget_usd (the cap), spent_usd (current-period spend), period (the period key, e.g. "2026-07", "2026-07-15", or "total"). Fired when a request is blocked because spend reached the cap — except for self-serve tenants with cap soft-degrade configured, where reaching the base allowance fires it once per billing period with degraded: true and stage: "primary" while requests continue on the efficient models; the true stop at the secondary cap then fires per-request with degraded: false and stage: "backstop". Payloads without degraded/stage are the unchanged manual-tenant / token / gateway hard stops.
budget_threshold Same fields as budget_exceeded, plus threshold_pct (the configured budget_alert_pct, e.g. 0.8). Fired once when spend crosses the soft threshold without being over the cap; the request is not blocked.

SIEM (gateway-level override)

A gateway-level siem key overrides the tenant-level SIEM (Security Information and Event Management) config for that specific gateway. All fields are identical to the tenant-level config.

Field Type Default Description
siem object | null null SIEM backend config for this gateway. Overrides the tenant default. Set null to remove the override and fall back to the tenant config.
siem.type string — Backend: splunk_hec, elasticsearch, vector, syslog.
siem.events array ["blocked"] Event filter. Values: blocked, guardrail, scrubbed, all.

See SIEM Integration for the full field reference and per-backend examples.


Tracing

Field Type Default Description
tracing object | null null Set to enable request tracing. null disables tracing entirely.
tracing.enabled boolean false Activate internal pipeline tracing (Traces API + Playground traces).
tracing.include_bodies boolean false Store raw content in the request and web-search trace-step fields (request message arrays, leg-1 response previews, and web-search queries/URLs/page previews). When off, those fields are omitted and only metadata (counts, sizes, status) is stored. Enable only for debugging.
tracing.otlp_endpoint string — Base URL of an OpenTelemetry collector (e.g. http://otel-collector:4318). Setting this enables OTLP span export.
tracing.service_name string "ai-gateway" service.name resource attribute on all emitted OTLP spans.
tracing.headers object {} Extra HTTP headers to include in the OTLP request (e.g. auth tokens for managed collectors).
tracing.sample_rate number 1.0 Fraction of requests to export via OTLP (0.0 = never, 1.0 = always).

💡 Note: Trace retention is not a per-gateway field. Stored traces are purged automatically by a deployment-wide background job; the window is a deployment setting configured by your operator (default 48 hours). See Request Tracing — Trace retention.

💡 Note: Conversation (chat content) retention is separate. It is disabled by default and enabled per deployment by your operator (default off), then per tenant via conversation_retention_days. See Conversation retention and deletion.

💡 Note: Request-log retention (deletion of request_log prompt/response rows) is separate again, disabled by default and enabled per deployment by your operator (default off), then per tenant via request_log_retention_days. The request_log_legs cost ledger is never deleted. See Data retention — Request-log retention.

💡 Note: Replay-record capture is separate again and default OFF. When a tenant sets replay_capture_enabled, the gateway persists a raw per-turn replay record — the client messages, the intra-turn tool calls/results, the available tool set, and the sampling params — i.e. the full context needed to faithfully re-run a turn through the gateway with a different model (for offline model evaluation). It is bounded by replay_capture_ttl_hours (default 48, reaped hourly) and replay_capture_cap (default 200 per tenant+gateway), residency-tagged, deleted on Art. 17 user erasure (by both the turn owner and the run-as driver) and on tenant purge, and is not captured on guardrail-blocked turns. It is a raw (unmasked) store — enable it only where that retention is acceptable.

See Request Tracing for the full pipeline step reference and OTLP integration guide.


Image generation

Field Type Default Description
image_generation object | null — (absent) Controls whether the EU image-generation tool (generate_image) is offered on this gateway. Absent by default — an admin- or API-created gateway has no key, and the tool is withheld until { "enabled": true } is set; self-serve signup is the one path that seeds it on. Set { "enabled": false } to withhold the tool (the admin UI's toggle persists null, which withholds it too). The tool is offered only when this is on and the request carries the x-aig-image-gen header (see the inference API reference).
image_generation.enabled boolean false Offer the generate_image tool to the model.
image_generation.model string — Optional override of the EU-hosted image model the tool calls (see Image generation); absent → the deployment's default image model.
Field Type Default Description
web_search object | null null Set to enable web-search augmentation. null disables the feature.
web_search.enabled boolean false Activate web search on this gateway.
web_search.api_key string — API key for the configured web_search.provider (Brave Search or linkup.so). Required for all providers except Google Gemini.
web_search.max_results integer 5 Maximum number of search results to retrieve per query.
web_search.mode string "opt-in" "opt-in" — only triggered when the client sends the x-aig-web-search: 1 request header. "always" — attempted on every request.
web_search.max_searches_per_turn integer 8 Per-turn cap on distinct web searches on the model-driven tool-loop path (near-duplicates are de-duplicated). API-managed; read fail-safe (invalid/zero/non-finite → default). See Web Search settings.
web_search.provider string "brave" Search backend: "brave" or "linkup" (EU-sovereign). When the gateway sets no provider of its own, the tenant web_search_provider default fills it (and, under an armed EU residency floor, a non-EU per-gateway provider is clamped up to an EU tenant default). See Web Search — Supported providers and tenant web_search_provider.

💡 Note: Web-search per-query cost is a DB catalog row, not a deployment setting. Migration 0226 seeds a model_price row for provider='linkup', model='web_search' with input_per_1k = 10.0 — the per-1,000-queries rate (10.0 = $0.01 per query), so a web-search leg reports pricing_source=db and is visible and editable on the Provider Costs page (Settings › Costs) like every other model. In the UI that row is flagged per query (its Input field is USD per 1,000 queries, not per token) so an edit is not a 1000× mis-price. Only linkup is metered (Brave is the historical unmetered default and has no cost row). A built-in default rate (0.01) is now only the last-resort fallback used when the catalog row is absent (a fresh DB before the migration, or an operator DELETE). Propagation: the resolved price is memoised per worker for up to 1 hour (uniform across all model prices; the cache is node-local and not invalidated on edit), so a live UI edit takes effect within 1 hour or on the next worker restart; the migration restarts workers, so the seed is immediate. Overshoot bound: the budget cap is checked before a request starts, but search cost is recorded during the turn, so a single turn can overshoot its cap by at most 100 linkup queries (25 tool-loop rounds × 4 parallel calls, ≈$1.00 at the default rate) before the next request is blocked. This is bounded and accepted, not a bug — see Web Search for the full provider notes.

See Web Search for provider support details and usage examples.


Semantic cache

The semantic cache serves a stored response when a new request is close in meaning to an earlier one, even when the request text is not identical. The semantic cache is separate from the exact-match cache that cache_ttl controls and requires an external embedding service.

Field Type Default Description
semantic_cache object | null null Set to enable the semantic cache. null disables it.
semantic_cache.enabled boolean false Activate the semantic cache for this gateway.
semantic_cache.embedding_url string — URL of the embedding service that produces the request vectors. Required; the semantic cache does nothing when this is unset.
semantic_cache.embedding_api_key string — Optional bearer token for the embedding service.
semantic_cache.embedding_model string "text-embedding-3-small" Model the embedding service uses to produce request vectors.
semantic_cache.max_candidates integer 100 Maximum number of stored vectors compared against an incoming request.
semantic_cache.threshold number 0.95 Minimum cosine similarity for a cached entry to be served.
semantic_cache.ttl integer 86400 Lifetime of a semantic cache entry in seconds.

Outbound PII egress guard

The egress guard scans a request's outbound tool traffic (web search, URL fetch, MCP tool arguments) for personal data and secrets before it leaves the gateway. It is on by default; any internal error in the guard fails closed (blocks). See Guardrails — threat model.

Field Type Default Description
egress_guard object | null on Configures the outbound PII egress guard. Omit to keep the default-on behaviour; set enabled: false to disable it.
egress_guard.enabled boolean true false disables the guard for this gateway.
egress_guard.mode string "block" "block" refuses a request whose outbound tool traffic carries flagged data; "flag" logs the near-miss and allows it (for an open-web-browsing gateway).
egress_guard.pii_sets array high-signal set The identifier/secret pattern sets to scan for. The default set omits loose numeric identifiers (SSN, phone, IP) that false-positive on ordinary URL IDs — opt those in explicitly.

Agentic fetch

Agentic fetch lets the model fetch a URL and read its content during a conversation.

Field Type Default Description
agentic_fetch object | null null Set to enable agentic fetch. null disables it.
agentic_fetch.enabled boolean false Activate agentic fetch for this gateway.
agentic_fetch.model string — Optional model for the fetch sub-request. When unset, the gateway uses the Myra fleet's canonical local chat model (qwen3.8-27b). A value naming a retired Myra-fleet id with a registered successor is upgraded to the successor at dispatch (logged; no response header — the inner leg does not serve the turn). See Retired model ids. Self-serve workspaces: the override must be in the plan's model list, and a Myra-hosted override needs the EU-Gov add-on — a PATCH setting a non-permitted value is refused (400 plan_model_not_allowed / eu_gov_model_not_allowed), and a stored non-permitted value runs on the default instead. The default itself is platform infrastructure and exempt from the EU-Gov rule.

URL-fetch per-turn budgets

The fetch_url/agentic_fetch tools are bounded per turn. Both budgets are API-managed (no UI control in the Edit modal, which preserves a stored section untouched) and read fail-safe: an invalid, zero, negative, or non-finite value falls back to the default — a config typo can never uncap the tool or brick it.

Field Type Default Description
url_fetch object | null null Optional per-turn fetch budgets. null/absent means the defaults apply.
url_fetch.max_fetches_per_turn integer 24 Total URL fetches (web pages and documents) the model may make in ONE turn. On reaching it the model gets a notice naming the limit and answers from what it has. Raise it for research-heavy workloads.
url_fetch.max_doc_fetches_per_turn integer 8 Sub-cap on document extractions (PDF/DOCX/XLSX — the expensive, possibly-OCR unit) within that total.

Upload malware scanning (ClamAV)

Scans uploaded file bytes for malware before they are processed or persisted, using a self-hosted ClamAV daemon (clamd). It is opt-in per gateway and OFF by default — malware scanning is never a global default. When a gateway opts in, its upload byte-ingress paths (chat file attachments, conversation attachments, anonymous form uploads) scan fail-closed: an infected file is rejected (422, code: malware_detected) and an unreachable/erroring scanner rejects rather than accepting unscanned bytes (503, code: scan_unavailable, with Retry-After). Tenant-scoped uploads that have no single gateway (project knowledge, personal My-PII lists, app-feedback evidence images — the last scanned on the reporter's tenant) are scanned when the tenant has scanning enabled on any of its gateways. That enablement read is itself fail-closed: if it cannot be determined whether scanning applies (the tenant's gateway list is unreadable, or a gateway config will not decode) the upload is refused with 503, code: scan_unavailable, not accepted unscanned. An account with no tenant at all (a platform admin's normal shape, or a tenantless user) owns no gateway, so no gateway can have opted in — such a caller's My-PII list / feedback image is knowably not scanned; a tenant id that is present but empty is a malformed row and refuses (503). The ClamAV endpoint is deployment infrastructure configured by your operator (a unix socket, or a host and port), not a config field; the daemon's StreamMaxLength must be at least the largest upload cap (100 MiB).

Field Type Default Description
av_scan object | null null Per-gateway upload malware scanning. null / absent = OFF.
av_scan.enabled boolean false When true, scan this gateway's upload byte-ingress fail-closed. Default OFF.

Code interpreter

The code interpreter lets the model run a short program in a sandboxed, network-isolated interpreter and read back its output — for calculation, data analysis, and transforming data during a conversation. It works on the open / EU model stack (the sandbox runs on a self-hostable runner service, independent of any single provider).

The feature is ON by default on every gateway; the per-gateway flag is an opt-OUT. It is available when the deployment-wide runner is configured and the gateway has not explicitly disabled it:

  1. Deployment-wide runner (required) — your operator must point the gateway at a running sandbox control plane over HTTPS (a non-https:// URL is treated as a misconfiguration and keeps the feature dormant — code egress must be encrypted). When no runner is configured, the code interpreter is fully dormant: the tool is never offered, the system prompt is unchanged, and no request is ever made.
  2. Per-gateway opt-OUT — code_interpreter.enabled defaults to true; set it to false on a gateway to disable the code interpreter there. An absent / null value leaves it on (the default). Only an exact boolean false disables it; a malformed value (non-boolean enabled, or a non-object code_interpreter) is rejected at the write boundary with 400, so a fumbled opt-out can never silently leave it enabled.

Runner authentication (mTLS or bearer). The gateway authenticates to the sandbox with a client certificate (mTLS) or a bearer token; the sandbox may additionally enforce an IP allowlist. Configure at most one client-auth method:

Deployment setting Purpose
Runner URL HTTPS base URL of the sandbox control plane. Required; a non-https URL keeps the feature dormant.
Client certificate (mTLS) The PEM client certificate.
Client private key (mTLS) The PEM client private key.
Bearer token Optional bearer token, sent as Authorization: Bearer ….
Session signing key Optional. Gateway-only secret for the persistent kernel (see Persistent kernel below). Unset ⇒ stateful sessions unavailable; stateless unaffected.
  • mTLS is all-or-nothing: set both cert and key, or neither. Setting only one, or a PEM that fails to parse, is a fail-closed misconfiguration — the tool is not advertised and no request is ever made (never an unauthenticated fall-through).
  • The PEM files are read and parsed once per worker at first use (loaded at start-up); a parse failure is sticky until the service restarts.
  • The gateway always verifies the sandbox's server certificate (TLS verification is never disabled). If the sandbox presents a private-CA certificate, add that CA to the gateway's trusted-certificate bundle; for an IP:port URL the server certificate must carry a matching IP SAN.

In addition, the tool is never offered when the active provider does not support native tool-calling, when the client supplied its own tools, or for a Tier-1 local_only project (the runner is an external egress and model-generated code could carry local data).

Field Type Default Description
code_interpreter object | null null (→ on) Container for the code-interpreter opt-out. Absent / null leaves the feature on (the default). A present, non-null, non-object value is rejected with 400.
code_interpreter.enabled boolean true On by default. Set to false to disable the code interpreter on this gateway (opt-out). Absent / null → on. A non-boolean value is rejected with 400. Still requires a runner configured for the deployment.
code_interpreter.stateful boolean false Opt into the persistent kernel (variables survive across turns), layered under enabled. Requires the deployment's session signing key. Default off ⇒ stateless one-shot. See Persistent kernel below.

Enabling it for a specific tenant/gateway

The code interpreter is on by default, so nothing has to be written to enable it — it is available on every gateway once the deployment-wide runner is configured. To disable it on a specific gateway, a tenant administrator PATCHes the gateway config — PATCH /admin/v1/gateways/{id} with { "config": { "code_interpreter": { "enabled": false } } }. The caller must be a tenant admin and the gateway must belong to their tenant (a plain user gets 403; another tenant's gateway gets 403). The merge is shallow at the top level: it sets code_interpreter and leaves every other config key intact — but it replaces the whole code_interpreter object, so always send the full sub-object. A malformed code_interpreter (a non-object, or a non-boolean enabled) is rejected with 400.

Two operational facts to plan around:

  • Dormant until the runner env is set. The flag alone does nothing: unless the deployment-wide runner points at a running HTTPS sandbox control plane, the capability is fully inert (tool never offered, no request made). Enabling the flag on a deployment whose runner is not configured is a no-op until the runner is provisioned.
  • Config-cache propagation. Gateway config is cached in a shared dictionary (across all workers of the instance) for config_cache_ttl (default 30 s). An admin PATCH of the gateway config invalidates that cache entry immediately, so enabling — and, more importantly, disabling — the flag on a gateway that is already serving traffic takes effect on the next request, not after the TTL. (A change written directly to the DB, bypassing the admin API, still propagates only within one cache TTL.)

PII / Datenschutz gateways — operator judgement, not a runtime block. Enabling the code interpreter on a PII-active gateway is not blocked at runtime (the only runtime egress block is a Tier-1 local_only project or a per-run no-egress profile; the PII axis is independent and is not consulted for the code-interpreter offer). Model-generated code egresses to the runner and can carry data the model saw in context — including PII it transcribes into code literals. The input_files raw-attachment path is separately PII-gated fail-closed (see below), but the code itself is not. So on a Datenschutz gateway this is an operator decision: enable it only when the runner is EU-hosted / in-house (there is no code-enforced residency tie between the runner and a gateway's eu_region_routing — the runner's location is an operator obligation) and the residual code-literal egress is acceptable for that gateway's data class.

Runner contract (trust boundary)

The runner is treated as an untrusted upstream. The gateway posts to POST <runner-url>/v1/execute. The request body is data-minimized — it carries only the program, optional stdin, and resource limits; never conversation history, PII, secrets, or tenant identity (the sandbox authenticates the gateway, not the end user):

{ "code": "<program>", "stdin": "<optional>",
  "limits": { "wall_seconds": 30, "memory_mb": 512, "output_bytes": 65536 } }

The only request headers are Content-Type, an opaque x-aig-request-id correlation id, and the client-auth header (Authorization, when a bearer token is set). The encoded job is capped at 128 KiB; an over-cap program is rejected before egress with a clear "too large" error.

Data-file analysis (input_files)

The code_interpreter tool accepts an optional input_files argument — the filenames of data files (spreadsheets / CSV) the user attached to the conversation. When the model passes one, the gateway looks up that attachment's raw bytes (scoped to the conversation and its owner — no cross-conversation or cross-user access), validates it, and stages it into the sandbox at /input/<filename> so the program can read the real file directly (e.g. pandas.read_excel("/input/data.xlsx")). This lets any model compute exact figures over an uploaded spreadsheet instead of estimating from extracted text, and the raw file is analysed inside the self-hosted sandbox — it is never uploaded to the model provider.

  • Accepted files: spreadsheet / tabular data only, by extension — xlsx, xlsm, ods, csv, tsv (an .xlsx uploaded as application/octet-stream is accepted; OOXML/ODS files are magic-verified as zip containers). Legacy .xls, images, PDFs and other types are not staged.
  • Which version is staged (filename resolution): a name may match a chat attachment the user uploaded, a file in the conversation's project knowledge, or a file the model itself generated on an earlier turn (so a follow-up like "add a total row to that spreadsheet you made" can re-open and edit it). When more than one of these shares the same filename, the gateway stages the most-recently-created version — a later generated edit supersedes the upload it was derived from, and a genuine re-upload supersedes an older generated file. If that newest version cannot be loaded (mid-ingest, corrupt, or over a cap) the gateway falls back to the next-available same-named version and notes the substitution to the model rather than failing the run. Resolution is always scoped to the caller's own conversation/project — no cross-conversation, cross-user, or cross-tenant access.
  • Bounds (fail-soft): ≤ 8 files, ≤ 8 MiB each, ≤ 16 MiB total. A file that is missing, unowned, the wrong type, corrupt, or over a cap is skipped with a note to the model — the rest of the run proceeds.
  • PII gating (fail-closed): raw bytes may enter the in-house sandbox, but the analysis result returned to the model can contain PII (names, salaries) and, on an external-provider gateway, would then leave to the provider. So on a PII-active gateway, data-file staging is allowed only when the gateway has a fail-closed general-PII detector (pii_protector) whose tool-result masking reliably masks that result before it egresses (the result passes the standard tool-result PII scrub, which reversibly tokenises PII for the provider and restores it in the final answer to the user). A PII-active gateway that masks only with custom_pii (keyword-only) or presidio (fails open on an outage outside a PII mandate) does not stage the raw file — the model still runs code, just without the attachment. A non-PII gateway stages freely.
  • Runner dependency: the gateway emits input_files in the /v1/execute request as [{ "name": "<bare filename>", "b64": "<standard-base64 raw bytes>" }]; the sandbox runner must stage each into the guest /input/<name> (see the internal sandbox design, ops/sandbox/DESIGN.md §1.1). Data-minimization holds: only the explicitly-requested, owned attachment(s) are sent — never conversation history, other attachments, or other-tenant data.

The gateway accepts a JSON object response:

{ "stdout": "string", "stderr": "string", "exit_code": 0, "timed_out": false,
  "truncated": false, "error": null,
  "artifacts": [ { "mime_type": "image/png",        "filename": "chart.png",   "b64": "<standard-base64 PNG>" },
                 { "mime_type": "text/csv",         "filename": "results.csv", "b64": "<standard-base64 CSV>" } ] }

Accepted / rejected shape (all response fields optional, but type-strict when present; the gateway fails closed on anything malformed):

  • stdout / stderr — must be strings when present; a non-string is rejected. Each stream is scrubbed of invalid UTF-8 and capped (64 KB) before being returned to the model.
  • exit_code — must be an integer when present (NaN, Infinity and non-integers are rejected). Absent → reported as “unknown”.
  • timed_out — boolean; anything else is treated as false.
  • truncated — boolean; when the sandbox trims a stream to output_bytes, the gateway propagates the truncation notice to the model.
  • error — a broker-level rejection reason (e.g. no_isolation_available, busy, job_too_large). Any present, non-null error fails the whole call closed — it is never rendered as an empty success, regardless of exit_code or HTTP status (a sandbox that could not isolate the job must not look like a job that ran and produced nothing).
  • artifacts — optional array of files the program produced (a chart, or a data file such as a CSV or .xlsx of results). Each element carries b64 (standard-alphabet base64; base64url is rejected) plus advisory mime_type / filename. Every artifact is validated fail-closed and any that fails is dropped (its bytes never reach the user), not rendered. Three kinds are accepted; the runner's declared mime_type is never trusted for the decision:
  • PNG chart — accepted by magic bytes (filename-independent; a non-PNG labelled image/png is rejected), then stored and rendered inline in the conversation, downloadable, and included in chat export via the existing generated-image path. Also gated on pixel dimensions ≤ 25 MP (a small PNG declaring gigapixel dimensions — a browser decompression bomb — is rejected; the same ceiling the document exporter enforces).
  • Text data file — a .csv / .txt (must be valid UTF-8 with no NUL byte) or .json (must parse as JSON). The filename must be a bare filename (a path-bearing or control-character name is rejected); its extension selects the kind and the gateway-pinned content type. Accepted files are delivered as a downloadable file card in the conversation (the same path the write_file tool uses); a duplicate name is auto-suffixed (results-2.csv) so it never overwrites an earlier file. Up to 4 data files are persisted per turn.
  • .xlsx workbook (binary) — a genuine Excel workbook the program produced (e.g. wb.save("report.xlsx")). Accepted by shape, not by the runner's declared type: the bytes must be a ZIP package (PK\x03\x04 magic) carrying the [Content_Types].xml part and an xl/ part (a .docx/.pptx OOXML zip, or a .xlsx name over non-ZIP bytes, is rejected). The filename must be a bare filename. The gateway stores and serves the bytes verbatim and never parses them (no server-side XXE / zip-bomb surface); it is delivered as a downloadable binary file card, served by an authz'd conversation download route. Shares the same per-artifact / total / count caps below.
  • Everything else is rejected: an unknown or absent extension, a non-PNG binary that is not a shape-valid .xlsx, a .docx/.pptx (or other non-spreadsheet OOXML) mislabeled .xlsx, unparseable JSON, a .txt/.csv containing binary, an oversize file, or more than 4 files.
  • Shared caps across all artifacts: per-artifact size ≤ 4 MB, total ≤ 8 MB, count ≤ 4 (extras rejected).
  • Degrades gracefully: a runner that returns no artifacts field (or only PNGs) is fully supported — the run still returns its text output; there is simply nothing extra to deliver.
  • The whole response body is size-bounded (12 MB; no OOM), and a non-200 status (including 400/401/413/429/503/500), a connection error, or a timeout yields a friendly “temporarily unavailable” tool result — never a crash or a leak of the sandbox's internal status to the model.

The model-supplied code and stdin are size-bounded (64 KB each, and ≤ 128 KB combined job) before egress and are never executed by the gateway — they run only inside the runner sandbox. The rendered runner output is fenced as untrusted content, so stdout cannot forge the gateway's content markers or override instructions.

The gateway integration is default-dormant and fail-closed; deploying the sandbox runner service itself is a separate operational step.

Persistent kernel (stateful sessions)

By default every code-interpreter run is stateless — a fresh sandbox per call, with no variables carried between turns. A gateway may opt into a persistent kernel: a live per-user Python kernel whose namespace survives across turns (load a spreadsheet once, then reference the dataframe in a later turn without re-loading), evicted after ~2 hours idle. This is a phase-2 capability enabled for demo / internal gateways; stateless remains the default everywhere else and is unaffected.

Enabling it requires two things beyond the stateless prerequisites:

Config / setting Purpose
code_interpreter.stateful (gateway config, boolean, default false) Per-gateway opt-in, layered under code_interpreter.enabled — a stateful-on gateway that isn't offering the tool at all can never mint a kernel.
Session signing key (deployment setting) Gateway-only secret from which the opaque session token is derived. Unset → stateful sessions are silently unavailable (the gateway falls back to the stateless one-shot path); stateless is unaffected either way.

When both are set and the request has an authenticated tenant, conversation, and user, the gateway mints a per-(tenant, conversation, user) session token and threads it to the runner; a missing principal falls back to stateless (never a mis-scoped kernel).

Isolation (per user, not per conversation). The token is "v1:" + HMAC-SHA256(key, "sess:v1:" + tenant + \x1f + conversation + \x1f + user). Because user is in the pre-image, a different participant in the same live-shared conversation derives a different token → a different kernel → they structurally cannot read each other's variables. The token is an opaque, unguessable map key — the runner never receives the tenant/conversation/user and does not (and cannot) verify the HMAC; the trust boundary remains mTLS + the client-CN pin exactly as for the stateless path. A tampered or guessed token simply misses the runner's session map and gets a fresh kernel — never someone else's.

Request additions (session mode). Alongside the stateless {code, stdin, limits} body, a session request carries session (the token) and, on a kernel-creating run, session_tags (three opaque HMAC correlators — conv_tag, user_tag, tenant_tag) the runner records so an erasure can destroy kernels by conversation / user / tenant. These are HMACs, never raw ids — the only widening of the data-minimized wire, EU-resident and mTLS-only. A request with no session is byte-for-byte the stateless body.

Response addition. The runner may return kernel_reset: true when it spawned a fresh kernel for this run (first use, idle-evicted, or a wall-clock timeout wiped the namespace — a timeout SIGKILLs the kernel, so any timeout loses all persistent state). The gateway surfaces a one-line honesty note to the model ("variables from earlier turns are not available; re-create what you need"). A stateless or older runner omits the field → treated as no reset.

GDPR teardown (Art. 17). Kernels are ephemeral (idle-evicted, never persisted to disk beyond the VM's discarded overlay), but the gateway also actively destroys them on the three erasure events by calling POST <runner>/v1/session/destroy with { "scope": "conversation" | "user" | "tenant", "tag": "<the matching HMAC tag>" } (same mTLS + CN pin). The tag is computed from the erasure event's tenant/conversation/user via the same derivation, so no gateway-side session tracking is needed and a cross-tenant destroy is impossible (every tag embeds the tenant in its pre-image). The call is best-effort / fail-soft: a runner-unreachable teardown is logged loudly for the retention audit but never blocks the database erasure.

Accepted / rejected (session inputs, at the runner boundary): the runner validates the token shape only (v1: + 64 lowercase hex) fail-closed — a malformed token is a 400; a well-formed but unknown token misses the map and spawns a fresh kernel. The destroy body is validated fail-closed: scope must be one of session/conversation/user/tenant, and tag must be a token (session scope) or a 64-hex tag (the others); anything else is a 400. Capacity is bounded (a small ceiling of live kernels sized to the runner's memory, plus per-tenant and per-user caps); exhaustion returns a clean "temporarily unavailable", never a queue or crash.

The side-channel residual of co-resident per-user micro-VMs on a shared host is a documented, accepted risk for the demo / internal gate (the feature is not enabled for production PII tenants without dedicated-host / SMT-off mitigation) — see ops/sandbox/DESIGN.md and docs/internal/code-exec-sandbox.md.


Context compaction

Context compaction automatically summarises older conversation turns when the estimated input token count exceeds a configured threshold. This keeps long conversations within model context limits while reducing the cost of each subsequent turn.

Context compaction is available for Anthropic models that support the native compact_20260112 strategy only. For other providers — and for Anthropic models without that strategy (for example claude-haiku-4-5, claude-sonnet-4-5, claude-opus-4-5) — the setting is present in the config but injects nothing.

Field Type Default Description
context_compaction object {enabled: true, threshold_tokens: 200000, keep_last_turns: 10} Context compaction settings. When the key is absent (or is not an object), the default applies. Set enabled: false — or the key to JSON null — to disable.
context_compaction.enabled boolean true Activate automatic context compaction for this gateway. Strictly boolean: only true enables; any other value (1, "true", null) leaves compaction off.
context_compaction.threshold_tokens integer 200000 Trigger compaction when the estimated input token count reaches this value. Minimum: 50000 (Anthropic API requirement).
context_compaction.keep_last_turns integer 10 Retained for compatibility. Compaction is performed by the native context management of the Anthropic API, which controls how many recent turns are kept. The gateway does not enforce this value separately.

See Context compaction on the Anthropic provider page for background and behavioural details.

Compaction retry threshold

compact_error_threshold (top-level integer, default unset) configures the consecutive provider-error count after which the gateway forces a context-compaction probe. Set to a positive integer to enable; leave unset or set to null to disable.


Prompt caching

Anthropic prompt caching writes frequently-used parts of the prompt (system instructions, large context blocks) to a cache and bills cache reads at a steep discount. It is on by default; configure per gateway:

Field Type Default Description
prompt_caching object {enabled: true, ttl: "1h"} Prompt-caching settings. When the key is absent (or is not an object), prompt caching is enabled by default. Set the key to JSON null to disable it.
prompt_caching.enabled boolean true Activate prompt caching for Anthropic requests. Send a real boolean: only false (or an omitted field) disables; any other value — including the string "false" — is read as enabled.
prompt_caching.ttl string "1h" Cache TTL for written entries. One of "5m" or "1h". The "1h" default applies when the whole prompt_caching key is absent; an explicit prompt_caching object that omits ttl (or carries an unrecognised value) caches with Anthropic's "5m" default.

Per-request header overrides

These headers can be sent on individual inference requests to override gateway config for that request only.

Header Type Description
x-aig-byok-alias string Use a non-default BYOK provider key alias for this request. Must match an alias stored for the resolved provider.
x-aig-meta-{key} string Attach a custom key-value pair to the request log entry and make it available in routing rule conditions as meta:{key} (colon notation) and load-balancer sticky fields as meta.{key}. See the caps below.
x-aig-collect-log "0", "false", or "1" "0" or "false" = skip writing this request to the log table entirely. "1" = log (default).
x-aig-collect-log-payload "0", "false", or "1" "0" or "false" = log request metadata but omit the prompt and response body. "1" = log body (default). Does not affect the gateway-level log_payloads setting.
x-aig-provider-{field} string Strip the x-aig-provider- prefix and forward the header to the upstream provider only if the stripped name is allow-listed (anthropic-beta, openai-organization, openai-project). Any other stripped name — a credential, a request-framing / hop-by-hop header, or an unrecognised name — is dropped and logged at warn, so a client can neither inject an upstream credential nor tamper with request framing. At most 16 overrides / 8 KB total per request. Useful for provider-specific beta flags.

💡 Note: An allow-listed x-aig-provider-* header is forwarded to whatever provider handles the request, not just the one the name belongs to. Sending openai-organization to an Anthropic request is harmless (the provider ignores the unknown header). Only the three allow-listed names are ever forwarded.

🔒 x-aig-meta-* accepted shape and caps. Client metadata is sanitized at the trust boundary before it becomes a routing input or is persisted to the request log: - Value — a single string. If the same x-aig-meta-{key} header is sent more than once, only the first value is used; a value longer than 256 characters is truncated on a UTF-8 character boundary (invalid byte sequences are scrubbed). - Key — the {key} suffix must be at most 64 characters; a longer key is rejected (dropped, not truncated, so it can never collide with another key or a routing rule). - Count — at most 16 client metadata keys are kept per request. When more are sent, a deterministic subset (the keys sorted lexicographically, first 16) is kept so the same request always routes identically. - Reserved namespace — keys beginning aig_ / aig- are gateway-owned and are dropped if sent by a client.

Over-cap or malformed metadata is dropped or clamped (never a 4xx), and the request log row is always written — an oversized meta object is bounded so it can never overflow the column and lose the row.


Knowledge-search reranking

Project knowledge search (hybrid dense + keyword retrieval) can optionally run a final cross-encoder rerank stage that re-scores the fused candidate passages for relevance to the query and keeps the most relevant few. It is a single deployment-wide setting configured by your operator — there is no per-gateway field:

Deployment setting Default Effect
Reranker model (unset) When unset (the default), the rerank stage is dormant and knowledge search returns exactly the hybrid-fused order — no rerank call, no behaviour change. When set to a reranker model id available on the EU model fleet (e.g. bge-reranker-v2-m3), retrieved passages are reranked and the top few kept.

The reranker response is treated as untrusted: if the reranker is unavailable, times out, or returns a malformed/incomplete response, retrieval fails closed to the pre-rerank hybrid order — it never blanks or drops results. Enabling it requires the reranker model to be deployed on the fleet first (an operations dependency).

💡 Note: The reranker model is a deployment setting configured by your operator. Leaving it unset keeps reranking off.


Config and BYOK cache TTLs

These deployment-wide settings tune how long the gateway caches resolved config and decrypted BYOK keys. Both apply per instance, not per gateway.

Deployment setting Default Effect
Config cache TTL 30 (seconds) How long a gateway's resolved config is cached in the shared dictionary across all workers of the instance. An admin PATCH of the config invalidates the entry immediately; a change written directly to the database propagates only within this TTL.
BYOK cache TTL 60 (seconds) How long a decrypted BYOK (provider key vault) key is held in the in-memory cache after decryption, so a burst of requests does not re-decrypt the stored key on every call.

Applying config changes

# Enable caching and set a budget
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
  -H "Content-Type: application/json" \
  -d '{
    "config": {
      "cache_ttl": 300,
      "budget_usd": 200.00
    }
  }'

# Disable authentication (development only)
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
  -H "Content-Type: application/json" \
  -d '{"config": {"auth_required": false}}'

# Add an IP allowlist
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
  -H "Content-Type: application/json" \
  -d '{"config": {"ip_allowlist": ["10.0.0.0/8"]}}'

See also