Request pipeline
Every request to the gateway passes through three processing phases. Each phase either short-circuits — returning a response to the client immediately — or passes the request to the next step.
Phase overview
flowchart TD
Client([Consumer])
Client --> A1
subgraph Access ["Access phase — before body is read"]
A1[Authenticate] --> A2[Rate limit] --> A3[Quota check] --> A4[IP allowlist]
end
A4 --> B1
subgraph Content ["Content phase — body available"]
B1[Cache check] --> B2[Transform & routing]
B2 --> B3[Guardrails — request]
B3 --> B4[Provider call]
B4 --> B5[Guardrails — response]
B5 --> B6["send response<br/>cost · cache store"]
end
B6 --> L1
subgraph Log ["Log phase — after response sent"]
L1["Structured log<br/>Prometheus metrics"]
end
L1 --> Response([Consumer Response])
Access phase
Runs before the request body is read. A rejection here is low-cost.
Authentication
Accepts a token from (in priority order):
x-aig-tokenheaderAuthorization: Bearer <token>x-api-keyheader (Anthropic SDK compatibility)
💡 Note: Returns
401 unauthorizedif no valid token is found,403 forbiddenif the role of the token does not permit inference requests. Skipped when theauth_requiredsetting of the gateway isfalse.
Rate limiting
Enforces the sliding-window request limit configured on the gateway or per-token. Returns 429 rate_limited with X-RateLimit-Limit, X-RateLimit-Remaining, and Retry-After headers. Both gateway-level and per-token limits are checked independently.
Quota check
Checks per-token, per-tenant, and per-gateway spend budgets against the spend_ledger. Returns 429 quota_exceeded with an actionable message if any budget is exhausted.
IP allowlist
Checks the client IP against the ip_allowlist CIDR list of the gateway. An empty list allows all traffic. Returns 403 forbidden on mismatch.
Content phase
Has access to the full request body.
Cache check
Looks up the request in the exact-match response cache. On a hit, returns the stored response immediately with X-AIG-Cache: HIT — the provider is never called. See Response caching.
Routing
Before the rules run, the request's model id is normalised (provider resolved, provider/ prefix stripped) and — for a retired Myra-fleet id with a registered successor — upgraded to that successor (see Retired model ids); a rule's model condition therefore sees the upgraded id. Evaluates the ordered routing rules. The first matching rule wins and can override the provider, model, and fallback chain. If no rule matches, the request's default provider and model are used: on a provider-native endpoint these come from the URL path, whereas on the compat endpoint the provider is inferred from the model field (by model-name prefix, falling back to OpenRouter) rather than the URL.
Guardrails (request)
Runs after routing, so the effective model leg — the resolved provider and its failover chain — is known before masking. Two-tier guardrail pipeline runs against the outbound request body:
- Tier 1 (in-process, sub-millisecond): regex and keyword guardrails
- Tier 2 (HTTP sidecar, milliseconds): NLP PII Detector, Prompt Guard, PII Protector — only if Tier 1 passes
A block verdict returns a synthetic error response to the client. scrub replaces matched content. flag records the match in the log. PII masking (scrub) is skipped when the request routes wholly to a first-party local (Myra/EU) model — the data never leaves Myra; content-safety block verdicts still apply. See Guardrail pipeline and PII Protector — Local model legs are not masked.
Provider call
Makes the HTTP call to the upstream provider. On a 5xx response, retries up to retry_count times (default retry_count = 2, so 1 initial attempt plus 2 retries — 3 attempts in total). If all retries fail, walks the fallback chain. Returns 502 all_providers_failed if every option is exhausted.
💡 Note: 4xx responses from the provider are returned to the client immediately with no retry.
For streaming requests ("stream": true), SSE chunks are flushed to the client as they arrive.
Guardrails (response)
Runs the guardrail pipeline against the inbound provider response before forwarding it to the client. Blocked responses are never sent to the client. PII masking (scrub) of the response is likewise skipped on a wholly-local model leg — the model output stays intact for the user — while response-phase content-safety block verdicts still apply.
Cost accounting
Extracts token counts from the provider response and computes the request cost. Increments the budget counter. See Cost attribution.
Cache store
Persists non-streaming 200 responses to the cache when cache_ttl > 0.
Log phase
Runs after the response has been sent. Failures here do not affect the client.
Writes a structured log entry containing identity, routing, status, cache state, token counts, cost, timing, guardrail results, and custom metadata. Payload logging is suppressed per gateway (log_payloads: false) or per request (x-aig-collect-log-payload: false). The log entry is skipped entirely with x-aig-collect-log: false.
See also
- Multi-tenancy — how tenant and gateway resolution works
- Response caching — cache key construction and TTL configuration
- Routing rules — rule engine conditions and actions
- Cost attribution — token counting and pricing lookup