Skip to content

Request pipeline

Every request to the gateway passes through three processing phases. Each phase either short-circuits — returning a response to the client immediately — or passes the request to the next step.


Phase overview

flowchart TD
    Client([Consumer])
    Client --> A1

    subgraph Access ["Access phase — before body is read"]
        A1[Authenticate] --> A2[Rate limit] --> A3[Quota check] --> A4[IP allowlist]
    end

    A4 --> B1

    subgraph Content ["Content phase — body available"]
        B1[Cache check] --> B2[Transform & routing]
        B2 --> B3[Guardrails — request]
        B3 --> B4[Provider call]
        B4 --> B5[Guardrails — response]
        B5 --> B6["send response<br/>cost · cache store"]
    end

    B6 --> L1

    subgraph Log ["Log phase — after response sent"]
        L1["Structured log<br/>Prometheus metrics"]
    end

    L1 --> Response([Consumer Response])

Access phase

Runs before the request body is read. A rejection here is low-cost.

Authentication

Accepts a token from (in priority order):

  1. x-aig-token header
  2. Authorization: Bearer <token>
  3. x-api-key header (Anthropic SDK compatibility)

💡 Note: Returns 401 unauthorized if no valid token is found, 403 forbidden if the role of the token does not permit inference requests. Skipped when the auth_required setting of the gateway is false.

Rate limiting

Enforces the sliding-window request limit configured on the gateway or per-token. Returns 429 rate_limited with X-RateLimit-Limit, X-RateLimit-Remaining, and Retry-After headers. Both gateway-level and per-token limits are checked independently.

Quota check

Checks per-token, per-tenant, and per-gateway spend budgets against the spend_ledger. Returns 429 quota_exceeded with an actionable message if any budget is exhausted.

IP allowlist

Checks the client IP against the ip_allowlist CIDR list of the gateway. An empty list allows all traffic. Returns 403 forbidden on mismatch.


Content phase

Has access to the full request body.

Cache check

Looks up the request in the exact-match response cache. On a hit, returns the stored response immediately with X-AIG-Cache: HIT — the provider is never called. See Response caching.

Routing

Before the rules run, the request's model id is normalised (provider resolved, provider/ prefix stripped) and — for a retired Myra-fleet id with a registered successor — upgraded to that successor (see Retired model ids); a rule's model condition therefore sees the upgraded id. Evaluates the ordered routing rules. The first matching rule wins and can override the provider, model, and fallback chain. If no rule matches, the request's default provider and model are used: on a provider-native endpoint these come from the URL path, whereas on the compat endpoint the provider is inferred from the model field (by model-name prefix, falling back to OpenRouter) rather than the URL.

Guardrails (request)

Runs after routing, so the effective model leg — the resolved provider and its failover chain — is known before masking. Two-tier guardrail pipeline runs against the outbound request body:

  • Tier 1 (in-process, sub-millisecond): regex and keyword guardrails
  • Tier 2 (HTTP sidecar, milliseconds): NLP PII Detector, Prompt Guard, PII Protector — only if Tier 1 passes

A block verdict returns a synthetic error response to the client. scrub replaces matched content. flag records the match in the log. PII masking (scrub) is skipped when the request routes wholly to a first-party local (Myra/EU) model — the data never leaves Myra; content-safety block verdicts still apply. See Guardrail pipeline and PII Protector — Local model legs are not masked.

Provider call

Makes the HTTP call to the upstream provider. On a 5xx response, retries up to retry_count times (default retry_count = 2, so 1 initial attempt plus 2 retries — 3 attempts in total). If all retries fail, walks the fallback chain. Returns 502 all_providers_failed if every option is exhausted.

💡 Note: 4xx responses from the provider are returned to the client immediately with no retry.

For streaming requests ("stream": true), SSE chunks are flushed to the client as they arrive.

Guardrails (response)

Runs the guardrail pipeline against the inbound provider response before forwarding it to the client. Blocked responses are never sent to the client. PII masking (scrub) of the response is likewise skipped on a wholly-local model leg — the model output stays intact for the user — while response-phase content-safety block verdicts still apply.

Cost accounting

Extracts token counts from the provider response and computes the request cost. Increments the budget counter. See Cost attribution.

Cache store

Persists non-streaming 200 responses to the cache when cache_ttl > 0.


Log phase

Runs after the response has been sent. Failures here do not affect the client.

Writes a structured log entry containing identity, routing, status, cache state, token counts, cost, timing, guardrail results, and custom metadata. Payload logging is suppressed per gateway (log_payloads: false) or per request (x-aig-collect-log-payload: false). The log entry is skipped entirely with x-aig-collect-log: false.


See also