Skip to content

Fallback and retry

AI providers occasionally become unavailable, apply rate limits, or return errors during high-traffic periods. Without a fallback, the gateway returns an error to the caller. With a fallback chain, the gateway silently retries against another provider and the caller receives a response.

Fallbacks are defined per routing rule and are transparent to the caller — the response format is the same regardless of which provider in the chain ultimately handled the request.

How the fallback chain works

A request follows this sequence:

  1. Primary provider — retry_count is the number of retries after the first attempt, so the primary is attempted up to retry_count + 1 times on server errors (HTTP 5xx) and rate-limit responses (HTTP 429, retried with back-off honouring Retry-After). The default retry_count of 2 therefore gives 3 attempts (1 initial attempt plus 2 retries).
  2. Fallback providers — each fallback is attempted once, in order. retry_count applies only to the primary provider. Fallbacks are defined in the routing rule that matched the request.
  3. If all providers are exhausted, the terminal status reflects the last failure in the chain. A server or connection error ends in 502 all_providers_failed, but a final 429 surfaces 429 rate_limited (forwarding the provider's Retry-After), a read timeout 504, and an oversized body 413. A deterministic policy refusal on the final attempt (data residency, provider allowlist, sub-processor objection, a deprecated/unpriced/disabled model, the role model-mask, or a provider auth-config error) surfaces its own status rather than a masking 502. See Terminal status by last failure.
flowchart LR
    Start([Request]) --> P1

    subgraph Primary ["Primary provider"]
        P1[Attempt 1] -- "5xx / 429" --> P2[Attempt 2]
        P2 -- "5xx / 429" --> P3[Attempt 3]
    end

    P1 -- "4xx (except 429)" --> Bail(["Return 4xx immediately<br/>no retry or fallback"])
    P3 -- "5xx / 429" --> F1[Fallback 1]
    F1 -- "5xx / 429" --> F2[Fallback 2]
    F2 -- "5xx (last failure)" --> E([502 all_providers_failed])
    F2 -- "429 (last failure)" --> R([429 rate_limited])

Client errors

A client error response from a provider (HTTP 4xx — for example, bad request, unauthorised, not found) is treated as a definitive failure. The gateway does not fall back to another model and, in general, does not retry. The error is returned to the caller immediately.

A narrow set of transparent, gateway-side self-recoveries are the exception to "no retry": when the gateway can see that its own request body caused the 4xx, it fixes the body and re-sends the same model once — stripping a deprecated parameter, dropping a forced tool_choice the backend cannot honour, removing gateway-injected tools the backend refuses, or stripping an unsupported reasoning parameter. These never change the model, never fire more than once per turn, and never fire after streaming has begun; if the single resend also fails, the original 4xx is surfaced.

HTTP 429 (rate limited) is the exception and is not treated as a definitive failure: the gateway retries it against the primary with back-off (honouring the provider's Retry-After header) and, once the primary's attempts are exhausted, falls through to the fallback chain just like a 5xx.

💡 Note: This prevents pointless retries: if your request is malformed or your API key is invalid, retrying against another provider will not help. A 429 is different — it is transient, so it is retried and does fall through the chain.

retry_count

retry_count controls the number of retries against the primary provider after the first attempt (not the total number of attempts). The primary is therefore attempted retry_count + 1 times. Only the primary provider uses retry_count; each fallback is tried exactly once. Set it in the gateway config:

curl -X PATCH https://your-gateway-host/admin/v1/gateways/{id} \
  -H "Cookie: aig_admin=<SESSION>" \
  -H "Content-Type: application/json" \
  -d '{"config": {"retry_count": 3}}'
retry_count Primary attempts Retries
1 2 1
2 3 2
3 4 3

Fallback configuration in routing rules

Fallbacks are defined in the actions object of a routing rule:

{
  "priority": 10,
  "conditions": [
    {"field": "model", "op": "prefix", "value": "gpt-"}
  ],
  "actions": {
    "provider": "openai",
    "model": "gpt-4o",
    "fallbacks": [
      {"provider": "anthropic", "model": "claude-sonnet-4-6"},
      {"provider": "gemini", "model": "gemini-2.0-flash"}
    ]
  },
  "enabled": true
}

Each entry in fallbacks is attempted once, in array order.

💡 Note: The per-rule timeout_ms action bounds every leg of this rule — the primary and each fallback leg alike. See the timeout_ms action in the routing rules API reference.

Configuring fallback providers

Before you begin, ensure the following conditions are met:

  • ☑ You have admin access.
  • ☑ A gateway with at least one routing rule exists.

Screenshot: Routing rule editor with fallbacks section visible The fallbacks section of the routing rule editor.

Proceed as follows to add fallback providers to a routing rule:

  1. Open Gateways in the left sidebar.
  2. The gateway list opens.
  3. Click on the gateway you want to configure.
  4. The gateway detail page opens.
  5. Scroll to the Routing Rules card.
  6. The rule list opens.
  7. Click on the routing rule you want to edit.
  8. The rule editor opens.
  9. Scroll to the Fallbacks section.
  10. The fallbacks list is visible.
  11. Click on the + Add button.
  12. A new fallback row appears.
  13. Select the fallback provider from the Provider drop-down list.
  14. The provider is set.
  15. Enter the model name in the Model text field.
  16. The model is set.
  17. Repeat steps 6–8 for each additional fallback provider, in the order the gateway should try them.
  18. Each fallback is added to the list.
  19. Click on the Save Rule button.

-> The routing rule is updated with the configured fallback chain.

To create the same configuration via the API:

curl -X POST https://your-gateway-host/admin/v1/gateways/{id}/rules \
  -H "Cookie: aig_admin=<SESSION>" \
  -H "Content-Type: application/json" \
  -d '{
    "priority": 10,
    "conditions": [{"field": "model", "op": "prefix", "value": "gpt-"}],
    "actions": {
      "provider": "openai",
      "model": "gpt-4o",
      "fallbacks": [
        {"provider": "anthropic", "model": "claude-sonnet-4-6"},
        {"provider": "gemini", "model": "gemini-2.0-flash"}
      ]
    },
    "enabled": true
  }'

BYOK key swap on provider change

When a fallback triggers a provider change, the gateway automatically selects the BYOK (bring your own key) key for the new provider. The x-aig-byok-alias header applies only to the primary (originally requested) provider; a fallback provider always uses its default-alias key (the header is not re-applied to fallback providers).

A BYOK key must be stored for each key-requiring provider in the fallback chain. Keyless providers (the Myra EU fleet) need no stored key — a fallback leg to them dispatches without one. If no key is found for a key-requiring fallback provider, that fallback leg is silently skipped (the gateway logs a warning) and the chain continues to the next fallback. No 424 is raised for a fallback leg — the 424 provider_key_missing error is emitted only for the primary (originally requested) provider, before the fallback loop begins.

⚠️ Caution: A missing key does not halt the fallback chain — a key-requiring fallback with no stored key is skipped and the gateway moves on to the next fallback. Store BYOK keys for every provider in your fallback chains: when every leg is skipped or fails, the request ends in 502 all_providers_failed, not a 424.

502 all_providers_failed

When every provider in the chain (primary plus all fallbacks) has been tried and none succeeded, the gateway returns:

{
  "error": {
    "code": "all_providers_failed",
    "message": "All configured providers failed"
  }
}

HTTP status code: 502.

💡 Note: The message field carries the last upstream error encountered while exhausting the chain, falling back to All configured providers failed when no more specific detail is available.

Terminal status by last failure

502 all_providers_failed is the outcome only when the chain's last attempt was a server or connection error. When the final attempt failed for a more specific reason, the gateway surfaces that reason's own status instead of masking it as a 502:

Last failure HTTP status Code
Server / connection error (default) 502 all_providers_failed
Provider rate limit 429 rate_limited (forwards Retry-After)
Read timeout 504 request_timeout
Request body too large 413 request_too_large
Data-residency block 403 (residency refusal)
Provider not on the gateway allowlist 403 (allowlist refusal)
Sub-processor objection 403 (objection refusal)
Model disabled on the gateway 403 model_disabled_on_gateway
Model blocked by the role model-mask 403 (role refusal)
Deprecated, unpriced, or capability-mismatched model 400 model_not_found / capability refusal
Provider auth misconfiguration 500 configuration_error

The deterministic policy refusals (residency, allowlist, objection, disabled/deprecated/unpriced model, role mask, auth-config) outrank the outage heuristics: a policy-blocked final attempt is reported as the policy refusal, never as a provider outage.

Logged fields

For every request that uses a fallback, the gateway logs the final provider used:

Log field Description
fallback_provider Name of the provider that ultimately served the request, if different from the primary.
fallback_model Model name used by the fallback provider.

These fields are visible in the request log table and in the admin UI log viewer.

See also