Skip to content

Circuit breaker

The circuit breaker automatically stops routing traffic to a provider that is consistently failing, then probes it after a cooldown period to detect recovery. This prevents cascading failures where a broken provider repeatedly consumes retries on every request.

State machine

The circuit breaker operates as a three-state machine:

stateDiagram-v2
    [*] --> CLOSED
    CLOSED --> OPEN      : failures ≥ threshold
    OPEN --> HALF_OPEN   : cooldown elapsed
    HALF_OPEN --> CLOSED : probe succeeds
    HALF_OPEN --> OPEN   : probe fails (restart cooldown)
State Behaviour
Closed (healthy) All requests route normally
Open Requests skip this provider entirely; the next target in the fallback chain is tried
Half-open (probation) Requests are admitted again to test recovery. The first attempt that delivers an answer closes the breaker; the first failure re-opens it and restarts the cooldown

The default state is Closed. No state is stored until the first failure is recorded.

💡 Note: While the breaker is half-open, every incoming request is admitted to probe recovery — not strictly a single one. The probe is judged on the attempt's delivered outcome, not on the HTTP status: an upstream that answers 200 with no answer at all has not proven recovery, so it neither closes the breaker nor clears the failure count.

Configuration fields

Field Type Default Description
enabled boolean false Set to true to activate the circuit breaker
failure_threshold integer 5 Number of failures within window_sec before the breaker opens
window_sec integer 60 Window in seconds over which failures are counted. The counter is created with a lifetime of window_sec x 2 and is not extended by later failures, so the window is tumbling, not sliding: failures accumulate for up to twice this value and then reset in one step
cooldown_ms integer 30000 Milliseconds to wait in the Open state before allowing a probe
failure_status_codes array [500,502,503,504] HTTP status codes that count as failures. Connection and timeout errors always count regardless of this list. Must be a JSON array of whole numbers; an empty array [] means "never trip on any HTTP status" (connection/timeout errors still count).

Validation & fail-safe. failure_status_codes is validated when a gateway is created or updated: a value that is not an array of whole numbers (a string, a number, a non-empty object, or an array containing a non-integer element) is rejected with 400. Should a malformed value ever be present in stored config (e.g. written by an older client), the runtime fails closed to the default [500,502,503,504] rather than erroring — the breaker always has a well-defined set.

What counts as a failure

The following events increment the failure counter:

  • HTTP 5xx from the provider — only codes listed in failure_status_codes count (default: 500, 502, 503, 504).
  • Connection errors — DNS failure, connection refused, TLS error — always counted regardless of failure_status_codes.
  • Timeouts — treated as connection errors.

  • A dead turn from a self-hosted (inline) model — a turn that finishes cleanly with 200 (finish_reason: stop) but produces no answer at all (an instant end-of-turn). The gateway retries such a turn once automatically; only a turn that is still dead after the retry counts as a failure. This bypasses failure_status_codes — the status was 200. A turn that produced no answer because it was length-capped (finish_reason: length, Anthropic stop_reason: max_tokens — a caller's tiny max_tokens, such as a one-token health probe, or a prompt that filled the window) is not a dead turn: it is neither retried nor counted (see the neither-success-nor-failure list below).

The following events do not increment the failure counter:

  • 4xx responses from the provider — bad request, auth failure, and similar client errors are not treated as provider failures, unless the status is listed in failure_status_codes. Note that 429 is counted for the built-in Myra fleet breaker.

💡 Note: The built-in Myra fleet breaker is separate from the per-gateway breaker and is not configurable. It trips after 3 failures in a 30-second window, stays open for 15 seconds, and counts HTTP 500, 502, 503, 504, and 429. It is keyed globally (shared across all tenants and gateways), because the Myra fleet is shared infrastructure — so its defaults deliberately differ from the per-gateway defaults (5 failures / 60 s / 30 s cooldown) described above.

What counts as a success — and what counts as neither

Only an attempt that actually delivered an answer records a success (and a success is the only thing that closes a breaker and clears its failure count).

The following outcomes record neither a success nor a failure. They can never trip a healthy breaker, and they can never close one that real failures opened:

  • A discarded attempt — one the gateway retried itself (an automatic dead-turn retry, an unparseable response, a failed streaming leg). The retry's own outcome decides.
  • A turn the client cancelled mid-stream — pressing Stop tells you nothing about the provider.
  • A stream that broke mid-flight, or one the provider terminated with an error event.
  • A turn that delivered zero visible characters — for example a reasoning model that spent its whole budget on internal reasoning, or a turn cut off by the caller's max_tokens before any visible token (a one-token health probe). (Applies to the OpenAI-compatible and buffered response paths; the Anthropic-native passthrough does not classify empty responses.)

Interaction with retries and fallbacks

The circuit breaker is checked once per target in the fallback chain — not before each same-provider retry. Consider a request with retry_count: 2 and one fallback, where the breaker of OpenAI is open:

  1. Check the breaker of OpenAI → Open → skip OpenAI entirely.
  2. Check the breaker of Anthropic → Closed → attempt Anthropic.
  3. Anthropic succeeds → record success → return response.

Without a circuit breaker, the gateway would exhaust two retry attempts on a failing OpenAI before trying Anthropic.

Configuring the circuit breaker

Before you begin, ensure the following conditions are met:

  • ☑ You have admin access.
  • ☑ A gateway exists.

Screenshot: Gateway configuration page with circuit breaker section The circuit breaker configuration on the gateway detail page.

Proceed as follows to configure the circuit breaker for a gateway:

  1. Open Gateways in the left sidebar.
  2. The gateway list opens.
  3. Click on the gateway you want to configure.
  4. The gateway detail page opens.
  5. Click on the Edit button in the Gateway card header.
  6. The configuration form opens.
  7. Toggle the Enable circuit breaker toggle on.
  8. The circuit breaker configuration fields appear.
  9. Enter a value in the Failure threshold text field.
  10. The failure threshold is set.
  11. Enter a value in the Window text field (in seconds).
  12. The counting window is set.
  13. Enter a value in the Cooldown text field (in milliseconds).
  14. The cooldown period is set.
  15. Click on the Save Changes button at the bottom of the modal.

-> The circuit breaker configuration is saved and takes effect immediately.

💡 Note: The editor exposes only the enable toggle, failure threshold, window, and cooldown. The set of HTTP status codes that count as failures (failure_status_codes) is configurable through the API only — see the example below.

To configure the circuit breaker via the API:

{
  "circuit_breaker": {
    "enabled": true,
    "failure_threshold": 5,
    "window_sec": 60,
    "cooldown_ms": 30000,
    "failure_status_codes": [500, 502, 503, 504]
  }
}

Status API

Check the current state of the breaker for all providers on a gateway:

GET /admin/v1/gateways/{id}/circuit-breaker

Response:

{
  "openai": {
    "state": "open",
    "failures": 7,
    "opened_at": 1748123456
  },
  "anthropic": {
    "state": "half_open",
    "failures": 5,
    "opened_at": 1748123401
  }
}

Only providers whose breaker is currently open or half-open appear in the response. A missing provider entry means the breaker is closed; when a breaker closes, its recorded state and failure count are cleared.

💡 Note: failures counts only while the breaker is closed — it is not incremented while a breaker is open or half-open, and the counter expires on its own after window_sec x 2. A row reading "state": "open", "failures": 0 is therefore normal for a breaker that has been open for a while, not a bug.

Every state transition is also logged with a [circuit_breaker] prefix (transition=closed->open reason=provider_5xx failures=3/3 ...) and counted in the aig_circuit_breaker_transitions_total metric, labelled by provider, scope (gateway or global), transition and reason.

The dashboard UI shows this status in a live table on the gateway detail page when the circuit breaker is enabled. The Circuit Breaker card is collapsed on first load; expand it to see the live status table.

Example configurations

Conservative — trip only on sustained outage

{
  "circuit_breaker": {
    "enabled": true,
    "failure_threshold": 10,
    "window_sec": 120,
    "cooldown_ms": 60000
  }
}

Opens after 10 failures in 2 minutes. Probes after 1 minute. Appropriate when providers have occasional transient errors.

Aggressive — trip fast, recover fast

{
  "circuit_breaker": {
    "enabled": true,
    "failure_threshold": 3,
    "window_sec": 30,
    "cooldown_ms": 10000
  }
}

Opens after 3 failures in 30 seconds. Probes after 10 seconds. Appropriate when you have multiple healthy fallbacks and want to shed load immediately.

See also