Skip to content

Routing rules API

Routing rules let you rewrite the provider and model for any request without changing the caller. Rules are evaluated in priority order — the first matching rule wins. Each rule can override the provider, rewrite the model name, and attach a fallback chain that is walked when the primary fails.

Base URL: https://<your-gateway-host>/admin/v1


Endpoints

Method Path Description
GET /gateways/{id}/rules List rules for a gateway, in evaluation order
POST /gateways/{id}/rules Create a rule
PATCH /gateways/{id}/rules/{rule_id} Update a rule
PUT /gateways/{id}/rules/order Reorder all rules for a gateway
DELETE /gateways/{id}/rules/{rule_id} Delete a rule

Rule structure

{
  "id": "rule_abc123",
  "priority": 10,
  "conditions": [
    {"field": "model", "op": "prefix", "value": "gpt-"}
  ],
  "actions": {
    "provider": "openai",
    "model": "gpt-4o",
    "fallbacks": [
      {"provider": "anthropic", "model": "claude-sonnet-4-6"}
    ]
  },
  "enabled": true
}

Rule fields

Field Type Description
priority integer Evaluation order. Higher numbers run first. Rules with equal priority are ordered by rule id (a random UUID) — deterministic, and identical to the order GET /rules returns, but not the order the rules were created in.
conditions array List of condition objects. All conditions must match (logical AND). An empty array matches every request.
actions object What to do when the rule matches. Specify a provider, a model rewrite, or both — a model-only rewrite that keeps the original provider is valid.
enabled boolean false disables the rule without deleting it. Default: true.

Conditions

Each condition object specifies a field, an op (operator), and a value to compare against.

Condition fields

Field Description
model The model name from the request body, after normalisation: the provider prefix is stripped, and a retired Myra-fleet id with a registered successor (e.g. qwen3.6-27b) has already been upgraded to that successor, so a condition keyed on a retired id no longer matches — see Retired model ids.
provider The provider from the request URL path (e.g. openai, compat).
tenant_id The tenant identifier resolved from the URL.
header:{name} The value of HTTP header {name} (e.g. header:x-customer-tier).
meta:{key} Any custom metadata field attached via x-aig-meta-{key} request header. For example, meta:env matches the value of x-aig-meta-env.

Condition operators

Operator Behaviour
eq Exact string match. Case-sensitive.
neq Exact string non-match. Case-sensitive.
prefix Value starts with the given string.
contains Value contains the given substring.
regex Value matches the given pattern. Patterns use Lua-pattern syntax, not PCRE. An invalid pattern never matches.

If a condition omits op, it defaults to eq (exact, case-sensitive match).

Examples

⭐ Example: The following examples show condition objects for common matching scenarios.

Match all requests for models starting with gpt-:

{"field": "model", "op": "prefix", "value": "gpt-"}

Match a specific model:

{"field": "model", "op": "eq", "value": "claude-opus-4-6"}

Match requests tagged with a custom header (x-aig-meta-env: production):

{"field": "meta:env", "op": "eq", "value": "production"}

Match all requests (catch-all rule, empty conditions):

[]

Actions

Field Type Description
provider string Override the inference provider (e.g. openai, anthropic, gemini).
model string Rewrite the model name sent to the provider. If omitted, the original model name is used. A rewrite onto a retired Myra-fleet id with a registered successor is upgraded to that successor exactly like a client's own request; the response then carries X-AIG-Model-Upgraded-From (but no Deprecation header — the client did not name the retired id). Each fallbacks[].model is upgraded the same way, logged only: the response describes the model that served the turn. A rewrite onto any other deprecated id is refused at serve time (400 model_not_found). See Retired model ids.
fallbacks array Ordered fallback chain. Each entry is {"provider": "...", "model": "..."}. Used when the primary fails after all retries.
load_balance object Distribute matching requests across several targets by weight. See below.
timeout_ms integer Per-rule upstream timeout in milliseconds for the leg this rule routes — including a load-balanced target and each fallback leg. Bounds every wait on the upstream: the connection, the request send, and — for both streamed and buffered responses — the time to the first response byte and every inter-chunk gap of a streamed response. When set, it replaces the built-in streaming stall budgets (20 s first byte / 300 s inter-chunk) in both directions: a tighter value cuts a stalled stream sooner, a wider value grants a slow model more time to its first byte. It does not cap the total duration of a stream that is actively delivering tokens. A stall past the bound mid-stream ends the stream with the standard stream-error surface; a stall before the first byte fails the leg over to the fallback chain. If omitted, the gateway (or global) default timeout bounds connect/send and the buffered read, and the built-in stall budgets bound streamed reads. Must be a positive integer no greater than 3600000 (1 hour).

💡 Note: Fallbacks are walked in order. Each fallback in the chain is attempted once. Only the primary provider uses retry_count. If all fallbacks fail, the gateway returns 502 all_providers_failed.

💡 Note: A fallback that names a different provider dispatches with that provider's key as stored on the gateway; a provider that requires no key (for example the platform's own local models) dispatches without one. A fallback whose provider has no stored key on the gateway is skipped with a logged warning and the chain moves on. A fallback that carries its own model is requested with exactly that model; a fallback that only names a provider re-requests the primary's model there.

model@provider shorthand

Anywhere a target takes a model, you may use the model@provider suffix form to set both the model and the provider in a single string. For example, "model": "gpt-4o@openai" is equivalent to {"provider": "openai", "model": "gpt-4o"}. This works in the top-level model, in each fallbacks entry, and in each load_balance target.

Load balancing

The load_balance action spreads matching requests across a set of weighted targets. It is the only action that carries per-target weight values.

Field Type Description
strategy string Selection strategy used to pick a target.
targets array Non-empty list of {"provider": "...", "model": "...", "weight": N} entries. A higher weight receives a proportionally larger share of traffic.
{
  "actions": {
    "load_balance": {
      "strategy": "weighted_random",
      "targets": [
        {"provider": "openai", "model": "gpt-4o", "weight": 3},
        {"provider": "anthropic", "model": "claude-sonnet-4-6", "weight": 1}
      ]
    }
  }
}

Validation

POST and PATCH validate the request body at the trust boundary and reject a malformed rule with 400 Bad Request and a descriptive error message. Nothing is persisted on rejection. The rules are:

conditions

Accepted Rejected → 400
Omitted, OR a JSON array of condition objects. An empty array ([]) is a valid catch-all. A JSON null, a scalar (number/string/boolean), or a non-array object (e.g. {"field":"model"} without the surrounding [...]).
Each element is an object with a non-empty string field. op and value, when present, must be strings. An element that is not an object (e.g. null or a number), a missing/empty/non-string field, or a non-string op/value.

actions

Accepted Rejected → 400
Omitted, OR a JSON object. provider/model are strings; an unknown provider or a model with more than one @ separator is rejected. A JSON null or a scalar.
fallbacks/load_balance, when present, must be an array / object respectively; load_balance.targets must be a non-empty array. A present-but-non-object load_balance (e.g. "load_balance": null) or non-array fallbacks; an empty or missing load_balance.targets.
load_balance.strategy, when present, must be a string; load_balance.sticky, when present, must be an object with a non-empty string field. A non-string strategy, a non-object sticky, or a missing/empty/non-string sticky.field.
timeout_ms, when present, must be a positive integer no greater than 3600000 (1 hour). A JSON null, a non-number, a fractional number, zero or negative, or a value above the ceiling. A stored 0 would disable the upstream timeout entirely (block indefinitely), and an out-of-range value would pin a worker — so the write is rejected rather than coerced.

enabled

Accepted Rejected → 400
Omitted, OR a JSON boolean (true/false). A JSON null, a number (including 0), a string (including "false"), or an object. Any of these is truthy internally and would silently store the rule as enabled — so the write is rejected rather than coerced.

💡 Note: These rules exist so a stored rule can never hold a value that would fail a later read. A rejected write is always safer than a persisted rule that breaks routing for the whole gateway. A weight that is absent or non-numeric on a load_balance target defaults to 1.

Response codes

Method Success Other
POST 201 Created with {"id": "<rule_id>"}. 400 invalid body (see Validation); 500 if the rule cannot be persisted — the response never reports success for a rule that was not stored.
PATCH 200 OK with {"ok": true}. PATCH is a merge — a field you omit keeps its stored value; only the fields you send are changed. (Sending {"enabled": false} alone disables the rule without touching its priority, conditions, or actions.) 400 invalid body; 404 Not Found if the rule_id does not exist or belongs to a different gateway; 500 if the update cannot be persisted.
DELETE 200 OK with {"ok": true}. 404 Not Found if the rule_id does not exist or belongs to a different gateway; 500 on a delete failure.

A 201/200 is returned only after the write is confirmed committed; a storage failure surfaces as 500, never as a fake success carrying an id for a rule that does not exist.


Examples

⭐ Example: The following examples show complete API requests for common rule management operations.

Listing rules

curl https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules

Response: a bare array of rule objects.

[
  {
    "id": "rule_abc123",
    "priority": 10,
    "conditions": [{"field": "model", "op": "prefix", "value": "gpt-"}],
    "actions": {
      "provider": "openai",
      "model": "gpt-4o",
      "fallbacks": [{"provider": "anthropic", "model": "claude-sonnet-4-6"}]
    },
    "enabled": true
  }
]

In every returned rule, conditions is always a JSON array and actions is always a JSON object. If a rule's stored value is corrupt (for example a literal null), conditions is served as an empty array [] and actions as an empty object {} — never null — so clients can iterate conditions and read actions fields without a null check.

Creating a rule — redirect GPT requests to OpenAI with Anthropic fallback

curl -X POST https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules \
  -H "Content-Type: application/json" \
  -d '{
    "priority": 10,
    "conditions": [
      {"field": "model", "op": "prefix", "value": "gpt-"}
    ],
    "actions": {
      "provider": "openai",
      "model": "gpt-4o",
      "fallbacks": [
        {"provider": "anthropic", "model": "claude-sonnet-4-6"}
      ]
    },
    "enabled": true
  }'

Creating a rule — route production traffic to a specific model

curl -X POST https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules \
  -H "Content-Type: application/json" \
  -d '{
    "priority": 5,
    "conditions": [
      {"field": "meta:env", "op": "eq", "value": "production"},
      {"field": "model", "op": "contains", "value": "claude"}
    ],
    "actions": {
      "provider": "anthropic",
      "model": "claude-opus-4-6"
    },
    "enabled": true
  }'

Creating a catch-all rule — default provider for all unmatched requests

curl -X POST https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules \
  -H "Content-Type: application/json" \
  -d '{
    "priority": 0,
    "conditions": [],
    "actions": {
      "provider": "openrouter",
      "fallbacks": []
    },
    "enabled": true
  }'

Updating a rule — disable without deleting

curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/{rule_id} \
  -H "Content-Type: application/json" \
  -d '{"enabled": false}'

Updating a rule — change priority

curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/{rule_id} \
  -H "Content-Type: application/json" \
  -d '{"priority": 20}'

Deleting a rule

curl -X DELETE https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/{rule_id}

Reordering rules

PUT /gateways/{id}/rules/order sets the evaluation order for all of a gateway's rules in one call. The body is the full list of the gateway's rule ids, first evaluated first:

curl -X PUT https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/order \
  -H "Content-Type: application/json" \
  -d '{"order": ["<rule_id_a>", "<rule_id_b>", "<rule_id_c>"]}'

The gateway assigns strictly-descending priorities in a single transaction, so the sent order becomes the evaluation order. order must be a permutation of exactly the gateway's current rule ids — a partial list, an unknown id, a rule id belonging to a different gateway, a duplicate, or a non-array is rejected with 400 and no priority is changed (a bad reorder can never gap or duplicate priorities). Returns 200 OK {"ok": true} on success. The admin UI's Routing Rules table drives this endpoint with per-row move-up / move-down controls.


Evaluation order

Rules are evaluated in descending priority order — the highest priority number is checked first, ties broken by rule id so the order is deterministic (and identical to what GET /rules lists). The first matching rule wins — subsequent rules are not evaluated. A last-resort catch-all therefore needs a low priority so every more specific rule is tried before it. Use PUT /rules/order (above) to set the order without hand-picking priority numbers.

If no rule matches the request, the provider and model from the original request URL or body are used as-is.

💡 Note: Rules are cached in-process for 30 seconds. Changes take effect within one cache TTL cycle. In a multi-worker deployment, each worker refreshes independently.


See also