Routing rules API
Routing rules let you rewrite the provider and model for any request without changing the caller. Rules are evaluated in priority order — the first matching rule wins. Each rule can override the provider, rewrite the model name, and attach a fallback chain that is walked when the primary fails.
Base URL: https://<your-gateway-host>/admin/v1
Endpoints
| Method | Path | Description |
|---|---|---|
GET |
/gateways/{id}/rules |
List rules for a gateway, in evaluation order |
POST |
/gateways/{id}/rules |
Create a rule |
PATCH |
/gateways/{id}/rules/{rule_id} |
Update a rule |
PUT |
/gateways/{id}/rules/order |
Reorder all rules for a gateway |
DELETE |
/gateways/{id}/rules/{rule_id} |
Delete a rule |
Rule structure
{
"id": "rule_abc123",
"priority": 10,
"conditions": [
{"field": "model", "op": "prefix", "value": "gpt-"}
],
"actions": {
"provider": "openai",
"model": "gpt-4o",
"fallbacks": [
{"provider": "anthropic", "model": "claude-sonnet-4-6"}
]
},
"enabled": true
}
Rule fields
| Field | Type | Description |
|---|---|---|
priority |
integer | Evaluation order. Higher numbers run first. Rules with equal priority are ordered by rule id (a random UUID) — deterministic, and identical to the order GET /rules returns, but not the order the rules were created in. |
conditions |
array | List of condition objects. All conditions must match (logical AND). An empty array matches every request. |
actions |
object | What to do when the rule matches. Specify a provider, a model rewrite, or both — a model-only rewrite that keeps the original provider is valid. |
enabled |
boolean | false disables the rule without deleting it. Default: true. |
Conditions
Each condition object specifies a field, an op (operator), and a value to compare against.
Condition fields
| Field | Description |
|---|---|
model |
The model name from the request body, after normalisation: the provider prefix is stripped, and a retired Myra-fleet id with a registered successor (e.g. qwen3.6-27b) has already been upgraded to that successor, so a condition keyed on a retired id no longer matches — see Retired model ids. |
provider |
The provider from the request URL path (e.g. openai, compat). |
tenant_id |
The tenant identifier resolved from the URL. |
header:{name} |
The value of HTTP header {name} (e.g. header:x-customer-tier). |
meta:{key} |
Any custom metadata field attached via x-aig-meta-{key} request header. For example, meta:env matches the value of x-aig-meta-env. |
Condition operators
| Operator | Behaviour |
|---|---|
eq |
Exact string match. Case-sensitive. |
neq |
Exact string non-match. Case-sensitive. |
prefix |
Value starts with the given string. |
contains |
Value contains the given substring. |
regex |
Value matches the given pattern. Patterns use Lua-pattern syntax, not PCRE. An invalid pattern never matches. |
If a condition omits op, it defaults to eq (exact, case-sensitive match).
Examples
⭐ Example: The following examples show condition objects for common matching scenarios.
Match all requests for models starting with gpt-:
Match a specific model:
Match requests tagged with a custom header (x-aig-meta-env: production):
Match all requests (catch-all rule, empty conditions):
Actions
| Field | Type | Description |
|---|---|---|
provider |
string | Override the inference provider (e.g. openai, anthropic, gemini). |
model |
string | Rewrite the model name sent to the provider. If omitted, the original model name is used. A rewrite onto a retired Myra-fleet id with a registered successor is upgraded to that successor exactly like a client's own request; the response then carries X-AIG-Model-Upgraded-From (but no Deprecation header — the client did not name the retired id). Each fallbacks[].model is upgraded the same way, logged only: the response describes the model that served the turn. A rewrite onto any other deprecated id is refused at serve time (400 model_not_found). See Retired model ids. |
fallbacks |
array | Ordered fallback chain. Each entry is {"provider": "...", "model": "..."}. Used when the primary fails after all retries. |
load_balance |
object | Distribute matching requests across several targets by weight. See below. |
timeout_ms |
integer | Per-rule upstream timeout in milliseconds for the leg this rule routes — including a load-balanced target and each fallback leg. Bounds every wait on the upstream: the connection, the request send, and — for both streamed and buffered responses — the time to the first response byte and every inter-chunk gap of a streamed response. When set, it replaces the built-in streaming stall budgets (20 s first byte / 300 s inter-chunk) in both directions: a tighter value cuts a stalled stream sooner, a wider value grants a slow model more time to its first byte. It does not cap the total duration of a stream that is actively delivering tokens. A stall past the bound mid-stream ends the stream with the standard stream-error surface; a stall before the first byte fails the leg over to the fallback chain. If omitted, the gateway (or global) default timeout bounds connect/send and the buffered read, and the built-in stall budgets bound streamed reads. Must be a positive integer no greater than 3600000 (1 hour). |
💡 Note: Fallbacks are walked in order. Each fallback in the chain is attempted once. Only the primary provider uses
retry_count. If all fallbacks fail, the gateway returns502 all_providers_failed.💡 Note: A fallback that names a different provider dispatches with that provider's key as stored on the gateway; a provider that requires no key (for example the platform's own local models) dispatches without one. A fallback whose provider has no stored key on the gateway is skipped with a logged warning and the chain moves on. A fallback that carries its own
modelis requested with exactly that model; a fallback that only names aproviderre-requests the primary's model there.
model@provider shorthand
Anywhere a target takes a model, you may use the model@provider suffix form to set both the model and the provider in a single string. For example, "model": "gpt-4o@openai" is equivalent to {"provider": "openai", "model": "gpt-4o"}. This works in the top-level model, in each fallbacks entry, and in each load_balance target.
Load balancing
The load_balance action spreads matching requests across a set of weighted targets. It is the only action that carries per-target weight values.
| Field | Type | Description |
|---|---|---|
strategy |
string | Selection strategy used to pick a target. |
targets |
array | Non-empty list of {"provider": "...", "model": "...", "weight": N} entries. A higher weight receives a proportionally larger share of traffic. |
{
"actions": {
"load_balance": {
"strategy": "weighted_random",
"targets": [
{"provider": "openai", "model": "gpt-4o", "weight": 3},
{"provider": "anthropic", "model": "claude-sonnet-4-6", "weight": 1}
]
}
}
}
Validation
POST and PATCH validate the request body at the trust boundary and reject a malformed rule with 400 Bad Request and a descriptive error message. Nothing is persisted on rejection. The rules are:
conditions
| Accepted | Rejected → 400 |
|---|---|
Omitted, OR a JSON array of condition objects. An empty array ([]) is a valid catch-all. |
A JSON null, a scalar (number/string/boolean), or a non-array object (e.g. {"field":"model"} without the surrounding [...]). |
Each element is an object with a non-empty string field. op and value, when present, must be strings. |
An element that is not an object (e.g. null or a number), a missing/empty/non-string field, or a non-string op/value. |
actions
| Accepted | Rejected → 400 |
|---|---|
Omitted, OR a JSON object. provider/model are strings; an unknown provider or a model with more than one @ separator is rejected. |
A JSON null or a scalar. |
fallbacks/load_balance, when present, must be an array / object respectively; load_balance.targets must be a non-empty array. |
A present-but-non-object load_balance (e.g. "load_balance": null) or non-array fallbacks; an empty or missing load_balance.targets. |
load_balance.strategy, when present, must be a string; load_balance.sticky, when present, must be an object with a non-empty string field. |
A non-string strategy, a non-object sticky, or a missing/empty/non-string sticky.field. |
timeout_ms, when present, must be a positive integer no greater than 3600000 (1 hour). |
A JSON null, a non-number, a fractional number, zero or negative, or a value above the ceiling. A stored 0 would disable the upstream timeout entirely (block indefinitely), and an out-of-range value would pin a worker — so the write is rejected rather than coerced. |
enabled
| Accepted | Rejected → 400 |
|---|---|
Omitted, OR a JSON boolean (true/false). |
A JSON null, a number (including 0), a string (including "false"), or an object. Any of these is truthy internally and would silently store the rule as enabled — so the write is rejected rather than coerced. |
💡 Note: These rules exist so a stored rule can never hold a value that would fail a later read. A rejected write is always safer than a persisted rule that breaks routing for the whole gateway. A
weightthat is absent or non-numeric on aload_balancetarget defaults to1.
Response codes
| Method | Success | Other |
|---|---|---|
POST |
201 Created with {"id": "<rule_id>"}. |
400 invalid body (see Validation); 500 if the rule cannot be persisted — the response never reports success for a rule that was not stored. |
PATCH |
200 OK with {"ok": true}. PATCH is a merge — a field you omit keeps its stored value; only the fields you send are changed. (Sending {"enabled": false} alone disables the rule without touching its priority, conditions, or actions.) |
400 invalid body; 404 Not Found if the rule_id does not exist or belongs to a different gateway; 500 if the update cannot be persisted. |
DELETE |
200 OK with {"ok": true}. |
404 Not Found if the rule_id does not exist or belongs to a different gateway; 500 on a delete failure. |
A 201/200 is returned only after the write is confirmed committed; a storage failure surfaces as 500, never as a fake success carrying an id for a rule that does not exist.
Examples
⭐ Example: The following examples show complete API requests for common rule management operations.
Listing rules
Response: a bare array of rule objects.
[
{
"id": "rule_abc123",
"priority": 10,
"conditions": [{"field": "model", "op": "prefix", "value": "gpt-"}],
"actions": {
"provider": "openai",
"model": "gpt-4o",
"fallbacks": [{"provider": "anthropic", "model": "claude-sonnet-4-6"}]
},
"enabled": true
}
]
In every returned rule, conditions is always a JSON array and actions is always a
JSON object. If a rule's stored value is corrupt (for example a literal null),
conditions is served as an empty array [] and actions as an empty object {} —
never null — so clients can iterate conditions and read actions fields without a
null check.
Creating a rule — redirect GPT requests to OpenAI with Anthropic fallback
curl -X POST https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules \
-H "Content-Type: application/json" \
-d '{
"priority": 10,
"conditions": [
{"field": "model", "op": "prefix", "value": "gpt-"}
],
"actions": {
"provider": "openai",
"model": "gpt-4o",
"fallbacks": [
{"provider": "anthropic", "model": "claude-sonnet-4-6"}
]
},
"enabled": true
}'
Creating a rule — route production traffic to a specific model
curl -X POST https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules \
-H "Content-Type: application/json" \
-d '{
"priority": 5,
"conditions": [
{"field": "meta:env", "op": "eq", "value": "production"},
{"field": "model", "op": "contains", "value": "claude"}
],
"actions": {
"provider": "anthropic",
"model": "claude-opus-4-6"
},
"enabled": true
}'
Creating a catch-all rule — default provider for all unmatched requests
curl -X POST https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules \
-H "Content-Type: application/json" \
-d '{
"priority": 0,
"conditions": [],
"actions": {
"provider": "openrouter",
"fallbacks": []
},
"enabled": true
}'
Updating a rule — disable without deleting
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/{rule_id} \
-H "Content-Type: application/json" \
-d '{"enabled": false}'
Updating a rule — change priority
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/{rule_id} \
-H "Content-Type: application/json" \
-d '{"priority": 20}'
Deleting a rule
Reordering rules
PUT /gateways/{id}/rules/order sets the evaluation order for all of a gateway's rules in one call. The body is the full list of the gateway's rule ids, first evaluated first:
curl -X PUT https://<your-gateway-host>/admin/v1/gateways/{gateway_id}/rules/order \
-H "Content-Type: application/json" \
-d '{"order": ["<rule_id_a>", "<rule_id_b>", "<rule_id_c>"]}'
The gateway assigns strictly-descending priorities in a single transaction, so the sent order becomes the evaluation order. order must be a permutation of exactly the gateway's current rule ids — a partial list, an unknown id, a rule id belonging to a different gateway, a duplicate, or a non-array is rejected with 400 and no priority is changed (a bad reorder can never gap or duplicate priorities). Returns 200 OK {"ok": true} on success. The admin UI's Routing Rules table drives this endpoint with per-row move-up / move-down controls.
Evaluation order
Rules are evaluated in descending priority order — the highest priority number is checked first, ties broken by rule id so the order is deterministic (and identical to what GET /rules lists). The first matching rule wins — subsequent rules are not evaluated. A last-resort catch-all therefore needs a low priority so every more specific rule is tried before it. Use PUT /rules/order (above) to set the order without hand-picking priority numbers.
If no rule matches the request, the provider and model from the original request URL or body are used as-is.
💡 Note: Rules are cached in-process for 30 seconds. Changes take effect within one cache TTL cycle. In a multi-worker deployment, each worker refreshes independently.