Models & pricing API
The Models API exposes the model catalogue of the gateway. The Pricing API lets you read and manage per-model cost data used for spend tracking and budget enforcement. Model pricing data is maintained by Myra Security.
Base URL: https://<your-gateway-host>/admin/v1
Endpoints
| Method | Path | Description |
|---|---|---|
GET |
/models |
List the model catalogue |
GET |
/providers |
List supported provider integrations with metadata |
GET |
/providers/health |
Per-provider configured status and live status-page health |
GET |
/model-prices |
List all stored model prices |
GET |
/model-prices/self-hosted |
List the self-hosted (Myra) chat model prices shown in the Providers admin UI (platform admin only) |
PUT |
/model-prices |
Upsert a model price (platform admin only) |
GET |
/model-prices/markup |
List external models with their region tag + markup % (platform admin only) |
PUT |
/model-prices/markup |
Set an external model's region tag + markup % (platform admin only) |
POST |
/model-prices/sync |
Synchronise model prices from upstream sources (admin only) |
DELETE |
/model-prices/{provider}/{model} |
Delete a model price (platform admin only) |
POST |
/model-prices/{provider}/{model}/deprecate |
Hide a model from the picker (admin only) |
POST |
/model-prices/{provider}/{model}/un-deprecate |
Restore a previously deprecated model (admin only) |
GET |
/model-prices/active-for-probing |
List active (provider, model) rows for the health prober (admin only) |
POST |
/model-prices/{provider}/{model}/probe-result |
Record one health-probe outcome for a model (admin only) |
GET /models
Returns the known model catalogue. This list is used for model picker dropdowns in the admin UI. This endpoint is open — no authentication is required.
The list is pruned for picking, not complete: deprecated rows and non-chat modalities (embedding, rerank, transcription, image) are dropped, near-duplicate ids are filtered out (dated checkpoints such as -20251001, :free / :beta / :batch routing shelves, quantisation variants), and the remainder is collapsed to one canonical row per display name. The full table is available through GET /model-prices.
A model your plan entitles is never pruned. When the request carries a session whose workspace is on a self-serve plan, every model id that plan explicitly entitles is kept — even a dated snapshot that the filters would otherwise remove, and even when a shorter alias shares its display name. A tier can therefore pin a specific model snapshot and its users can still select it. Callers on other plans, and unauthenticated callers, get the ordinary pruned list.
Filter by provider:
Response:
A bare JSON array of model rows.
[
{
"provider": "openai",
"model": "gpt-4o",
"display_name": "GPT-4o",
"hosted_in": "US",
"display_rank": 0,
"max_input_tokens": 128000,
"max_output_tokens": 16384,
"input_per_1k": 0.005,
"output_per_1k": 0.015,
"cache_write_per_1k": null,
"cache_read_per_1k": null,
"cache_write_1h_per_1k": null,
"supports_thinking": false,
"supports_vision": true,
"supports_pdf_input": false,
"supports_files_api": false
},
{
"provider": "anthropic",
"model": "claude-opus-4-6",
"display_name": "Opus 4.6",
"hosted_in": "US",
"display_rank": 0,
"max_input_tokens": 200000,
"max_output_tokens": 64000,
"input_per_1k": 0.005,
"output_per_1k": 0.025,
"cache_write_per_1k": 0.00625,
"cache_read_per_1k": 0.0005,
"cache_write_1h_per_1k": 0.01,
"supports_thinking": true,
"thinking_always_on": false,
"supports_vision": true,
"supports_pdf_input": true,
"supports_files_api": true
}
]
Tagline resolution (tagline, tagline_de, tagline_source)
On this route the tagline / tagline_de of each row are resolved for the caller, and every row
carries tagline_source — where the served text came from:
tagline_source |
Meaning |
|---|---|
plan_copy |
Written for the caller's offered model set: a short line that names a property true only for this model within that set (e.g. "Uses the least trial credit", "Handles the longest texts", "Only one here that reads PDFs directly", "A model by Mistral AI"). |
catalog |
The model's global catalog tagline (model_price.tagline*, the same text GET /model-prices shows). |
name |
No tagline — show the display name only. tagline and tagline_de are absent. |
When plan copy applies. Only when all of these hold; otherwise every row keeps today's catalog
tagline (catalog) or none (name):
- the session's workspace is on a self-serve plan (trial or paid). Manual / enterprise workspaces and unauthenticated callers are unaffected;
- the platform setting
plan_model_copy_enabledistrue(default off — v1 ships dark until a quality fact source is licensed or measured); - the call carries no
?provider=filter (a filtered list is a subset, so relative copy would be judged against the wrong set); - the offered set has at least two models. The offered set is the listed rows the workspace can actually use: the plan's model list intersected with the EU-Gov add-on gate — the same decision the gateway enforces when a request is sent.
What an offered model shows. The copy is produced by a deterministic generator from facts only
(price order on both input and output price, context window, native image / tool support (never PDF — the workspace extracts PDFs itself), EU
hosting, the model's maker, and a measured speed lead over the last complete ISO week — at least 200
answers per model, a runner-up whose median time-to-first-token is at least 20 % slower, and no
lower throughput). A capability "no" counts only where it is certain — the capability catalog says so
explicitly; an unlisted or
unreadable flag blocks the matching "only one here that …" line for the whole set. It never
uses age words ("older", "newer") or quality words ("best", "smartest"), never mentions a capability
every model of the set shares, and never gives two models of a set the same line. The house rules are
in docs/internal/model-copy-house-style.md.
- A model with a distinguishing fact gets its generated line (
plan_copy). - A model no fact distinguishes from the others (for example a second model by the same maker that
is neither the cheapest, the largest, nor the only one with a capability) gets
name— no line is better than a line that is not true for it. - While no generated copy is available for the set (the telemetry could not be read for this request,
another request is computing it at that moment, or the set's generation was rejected), a model shows
its hand-written seed line for that set if one exists and the set has never had generated copy
(
plan_copy; the seed only bridges the rollout), else its catalog tagline if that text passes the same house rules and makes no superlative ("fastest", "cheapest", "only", "günstigste", …), comparative ("more", "than", "mehr"), version number or other-model-name claim (catalog), elsename. If the copy table itself cannot be read, the seed lines (stored there) are skipped: checked catalog tagline, elsename. An unexpected internal fault while resolving has the same outcome for every row that does not already carry plan copy. - The offered set holds one row per model id — the first row the list returns for that id; another provider's row for the same id keeps its catalog tagline.
Copy is generated lazily on the first request that sees a new combination of facts (no timer), stored,
and reused; the same facts never regenerate. tagline_source is emitted on GET /models only — the
/model-prices admin feeds carry the raw catalog columns. A repeated ?provider= parameter is refused
with 400 (provider must be a single string).
Each row also carries the model's native capability flags, resolved server-side so the client never hardcodes model lists:
supports_thinking— the model exposes extended-thinking/reasoning output. For a bareclaude-*id this is resolved by the same server-side resolver the inference path uses (core.model_identity), so the capability a client is offered and the request the gateway actually makes agree by construction. That resolver reads two sources in precedence order: the gateway's curated capability registry (an override), and — when the registry is silent for that id — a provider-synced flag imported from the first-party Anthropic Models API (see Capability sources & the trust stamp below). A bare anthropic-cloudclaude-*id the registry does not enumerate falls back to a name heuristic, which is broader: it can reporttruefor a model whose thinking parameter the gateway will not send. Treat it as "offer the control", not as a guarantee.- Azure- and Bedrock-hosted Claude (
azure_ai/claude-*; Bedrock ids such asanthropic.claude-…,eu.anthropic.claude-…,bedrock/<region>/…, and bareclaude-…-vN:M) are resolved only from the registry, in their own per-platform key space — never via the anthropic-cloud name heuristic. What Azure-hosted and Bedrock-hosted Claude accept (thinking shape, real output ceiling, native-document types) is a platform fact, not inheritable from the anthropic-cloud entry, and each must be verified against the real platform before an entry is added. Until then these ids are fail-closed:supports_thinking(andthinking_always_on) reportfalse— matching the backend, which likewise sends nothinkingparameter — so a client is never offered an Adaptive-thinking toggle the platform will silently ignore. thinking_always_on— the model reasons by default and cannot be turned off. Some always-adaptive Claude models (e.g.claude-fable-5) run extended thinking whenever thethinkingparameter is omitted and reject an explicit disable, sosupports_thinkingalone (which only says "thinking is available") cannot express this. Whentrue, a client should render the thinking control as locked on rather than a toggle that pretends "off" works. Fail-closed (absent →false, i.e. a normal on/off toggle). This is distinct from a model likeclaude-opus-5, which also reasons by default but can be disabled — for it the gateway emits an explicitthinking = {"type": "disabled"}when the client turns thinking off, and this flag staysfalse.supports_vision— accepts native image input (image blocks); whenfalse, attached images are extracted to text server-side instead.supports_pdf_input— accepts a native PDF document block; whenfalse, PDFs are extracted to text.supports_files_api— supports the provider Files-API upload + document skill (Anthropic).
supports_pdf_input and supports_files_api are additionally constrained by the route the row
would actually take: a model is only reported as accepting a native document block when the resolved
provider's request path accepts one (Anthropic always; Bedrock only for inline base64 on an
Anthropic-family model). The same Claude model reached through a re-selling provider therefore
reports false and its documents are extracted to text — which is what that route supports — rather
than advertising a native block the gateway would reject.
- supports_function_calling — the model can use gateway-injected function tools (web search, URL
fetch, connectors, sub-agents, knowledge search). Unlike the flags above this is permissive on the
unknown space: it is true for every model except one curated as tool-incapable (e.g.
Perplexity Sonar, which exposes no function-tools endpoint). An uncatalogued model reports true, so
the client never over-blocks a tool-capable model whose flag it doesn't know; a false here is the
authoritative signal the agent editor uses to disable all tool toggles. Sourced from the capability
registry (the same one the runtime tool-strip reads), not the sparse model_price DB column.
Every flag above except supports_function_calling is a boolean that fails closed (absent →
false); supports_function_calling is the sole fail-open flag (absent → true, tool-capable). supports_vision is resolved by a single shared function
(capability.effective_vision) that the gateway ALSO uses when it decides whether to forward image
pixels or extract them to text — so the value in this payload and the gateway's runtime behaviour can
never disagree. It reports true when either:
- the model is affirmatively vision-capable in the curated capability catalog/registry, or
- the model id matches the curated vision name heuristic: a bare
claude-*id (covers Claude models not individually enumerated in the registry, e.g. a freshly releasedclaude-opus-4-6) or a*-vl-*id (Qwen-VL / ERNIE-VL / Nemotron-VL pass-through vision models). Matching is anchored and case-insensitive;azure_ai/claude-*is intentionally excluded (it reportsfalse).
Otherwise supports_vision is false — an uncurated pass-through model (for example x-ai/grok-4.x
on OpenRouter) whose vision capability the gateway cannot confirm. For those, attached images are
extracted to text server-side by the gateway before dispatch; it never forwards raw pixels to a
model it cannot confirm sees them, because a non-vision or picky upstream rejects an image block with
an opaque 400. To force native image input for a specific uncurated pass-through vision model, register
it in the capability catalog with supports_vision: true; it then receives pixels on both sides.
Native claude vision requires the bare model id (
claude-opus-4-6). A provider-prefixed id (anthropic/claude-…,azure_ai/claude-…) is not matched by the name heuristic and its images are text-extracted unless the model is curated in the catalog.
*-vl-*caveat: if a curated*-vl-*pass-through model enforces a per-request image cap upstream (e.g. "at most 1 image"), that per-model image cap must be registered in the model's capability entry; otherwise the gateway forwards every image and the upstream may400on the excess.
Capability sources & the trust stamp
Behavioural capabilities resolve from two sources, in strict precedence: the in-image capability registry (an override) first; when it is silent for a model, a provider-synced flag imported by the daily model importer. The registry is no longer the only source — but it always wins where it speaks, and the synced layer only ever fills a registry silence, never downgrades a curated value.
Only first-party Anthropic models carry a provider-synced source: the importer reads
GET /v1/models's capabilities tree (thinking.types.adaptive|enabled.supported,
context_management.compact_20260112.supported) and writes three model_price columns —
supports_reasoning, uses_adaptive_thinking, supports_native_compaction — stamped with a trust
column, capability_source. Bedrock/Vertex-hosted Claude, Gemini, and OpenRouter expose no such tree,
so they keep hand-curation and are never stamped.
capability_source — accepted values: the exact string provider, or SQL NULL. The resolver
trusts a synced flag ONLY when capability_source = 'provider' (a strict allowlist). Any other
value — NULL, empty string, litellm_json, catalog, or anything else — is treated as untrusted
and the synced columns are ignored (the model falls back to the registry, else OFF). A synced
supports_reasoning additionally takes effect only when the provider also declared a determinate
thinking shape (uses_adaptive_thinking is a real 0/1, never NULL) and a
provider-sourced output cap large enough to leave room for a visible answer; otherwise reasoning stays
OFF. This is enforced DB-side: capability_source is only ever written to provider by the
first-party Anthropic importer (community/litellm/catalog sync paths physically cannot write it), and
uses_adaptive_thinking is tri-state so an undetermined shape can never trigger a hard 400.
A synced capability flip (on↔off) raises an operator alert, and a routable Anthropic model with no
resolvable capability record at all (neither registry nor a provider stamp) is flagged by a standing
drift guard.
A native document block ({"type":"document","source":{...}}, used for native PDF / Files-API
attachments) is an Anthropic-wire shape: only providers whose request serializer emits the Anthropic
Messages wire carry it to a backend that accepts it — Anthropic, and AWS Bedrock for an
anthropic-family model. Every OpenAI-compatible provider forwards it raw and the upstream rejects it
with an opaque 400. The gateway therefore fails closed on the provider wire-shape (not on
supports_pdf_input): a native document sent to any other provider/model is rejected before dispatch
with model_capability_mismatch (HTTP 400) rather than forwarded — pick a document-capable model
(e.g. Claude) or remove the attachment. The Anthropic Files-API (source.type: "file") is
Anthropic-cloud-only and is not accepted via Bedrock. (Clients should extract non-PDF-capable
attachments to text before sending; the web app does this automatically and folds a history document to
a text reference when the selected model cannot read documents.)
GET /model-prices
Returns all stored model pricing records. Used internally for cost calculation on every inference request. Required role: admin or tenant_admin.
Response:
A bare JSON array of pricing rows.
[
{
"provider": "openai",
"model": "gpt-4o",
"input_per_1k": 0.005,
"output_per_1k": 0.015,
"cache_write_per_1k": null,
"cache_read_per_1k": null,
"cache_write_1h_per_1k": null,
"max_input_tokens": 128000,
"max_output_tokens": 16384,
"max_input_tokens_source": "litellm_json",
"deprecated_at": null,
"deprecated_source": null,
"display_name": "GPT-4o",
"tagline": null,
"tagline_de": null,
"metadata_source": "litellm",
"hosted_in": "US",
"updated_at": 1742551232
},
{
"provider": "anthropic",
"model": "claude-opus-4-6",
"input_per_1k": 0.005,
"output_per_1k": 0.025,
"cache_write_per_1k": 0.00625,
"cache_read_per_1k": 0.0005,
"cache_write_1h_per_1k": 0.01,
"max_input_tokens": 200000,
"max_output_tokens": 64000,
"max_input_tokens_source": "provider",
"deprecated_at": null,
"deprecated_source": null,
"display_name": "Opus 4.6",
"tagline": null,
"tagline_de": null,
"metadata_source": "manual",
"hosted_in": "US",
"updated_at": 1742551232
}
]
ModelPrice fields
| Field | Type | Description |
|---|---|---|
provider |
string | Provider identifier (e.g. openai, anthropic). |
model |
string | Exact model name as used in requests. |
input_per_1k |
number | Cost in USD per 1,000 input (prompt) tokens. |
output_per_1k |
number | Cost in USD per 1,000 output (completion) tokens. |
cache_write_per_1k |
number | null | Cost per 1,000 tokens written to provider prompt cache, 5-minute TTL (Anthropic). null if not applicable. |
cache_read_per_1k |
number | null | Cost per 1,000 tokens read from provider prompt cache (Anthropic). null if not applicable. |
cache_write_1h_per_1k |
number | null | Cost per 1,000 tokens written to provider prompt cache, 1-hour TTL (Anthropic). null if not applicable. |
max_input_tokens |
integer | null | Maximum input (context) window in tokens. null if unknown. |
max_output_tokens |
integer | null | Maximum output tokens the model can generate. null if unknown. |
max_input_tokens_source |
string | null | Provenance of the token-limit figures (max_input_tokens and max_output_tokens): provider (the provider's own model API, e.g. Anthropic GET /v1/models — authoritative), catalog (a self-hosted Myra pod's own configuration), or litellm_json (the community LiteLLM snapshot — a hint). A provider value is authoritative and is never overwritten by the LiteLLM hint; a routed alias (e.g. claude-opus-4-5) inherits the newest provider snapshot's caps. null for legacy rows whose source was not recorded. |
deprecated_at |
integer | null | Unix timestamp (seconds) when the model was deprecated, or null if active. |
deprecated_source |
string | null | Who deprecated the row: admin (an operator, through the deprecate endpoint), prober (the health prober's automatic flip after repeated not-found results), or myra_sync / openrouter_sync (a catalog reconciler, because the provider stopped offering the model). A reconciler clears only its own deprecation when the model reappears, so an operator's decision is never silently undone; an operator's deprecate call takes over a reconciler's, keeping the original timestamp. null for rows deprecated before provenance was recorded — those are likewise never cleared automatically. |
display_name |
string | null | Human-friendly model name shown in the picker. Derived server-side by a brand-stripped normaliser: the vendor prefix is dropped and the tier + version kept (claude-opus-4-8 → Opus 4.8, claude-opus-5 → Opus 5, claude-fable-5-1 → Fable 5.1). Any Claude tier is handled (opus/sonnet/haiku/fable/…), not a fixed list. A manual row (metadata_source = "manual") is never overwritten. |
tagline |
string | null | Short English description of the model, positioned by recency decided server-side: the newest model in a family reads as capable, and an older sibling is framed plainly as older (e.g. once a newer one is added). This self-heals — an older model's tagline is regenerated to an "older" framing when a newer sibling appears, and the newest never reads "older". manual rows are left untouched. |
tagline_de |
string | null | Short German description of the model (co-generated with tagline, same recency positioning). |
tagline_source |
string | GET /models only: plan_copy, catalog or name — see Tagline resolution above. On GET /models the tagline / tagline_de values are the resolved ones (for a self-serve workspace they can be plan copy instead of the catalog columns). |
metadata_source |
string | null | Where the row's metadata came from (e.g. litellm, openrouter, manual). |
hosted_in |
string | null | Region/jurisdiction the model is hosted in: "Myra" (self-hosted), "EU", "US", "China", or null when no rule applies (e.g. meta-routers). Derived server-side; drives the region indicator in the chat model picker. |
display_rank |
integer | null | Per-provider display order for the chat model picker, computed server-side on read: models are grouped into families and ordered by tier (a family's highest price) then newest version first, so 0 is shown first within its provider group. This makes a newer-but-cheaper flagship (e.g. Opus 5.5) sort above a pricier older sibling (Opus 4.8) — which a raw price sort gets wrong. Advisory display metadata only, never an entitlement or routing signal. May be null/absent on an older response; clients MUST tolerate its absence and fall back to their own ordering. Present on GET /models; not emitted on the /model-prices admin feeds. |
updated_at |
integer | null | Unix timestamp (seconds) of the last update to the row. |
PUT /model-prices
Required role: platform admin (the MODEL_PRICES_MANAGE permission). A tenant_admin may read model prices but not change them.
Create or update a price for a model. The (provider, model) pair is the unique key. If a record already exists it is replaced.
⭐ Example: The following examples show how to upsert pricing for different provider and model combinations.
curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
-H "Content-Type: application/json" \
-d '{
"provider": "openai",
"model": "gpt-4o",
"input_per_1k": 0.005,
"output_per_1k": 0.015
}'
With prompt cache pricing (Anthropic):
curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
-H "Content-Type: application/json" \
-d '{
"provider": "anthropic",
"model": "claude-opus-4-6",
"input_per_1k": 0.005,
"output_per_1k": 0.025,
"cache_write_per_1k": 0.00625,
"cache_read_per_1k": 0.0005
}'
Accepted body
| Field | Required | Accepted |
|---|---|---|
provider, model |
yes | Non-empty strings. model must be a well-formed model id (no empty final path segment). |
input_per_1k, output_per_1k |
yes | A finite number >= 0. |
cache_write_per_1k, cache_read_per_1k, cache_write_1h_per_1k |
no | A finite number >= 0, or null / omitted to clear. |
Anything else is rejected with 400 and nothing is written: a missing or null required
price, a non-numeric value (including a numeric string), NaN, Infinity, or a negative
price. Negative is rejected rather than stored because spend is only recorded for a request that
costs more than nothing — a negative price would switch metering off for the model entirely, the
same way a zero price does, and budgets, caps, budget-alert e-mails and trial credit would stop
accruing silently.
A price of exactly 0 is accepted: some self-hosted models are billed per a different unit and
are priced at zero on purpose. A zero-priced chat model does raise a daily platform alert —
see Provider Costs.
Custom internal model (e.g. a fine-tuned model on Azure):
curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
-H "Content-Type: application/json" \
-d '{
"provider": "azure",
"model": "my-ft-gpt4o-deployment",
"input_per_1k": 0.008,
"output_per_1k": 0.020
}'
GET /model-prices/markup
Required role: platform admin (the MODEL_PRICES_MANAGE permission). Not readable by a
tenant_admin — the markup is Myra's confidential charge-margin over raw provider cost and is
deliberately absent from the tenant-readable GET /model-prices and the picker GET /models.
Lists the active external model rows (provider other than myra) with their billing region tag
and markup percentage:
[
{ "provider": "anthropic", "model": "claude-sonnet-4-6", "display_name": "Sonnet 4.6",
"markup_region": null, "markup_pct": null },
{ "provider": "mistral", "model": "mistral/mistral-small-latest", "display_name": "Mistral Small",
"markup_region": "eu", "markup_pct": null }
]
markup_region is "eu", "non_eu", or null (unset → treated as non-EU +20 %). markup_pct is a
number in [0,100] or null (unset → the region default: EU 10 %, non-EU 20 %).
PUT /model-prices/markup
Required role: platform admin (the MODEL_PRICES_MANAGE permission); a tenant_admin receives
403.
curl -X PUT https://<your-gateway-host>/admin/v1/model-prices/markup \
-H 'Content-Type: application/json' \
-d '{ "provider": "mistral", "model": "mistral/mistral-small-latest",
"markup_region": "eu", "markup_pct": 10 }'
Accepted body
provider,model(required) — must identify an existing external model-price row. A model with no price row →404(a price-less external row would bill0, so markup is never set on one).markup_region—"eu"or"non_eu", ornull/absent to clear (→ non-EU default). Any other string →400, nothing written.markup_pct— a finite number in[0,100], ornull/absent to clear (→ region default). A non-number,NaN, infinity, negative, or> 100→400, nothing written. The value is never silently coerced.
The effective markup is applied to external-provider leg costs at cost-computation time, before the
budget_usd/wallet debit — see Cost attribution.
Self-hosted (myra) models are never marked up.
DELETE /model-prices/{provider}/{model}
Required role: platform admin (the MODEL_PRICES_MANAGE permission).
Remove a pricing record. Subsequent requests for this model fall back to the defaults of the hardcoded cost table of the gateway, or report zero cost if the model is not in the defaults.
💡 Note: URL-encode the model name if it contains slashes. For example,
accounts/fireworks/models/llama-v3-70bbecomesaccounts%2Ffireworks%2Fmodels%2Fllama-v3-70b.
POST /model-prices/{provider}/{model}/deprecate
Marks a model as deprecated, recording deprecated_source: "admin". This takes over a
deprecation a catalog reconciler made (the original deprecated_at is kept), which is what makes
an operator's decision durable: a reconciler only ever clears its own deprecation, so a model an
operator deprecated stays deprecated even when the provider starts offering it again. A deprecated model is dropped from the chat-model catalogue served by GET /models (the picker feed) and from the Auto default pick, and is refused at serve time: an inference request that resolves to a deprecated (provider, model) — whether named explicitly by the client, pinned by a project, or reached through a routing rule or fallback chain — is rejected before any upstream dispatch with 400 model_not_found (header X-AIG-Error: model_not_found), rather than being routed to the provider. Its pricing row is kept so historical cost figures still resolve. This is the supported way to disable a model. Required role: platform admin only.
One exception: a deprecated Myra-fleet id that has a registered successor (a rollover such as qwen3.6-27b → qwen3.8-27b) is upgraded to the successor instead of refused when a client names it explicitly — see Inference › Retired model ids. Un-deprecating the old row switches that off (the request then goes to the old id as named).
curl -X POST "https://<your-gateway-host>/admin/v1/model-prices/anthropic/claude-opus-4-6/deprecate"
Response
| Field | Type | Description |
|---|---|---|
ok |
boolean | Always true on success. |
flipped |
boolean | true if the model was active and is now deprecated; false if it was already deprecated (no change). |
💡 Note: URL-encode the model name if it contains slashes (e.g.
accounts/fireworks/models/llama-v3-70bbecomesaccounts%2Ffireworks%2Fmodels%2Fllama-v3-70b).
POST /model-prices/{provider}/{model}/un-deprecate
Clears the deprecation flag, returning the model to the picker feed and the Auto default pick. Required role: platform admin only.
curl -X POST "https://<your-gateway-host>/admin/v1/model-prices/anthropic/claude-opus-4-6/un-deprecate"
Response
Returns 400 { "error": "provider '<provider>' is retired from the catalog" } for providers retired from the catalogue (none currently), and 500 { "error": ... } on a storage failure.
GET /model-prices/active-for-probing
Lists the active (provider, model) rows (those with deprecated_at IS NULL) for the scheduled health prober to ping. Optionally filtered by ?provider=<name>. Required role: platform admin only.
Response
[ { "provider": "anthropic", "model": "claude-opus-4-6" }, { "provider": "gemini", "model": "gemini-2.5-pro" } ]
POST /model-prices/{provider}/{model}/probe-result
Records one health-probe outcome for a model (used by the scheduled deprecation prober). Required role: platform admin only.
Request body
| Field | Type | Description |
|---|---|---|
outcome |
string | Required. One of the accepted values below. Any other value is rejected with 400 { "error": "invalid outcome" } and nothing is recorded. |
http_status |
integer | null | The upstream HTTP status observed (optional). |
upstream_error |
string | null | Upstream error text, truncated to 255 chars (optional). |
duration_ms |
integer | null | Probe round-trip time in ms (optional; stored as 0 when absent or negative, capped at 2147483647; NaN counts as absent). |
consecutive_threshold |
integer | Consecutive-failure threshold for auto-deprecation (optional, default 3; accepted range 1..100 — larger values are clamped to 100, values below 1, non-numeric or NaN values fall back to the default). The probe log is kept for 30 days, so a threshold cannot exceed what the window can ever hold. |
Accepted outcome values (any other value → 400):
| Value | Meaning | Auto-deprecates? |
|---|---|---|
pass |
2xx — the model served the probe. | No |
fail_not_found |
4xx model-not-found / deprecated shape. | Yes — flips deprecated_at after consecutive_threshold consecutive fail_not_found outcomes. |
fail_quota |
429 / rate-limit / quota block. Record-only — a quota block is a transient route-around signal, not a deprecation, so it is logged for observability + human review but never auto-deprecates. | No |
fail_other |
Any other non-2xx (transient). | No |
driver_error |
The prober itself failed before reaching the upstream. | No |
fail_guided |
The hourly prober's guided-output leg (a tiny json_schema request against the in-house fleet) failed or hung — the vLLM guided-path / engine-death signature. Record-only — an engine crash is an ops incident, not a deprecation; also excluded from the flip window so interleaved rows can never mask a genuine fail_not_found streak. |
No |
deprecated_at is a per-(provider, model) row, so deprecating one provider's copy of a model never affects another provider's copy (e.g. deprecating gemini/gemini-2.5-pro leaves openrouter/google/gemini-2.5-pro untouched).
Response
| Field | Type | Description |
|---|---|---|
ok |
boolean | Always true on success. |
flipped |
boolean | true only when this fail_not_found outcome tipped the model into deprecation. |
POST /model-prices/sync
Synchronises pricing data from the configured upstream sources (LiteLLM JSON and OpenRouter API) into the local model_price table. Required role: platform admin only.
# Sync every provider
curl -X POST https://<your-gateway-host>/admin/v1/model-prices/sync
# Sync one provider
curl -X POST 'https://<your-gateway-host>/admin/v1/model-prices/sync?provider=openrouter'
The response is a per-provider summary of inserted, updated, and skipped records.
GET /providers
Returns the list of provider integrations supported by the gateway with metadata such as requires_key. Used by the admin UI to populate the provider drop-down in the Add Model dialog.
GET /providers/health
Returns one row per supported provider with these pieces of state merged together:
name— the provider identifier (e.g.openai,anthropic).requires_key— whether the provider requires a BYOK API key.configured—truewhen at least one BYOK key for the provider is stored on any gateway in the tenant;falsewhen keys are required but absent;nullwhen no key is needed.status/message/checked_at/has_status_page— the live availability state, polled from the public Atlassian Statuspage of the provider every five minutes. Providers without a Statuspage havehas_status_page = false.latency_ms— the latency of the last status-page poll, if recorded.