Skip to content

Models & pricing API

The Models API exposes the model catalogue of the gateway. The Pricing API lets you read and manage per-model cost data used for spend tracking and budget enforcement. Model pricing data is maintained by Myra Security.

Base URL: https://<your-gateway-host>/admin/v1


Endpoints

Method Path Description
GET /models List the model catalogue
GET /providers List supported provider integrations with metadata
GET /providers/health Per-provider configured status and live status-page health
GET /model-prices List all stored model prices
GET /model-prices/self-hosted List the self-hosted (Myra) chat model prices shown in the Providers admin UI (platform admin only)
PUT /model-prices Upsert a model price (platform admin only)
GET /model-prices/markup List external models with their region tag + markup % (platform admin only)
PUT /model-prices/markup Set an external model's region tag + markup % (platform admin only)
POST /model-prices/sync Synchronise model prices from upstream sources (admin only)
DELETE /model-prices/{provider}/{model} Delete a model price (platform admin only)
POST /model-prices/{provider}/{model}/deprecate Hide a model from the picker (admin only)
POST /model-prices/{provider}/{model}/un-deprecate Restore a previously deprecated model (admin only)
GET /model-prices/active-for-probing List active (provider, model) rows for the health prober (admin only)
POST /model-prices/{provider}/{model}/probe-result Record one health-probe outcome for a model (admin only)

GET /models

Returns the known model catalogue. This list is used for model picker dropdowns in the admin UI. This endpoint is open — no authentication is required.

The list is pruned for picking, not complete: deprecated rows and non-chat modalities (embedding, rerank, transcription, image) are dropped, near-duplicate ids are filtered out (dated checkpoints such as -20251001, :free / :beta / :batch routing shelves, quantisation variants), and the remainder is collapsed to one canonical row per display name. The full table is available through GET /model-prices.

A model your plan entitles is never pruned. When the request carries a session whose workspace is on a self-serve plan, every model id that plan explicitly entitles is kept — even a dated snapshot that the filters would otherwise remove, and even when a shorter alias shares its display name. A tier can therefore pin a specific model snapshot and its users can still select it. Callers on other plans, and unauthenticated callers, get the ordinary pruned list.

curl https://<your-gateway-host>/admin/v1/models

Filter by provider:

curl "https://<your-gateway-host>/admin/v1/models?provider=anthropic"

Response:

A bare JSON array of model rows.

[
  {
    "provider": "openai",
    "model": "gpt-4o",
    "display_name": "GPT-4o",
    "hosted_in": "US",
    "display_rank": 0,
    "max_input_tokens": 128000,
    "max_output_tokens": 16384,
    "input_per_1k": 0.005,
    "output_per_1k": 0.015,
    "cache_write_per_1k": null,
    "cache_read_per_1k": null,
    "cache_write_1h_per_1k": null,
    "supports_thinking": false,
    "supports_vision": true,
    "supports_pdf_input": false,
    "supports_files_api": false
  },
  {
    "provider": "anthropic",
    "model": "claude-opus-4-6",
    "display_name": "Opus 4.6",
    "hosted_in": "US",
    "display_rank": 0,
    "max_input_tokens": 200000,
    "max_output_tokens": 64000,
    "input_per_1k": 0.005,
    "output_per_1k": 0.025,
    "cache_write_per_1k": 0.00625,
    "cache_read_per_1k": 0.0005,
    "cache_write_1h_per_1k": 0.01,
    "supports_thinking": true,
    "thinking_always_on": false,
    "supports_vision": true,
    "supports_pdf_input": true,
    "supports_files_api": true
  }
]

Tagline resolution (tagline, tagline_de, tagline_source)

On this route the tagline / tagline_de of each row are resolved for the caller, and every row carries tagline_source — where the served text came from:

tagline_source Meaning
plan_copy Written for the caller's offered model set: a short line that names a property true only for this model within that set (e.g. "Uses the least trial credit", "Handles the longest texts", "Only one here that reads PDFs directly", "A model by Mistral AI").
catalog The model's global catalog tagline (model_price.tagline*, the same text GET /model-prices shows).
name No tagline — show the display name only. tagline and tagline_de are absent.

When plan copy applies. Only when all of these hold; otherwise every row keeps today's catalog tagline (catalog) or none (name):

  • the session's workspace is on a self-serve plan (trial or paid). Manual / enterprise workspaces and unauthenticated callers are unaffected;
  • the platform setting plan_model_copy_enabled is true (default off — v1 ships dark until a quality fact source is licensed or measured);
  • the call carries no ?provider= filter (a filtered list is a subset, so relative copy would be judged against the wrong set);
  • the offered set has at least two models. The offered set is the listed rows the workspace can actually use: the plan's model list intersected with the EU-Gov add-on gate — the same decision the gateway enforces when a request is sent.

What an offered model shows. The copy is produced by a deterministic generator from facts only (price order on both input and output price, context window, native image / tool support (never PDF — the workspace extracts PDFs itself), EU hosting, the model's maker, and a measured speed lead over the last complete ISO week — at least 200 answers per model, a runner-up whose median time-to-first-token is at least 20 % slower, and no lower throughput). A capability "no" counts only where it is certain — the capability catalog says so explicitly; an unlisted or unreadable flag blocks the matching "only one here that …" line for the whole set. It never uses age words ("older", "newer") or quality words ("best", "smartest"), never mentions a capability every model of the set shares, and never gives two models of a set the same line. The house rules are in docs/internal/model-copy-house-style.md.

  • A model with a distinguishing fact gets its generated line (plan_copy).
  • A model no fact distinguishes from the others (for example a second model by the same maker that is neither the cheapest, the largest, nor the only one with a capability) gets name — no line is better than a line that is not true for it.
  • While no generated copy is available for the set (the telemetry could not be read for this request, another request is computing it at that moment, or the set's generation was rejected), a model shows its hand-written seed line for that set if one exists and the set has never had generated copy (plan_copy; the seed only bridges the rollout), else its catalog tagline if that text passes the same house rules and makes no superlative ("fastest", "cheapest", "only", "günstigste", …), comparative ("more", "than", "mehr"), version number or other-model-name claim (catalog), else name. If the copy table itself cannot be read, the seed lines (stored there) are skipped: checked catalog tagline, else name. An unexpected internal fault while resolving has the same outcome for every row that does not already carry plan copy.
  • The offered set holds one row per model id — the first row the list returns for that id; another provider's row for the same id keeps its catalog tagline.

Copy is generated lazily on the first request that sees a new combination of facts (no timer), stored, and reused; the same facts never regenerate. tagline_source is emitted on GET /models only — the /model-prices admin feeds carry the raw catalog columns. A repeated ?provider= parameter is refused with 400 (provider must be a single string).

Each row also carries the model's native capability flags, resolved server-side so the client never hardcodes model lists:

  • supports_thinking — the model exposes extended-thinking/reasoning output. For a bare claude-* id this is resolved by the same server-side resolver the inference path uses (core.model_identity), so the capability a client is offered and the request the gateway actually makes agree by construction. That resolver reads two sources in precedence order: the gateway's curated capability registry (an override), and — when the registry is silent for that id — a provider-synced flag imported from the first-party Anthropic Models API (see Capability sources & the trust stamp below). A bare anthropic-cloud claude-* id the registry does not enumerate falls back to a name heuristic, which is broader: it can report true for a model whose thinking parameter the gateway will not send. Treat it as "offer the control", not as a guarantee.
  • Azure- and Bedrock-hosted Claude (azure_ai/claude-*; Bedrock ids such as anthropic.claude-…, eu.anthropic.claude-…, bedrock/<region>/…, and bare claude-…-vN:M) are resolved only from the registry, in their own per-platform key space — never via the anthropic-cloud name heuristic. What Azure-hosted and Bedrock-hosted Claude accept (thinking shape, real output ceiling, native-document types) is a platform fact, not inheritable from the anthropic-cloud entry, and each must be verified against the real platform before an entry is added. Until then these ids are fail-closed: supports_thinking (and thinking_always_on) report false — matching the backend, which likewise sends no thinking parameter — so a client is never offered an Adaptive-thinking toggle the platform will silently ignore.
  • thinking_always_on — the model reasons by default and cannot be turned off. Some always-adaptive Claude models (e.g. claude-fable-5) run extended thinking whenever the thinking parameter is omitted and reject an explicit disable, so supports_thinking alone (which only says "thinking is available") cannot express this. When true, a client should render the thinking control as locked on rather than a toggle that pretends "off" works. Fail-closed (absent → false, i.e. a normal on/off toggle). This is distinct from a model like claude-opus-5, which also reasons by default but can be disabled — for it the gateway emits an explicit thinking = {"type": "disabled"} when the client turns thinking off, and this flag stays false.
  • supports_vision — accepts native image input (image blocks); when false, attached images are extracted to text server-side instead.
  • supports_pdf_input — accepts a native PDF document block; when false, PDFs are extracted to text.
  • supports_files_api — supports the provider Files-API upload + document skill (Anthropic).

supports_pdf_input and supports_files_api are additionally constrained by the route the row would actually take: a model is only reported as accepting a native document block when the resolved provider's request path accepts one (Anthropic always; Bedrock only for inline base64 on an Anthropic-family model). The same Claude model reached through a re-selling provider therefore reports false and its documents are extracted to text — which is what that route supports — rather than advertising a native block the gateway would reject. - supports_function_calling — the model can use gateway-injected function tools (web search, URL fetch, connectors, sub-agents, knowledge search). Unlike the flags above this is permissive on the unknown space: it is true for every model except one curated as tool-incapable (e.g. Perplexity Sonar, which exposes no function-tools endpoint). An uncatalogued model reports true, so the client never over-blocks a tool-capable model whose flag it doesn't know; a false here is the authoritative signal the agent editor uses to disable all tool toggles. Sourced from the capability registry (the same one the runtime tool-strip reads), not the sparse model_price DB column.

Every flag above except supports_function_calling is a boolean that fails closed (absent → false); supports_function_calling is the sole fail-open flag (absent → true, tool-capable). supports_vision is resolved by a single shared function (capability.effective_vision) that the gateway ALSO uses when it decides whether to forward image pixels or extract them to text — so the value in this payload and the gateway's runtime behaviour can never disagree. It reports true when either:

  • the model is affirmatively vision-capable in the curated capability catalog/registry, or
  • the model id matches the curated vision name heuristic: a bare claude-* id (covers Claude models not individually enumerated in the registry, e.g. a freshly released claude-opus-4-6) or a *-vl-* id (Qwen-VL / ERNIE-VL / Nemotron-VL pass-through vision models). Matching is anchored and case-insensitive; azure_ai/claude-* is intentionally excluded (it reports false).

Otherwise supports_vision is false — an uncurated pass-through model (for example x-ai/grok-4.x on OpenRouter) whose vision capability the gateway cannot confirm. For those, attached images are extracted to text server-side by the gateway before dispatch; it never forwards raw pixels to a model it cannot confirm sees them, because a non-vision or picky upstream rejects an image block with an opaque 400. To force native image input for a specific uncurated pass-through vision model, register it in the capability catalog with supports_vision: true; it then receives pixels on both sides.

Native claude vision requires the bare model id (claude-opus-4-6). A provider-prefixed id (anthropic/claude-…, azure_ai/claude-…) is not matched by the name heuristic and its images are text-extracted unless the model is curated in the catalog.

*-vl-* caveat: if a curated *-vl-* pass-through model enforces a per-request image cap upstream (e.g. "at most 1 image"), that per-model image cap must be registered in the model's capability entry; otherwise the gateway forwards every image and the upstream may 400 on the excess.

Capability sources & the trust stamp

Behavioural capabilities resolve from two sources, in strict precedence: the in-image capability registry (an override) first; when it is silent for a model, a provider-synced flag imported by the daily model importer. The registry is no longer the only source — but it always wins where it speaks, and the synced layer only ever fills a registry silence, never downgrades a curated value.

Only first-party Anthropic models carry a provider-synced source: the importer reads GET /v1/models's capabilities tree (thinking.types.adaptive|enabled.supported, context_management.compact_20260112.supported) and writes three model_price columns — supports_reasoning, uses_adaptive_thinking, supports_native_compaction — stamped with a trust column, capability_source. Bedrock/Vertex-hosted Claude, Gemini, and OpenRouter expose no such tree, so they keep hand-curation and are never stamped.

capability_source — accepted values: the exact string provider, or SQL NULL. The resolver trusts a synced flag ONLY when capability_source = 'provider' (a strict allowlist). Any other value — NULL, empty string, litellm_json, catalog, or anything else — is treated as untrusted and the synced columns are ignored (the model falls back to the registry, else OFF). A synced supports_reasoning additionally takes effect only when the provider also declared a determinate thinking shape (uses_adaptive_thinking is a real 0/1, never NULL) and a provider-sourced output cap large enough to leave room for a visible answer; otherwise reasoning stays OFF. This is enforced DB-side: capability_source is only ever written to provider by the first-party Anthropic importer (community/litellm/catalog sync paths physically cannot write it), and uses_adaptive_thinking is tri-state so an undetermined shape can never trigger a hard 400. A synced capability flip (on↔off) raises an operator alert, and a routable Anthropic model with no resolvable capability record at all (neither registry nor a provider stamp) is flagged by a standing drift guard.

A native document block ({"type":"document","source":{...}}, used for native PDF / Files-API attachments) is an Anthropic-wire shape: only providers whose request serializer emits the Anthropic Messages wire carry it to a backend that accepts it — Anthropic, and AWS Bedrock for an anthropic-family model. Every OpenAI-compatible provider forwards it raw and the upstream rejects it with an opaque 400. The gateway therefore fails closed on the provider wire-shape (not on supports_pdf_input): a native document sent to any other provider/model is rejected before dispatch with model_capability_mismatch (HTTP 400) rather than forwarded — pick a document-capable model (e.g. Claude) or remove the attachment. The Anthropic Files-API (source.type: "file") is Anthropic-cloud-only and is not accepted via Bedrock. (Clients should extract non-PDF-capable attachments to text before sending; the web app does this automatically and folds a history document to a text reference when the selected model cannot read documents.)


GET /model-prices

Returns all stored model pricing records. Used internally for cost calculation on every inference request. Required role: admin or tenant_admin.

curl https://<your-gateway-host>/admin/v1/model-prices

Response:

A bare JSON array of pricing rows.

[
  {
    "provider": "openai",
    "model": "gpt-4o",
    "input_per_1k": 0.005,
    "output_per_1k": 0.015,
    "cache_write_per_1k": null,
    "cache_read_per_1k": null,
    "cache_write_1h_per_1k": null,
    "max_input_tokens": 128000,
    "max_output_tokens": 16384,
    "max_input_tokens_source": "litellm_json",
    "deprecated_at": null,
    "deprecated_source": null,
    "display_name": "GPT-4o",
    "tagline": null,
    "tagline_de": null,
    "metadata_source": "litellm",
    "hosted_in": "US",
    "updated_at": 1742551232
  },
  {
    "provider": "anthropic",
    "model": "claude-opus-4-6",
    "input_per_1k": 0.005,
    "output_per_1k": 0.025,
    "cache_write_per_1k": 0.00625,
    "cache_read_per_1k": 0.0005,
    "cache_write_1h_per_1k": 0.01,
    "max_input_tokens": 200000,
    "max_output_tokens": 64000,
    "max_input_tokens_source": "provider",
    "deprecated_at": null,
    "deprecated_source": null,
    "display_name": "Opus 4.6",
    "tagline": null,
    "tagline_de": null,
    "metadata_source": "manual",
    "hosted_in": "US",
    "updated_at": 1742551232
  }
]

ModelPrice fields

Field Type Description
provider string Provider identifier (e.g. openai, anthropic).
model string Exact model name as used in requests.
input_per_1k number Cost in USD per 1,000 input (prompt) tokens.
output_per_1k number Cost in USD per 1,000 output (completion) tokens.
cache_write_per_1k number | null Cost per 1,000 tokens written to provider prompt cache, 5-minute TTL (Anthropic). null if not applicable.
cache_read_per_1k number | null Cost per 1,000 tokens read from provider prompt cache (Anthropic). null if not applicable.
cache_write_1h_per_1k number | null Cost per 1,000 tokens written to provider prompt cache, 1-hour TTL (Anthropic). null if not applicable.
max_input_tokens integer | null Maximum input (context) window in tokens. null if unknown.
max_output_tokens integer | null Maximum output tokens the model can generate. null if unknown.
max_input_tokens_source string | null Provenance of the token-limit figures (max_input_tokens and max_output_tokens): provider (the provider's own model API, e.g. Anthropic GET /v1/models — authoritative), catalog (a self-hosted Myra pod's own configuration), or litellm_json (the community LiteLLM snapshot — a hint). A provider value is authoritative and is never overwritten by the LiteLLM hint; a routed alias (e.g. claude-opus-4-5) inherits the newest provider snapshot's caps. null for legacy rows whose source was not recorded.
deprecated_at integer | null Unix timestamp (seconds) when the model was deprecated, or null if active.
deprecated_source string | null Who deprecated the row: admin (an operator, through the deprecate endpoint), prober (the health prober's automatic flip after repeated not-found results), or myra_sync / openrouter_sync (a catalog reconciler, because the provider stopped offering the model). A reconciler clears only its own deprecation when the model reappears, so an operator's decision is never silently undone; an operator's deprecate call takes over a reconciler's, keeping the original timestamp. null for rows deprecated before provenance was recorded — those are likewise never cleared automatically.
display_name string | null Human-friendly model name shown in the picker. Derived server-side by a brand-stripped normaliser: the vendor prefix is dropped and the tier + version kept (claude-opus-4-8 → Opus 4.8, claude-opus-5 → Opus 5, claude-fable-5-1 → Fable 5.1). Any Claude tier is handled (opus/sonnet/haiku/fable/…), not a fixed list. A manual row (metadata_source = "manual") is never overwritten.
tagline string | null Short English description of the model, positioned by recency decided server-side: the newest model in a family reads as capable, and an older sibling is framed plainly as older (e.g. once a newer one is added). This self-heals — an older model's tagline is regenerated to an "older" framing when a newer sibling appears, and the newest never reads "older". manual rows are left untouched.
tagline_de string | null Short German description of the model (co-generated with tagline, same recency positioning).
tagline_source string GET /models only: plan_copy, catalog or name — see Tagline resolution above. On GET /models the tagline / tagline_de values are the resolved ones (for a self-serve workspace they can be plan copy instead of the catalog columns).
metadata_source string | null Where the row's metadata came from (e.g. litellm, openrouter, manual).
hosted_in string | null Region/jurisdiction the model is hosted in: "Myra" (self-hosted), "EU", "US", "China", or null when no rule applies (e.g. meta-routers). Derived server-side; drives the region indicator in the chat model picker.
display_rank integer | null Per-provider display order for the chat model picker, computed server-side on read: models are grouped into families and ordered by tier (a family's highest price) then newest version first, so 0 is shown first within its provider group. This makes a newer-but-cheaper flagship (e.g. Opus 5.5) sort above a pricier older sibling (Opus 4.8) — which a raw price sort gets wrong. Advisory display metadata only, never an entitlement or routing signal. May be null/absent on an older response; clients MUST tolerate its absence and fall back to their own ordering. Present on GET /models; not emitted on the /model-prices admin feeds.
updated_at integer | null Unix timestamp (seconds) of the last update to the row.

PUT /model-prices

Required role: platform admin (the MODEL_PRICES_MANAGE permission). A tenant_admin may read model prices but not change them.

Create or update a price for a model. The (provider, model) pair is the unique key. If a record already exists it is replaced.

⭐ Example: The following examples show how to upsert pricing for different provider and model combinations.

curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
  -H "Content-Type: application/json" \
  -d '{
    "provider": "openai",
    "model": "gpt-4o",
    "input_per_1k": 0.005,
    "output_per_1k": 0.015
  }'

With prompt cache pricing (Anthropic):

curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
  -H "Content-Type: application/json" \
  -d '{
    "provider": "anthropic",
    "model": "claude-opus-4-6",
    "input_per_1k": 0.005,
    "output_per_1k": 0.025,
    "cache_write_per_1k": 0.00625,
    "cache_read_per_1k": 0.0005
  }'

Accepted body

Field Required Accepted
provider, model yes Non-empty strings. model must be a well-formed model id (no empty final path segment).
input_per_1k, output_per_1k yes A finite number >= 0.
cache_write_per_1k, cache_read_per_1k, cache_write_1h_per_1k no A finite number >= 0, or null / omitted to clear.

Anything else is rejected with 400 and nothing is written: a missing or null required price, a non-numeric value (including a numeric string), NaN, Infinity, or a negative price. Negative is rejected rather than stored because spend is only recorded for a request that costs more than nothing — a negative price would switch metering off for the model entirely, the same way a zero price does, and budgets, caps, budget-alert e-mails and trial credit would stop accruing silently.

A price of exactly 0 is accepted: some self-hosted models are billed per a different unit and are priced at zero on purpose. A zero-priced chat model does raise a daily platform alert — see Provider Costs.

Custom internal model (e.g. a fine-tuned model on Azure):

curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
  -H "Content-Type: application/json" \
  -d '{
    "provider": "azure",
    "model": "my-ft-gpt4o-deployment",
    "input_per_1k": 0.008,
    "output_per_1k": 0.020
  }'

GET /model-prices/markup

Required role: platform admin (the MODEL_PRICES_MANAGE permission). Not readable by a tenant_admin — the markup is Myra's confidential charge-margin over raw provider cost and is deliberately absent from the tenant-readable GET /model-prices and the picker GET /models.

Lists the active external model rows (provider other than myra) with their billing region tag and markup percentage:

[
  { "provider": "anthropic", "model": "claude-sonnet-4-6", "display_name": "Sonnet 4.6",
    "markup_region": null, "markup_pct": null },
  { "provider": "mistral", "model": "mistral/mistral-small-latest", "display_name": "Mistral Small",
    "markup_region": "eu", "markup_pct": null }
]

markup_region is "eu", "non_eu", or null (unset → treated as non-EU +20 %). markup_pct is a number in [0,100] or null (unset → the region default: EU 10 %, non-EU 20 %).

PUT /model-prices/markup

Required role: platform admin (the MODEL_PRICES_MANAGE permission); a tenant_admin receives 403.

curl -X PUT https://<your-gateway-host>/admin/v1/model-prices/markup \
  -H 'Content-Type: application/json' \
  -d '{ "provider": "mistral", "model": "mistral/mistral-small-latest",
        "markup_region": "eu", "markup_pct": 10 }'

Accepted body

  • provider, model (required) — must identify an existing external model-price row. A model with no price row → 404 (a price-less external row would bill 0, so markup is never set on one).
  • markup_region — "eu" or "non_eu", or null/absent to clear (→ non-EU default). Any other string → 400, nothing written.
  • markup_pct — a finite number in [0,100], or null/absent to clear (→ region default). A non-number, NaN, infinity, negative, or > 100 → 400, nothing written. The value is never silently coerced.

The effective markup is applied to external-provider leg costs at cost-computation time, before the budget_usd/wallet debit — see Cost attribution. Self-hosted (myra) models are never marked up.


DELETE /model-prices/{provider}/{model}

Required role: platform admin (the MODEL_PRICES_MANAGE permission).

Remove a pricing record. Subsequent requests for this model fall back to the defaults of the hardcoded cost table of the gateway, or report zero cost if the model is not in the defaults.

curl -X DELETE "https://<your-gateway-host>/admin/v1/model-prices/openai/gpt-4o"

💡 Note: URL-encode the model name if it contains slashes. For example, accounts/fireworks/models/llama-v3-70b becomes accounts%2Ffireworks%2Fmodels%2Fllama-v3-70b.


POST /model-prices/{provider}/{model}/deprecate

Marks a model as deprecated, recording deprecated_source: "admin". This takes over a deprecation a catalog reconciler made (the original deprecated_at is kept), which is what makes an operator's decision durable: a reconciler only ever clears its own deprecation, so a model an operator deprecated stays deprecated even when the provider starts offering it again. A deprecated model is dropped from the chat-model catalogue served by GET /models (the picker feed) and from the Auto default pick, and is refused at serve time: an inference request that resolves to a deprecated (provider, model) — whether named explicitly by the client, pinned by a project, or reached through a routing rule or fallback chain — is rejected before any upstream dispatch with 400 model_not_found (header X-AIG-Error: model_not_found), rather than being routed to the provider. Its pricing row is kept so historical cost figures still resolve. This is the supported way to disable a model. Required role: platform admin only.

One exception: a deprecated Myra-fleet id that has a registered successor (a rollover such as qwen3.6-27b → qwen3.8-27b) is upgraded to the successor instead of refused when a client names it explicitly — see Inference › Retired model ids. Un-deprecating the old row switches that off (the request then goes to the old id as named).

curl -X POST "https://<your-gateway-host>/admin/v1/model-prices/anthropic/claude-opus-4-6/deprecate"

Response

{ "ok": true, "flipped": true }
Field Type Description
ok boolean Always true on success.
flipped boolean true if the model was active and is now deprecated; false if it was already deprecated (no change).

💡 Note: URL-encode the model name if it contains slashes (e.g. accounts/fireworks/models/llama-v3-70b becomes accounts%2Ffireworks%2Fmodels%2Fllama-v3-70b).


POST /model-prices/{provider}/{model}/un-deprecate

Clears the deprecation flag, returning the model to the picker feed and the Auto default pick. Required role: platform admin only.

curl -X POST "https://<your-gateway-host>/admin/v1/model-prices/anthropic/claude-opus-4-6/un-deprecate"

Response

{ "ok": true }

Returns 400 { "error": "provider '<provider>' is retired from the catalog" } for providers retired from the catalogue (none currently), and 500 { "error": ... } on a storage failure.


GET /model-prices/active-for-probing

Lists the active (provider, model) rows (those with deprecated_at IS NULL) for the scheduled health prober to ping. Optionally filtered by ?provider=<name>. Required role: platform admin only.

Response

[ { "provider": "anthropic", "model": "claude-opus-4-6" }, { "provider": "gemini", "model": "gemini-2.5-pro" } ]

POST /model-prices/{provider}/{model}/probe-result

Records one health-probe outcome for a model (used by the scheduled deprecation prober). Required role: platform admin only.

Request body

Field Type Description
outcome string Required. One of the accepted values below. Any other value is rejected with 400 { "error": "invalid outcome" } and nothing is recorded.
http_status integer | null The upstream HTTP status observed (optional).
upstream_error string | null Upstream error text, truncated to 255 chars (optional).
duration_ms integer | null Probe round-trip time in ms (optional; stored as 0 when absent or negative, capped at 2147483647; NaN counts as absent).
consecutive_threshold integer Consecutive-failure threshold for auto-deprecation (optional, default 3; accepted range 1..100 — larger values are clamped to 100, values below 1, non-numeric or NaN values fall back to the default). The probe log is kept for 30 days, so a threshold cannot exceed what the window can ever hold.

Accepted outcome values (any other value → 400):

Value Meaning Auto-deprecates?
pass 2xx — the model served the probe. No
fail_not_found 4xx model-not-found / deprecated shape. Yes — flips deprecated_at after consecutive_threshold consecutive fail_not_found outcomes.
fail_quota 429 / rate-limit / quota block. Record-only — a quota block is a transient route-around signal, not a deprecation, so it is logged for observability + human review but never auto-deprecates. No
fail_other Any other non-2xx (transient). No
driver_error The prober itself failed before reaching the upstream. No
fail_guided The hourly prober's guided-output leg (a tiny json_schema request against the in-house fleet) failed or hung — the vLLM guided-path / engine-death signature. Record-only — an engine crash is an ops incident, not a deprecation; also excluded from the flip window so interleaved rows can never mask a genuine fail_not_found streak. No

deprecated_at is a per-(provider, model) row, so deprecating one provider's copy of a model never affects another provider's copy (e.g. deprecating gemini/gemini-2.5-pro leaves openrouter/google/gemini-2.5-pro untouched).

Response

{ "ok": true, "flipped": false }
Field Type Description
ok boolean Always true on success.
flipped boolean true only when this fail_not_found outcome tipped the model into deprecation.

POST /model-prices/sync

Synchronises pricing data from the configured upstream sources (LiteLLM JSON and OpenRouter API) into the local model_price table. Required role: platform admin only.

# Sync every provider
curl -X POST https://<your-gateway-host>/admin/v1/model-prices/sync

# Sync one provider
curl -X POST 'https://<your-gateway-host>/admin/v1/model-prices/sync?provider=openrouter'

The response is a per-provider summary of inserted, updated, and skipped records.


GET /providers

Returns the list of provider integrations supported by the gateway with metadata such as requires_key. Used by the admin UI to populate the provider drop-down in the Add Model dialog.

curl https://<your-gateway-host>/admin/v1/providers

GET /providers/health

Returns one row per supported provider with these pieces of state merged together:

  • name — the provider identifier (e.g. openai, anthropic).
  • requires_key — whether the provider requires a BYOK API key.
  • configured — true when at least one BYOK key for the provider is stored on any gateway in the tenant; false when keys are required but absent; null when no key is needed.
  • status / message / checked_at / has_status_page — the live availability state, polled from the public Atlassian Statuspage of the provider every five minutes. Providers without a Statuspage have has_status_page = false.
  • latency_ms — the latency of the last status-page poll, if recorded.
curl https://<your-gateway-host>/admin/v1/providers/health

See also