Skip to content

Cost attribution

Myra AI Workspace computes and records the cost of every inference request, accumulates spend against configurable budget caps, and exposes aggregated cost data through the Stats API and admin UI.

💡 Display currency. Costs are stored and enforced in USD, but the UI shows them in the viewer's display currency — EUR by default, or USD if the user set that preference. Only the presentation converts (at the display boundary); budgets, caps, and stored prices stay in USD.


Cost calculation formula

cost_usd = ( (input_tokens             / 1000) × input_per_1k
           + (output_tokens            / 1000) × output_per_1k
           + (cache_write_5m_tokens    / 1000) × cache_write_per_1k
           + (cache_write_1h_tokens    / 1000) × cache_write_1h_per_1k
           + (cache_read_tokens        / 1000) × cache_read_per_1k
           + (cache_deletion_tokens    / 1000) × cache_delete_per_1k )
         × tier_multiplier

For Anthropic requests that use prompt caching, the cache-write, cache-read and cache-deletion terms apply; for all other requests those token counts are zero and only the input and output terms contribute.

Token counts come from the provider response. For streaming requests, they are accumulated across all chunks and recorded when the stream ends.

When a single request makes several provider rounds in one turn — a server-side tool loop, or a length-cap auto-continuation — each round is a separate provider dispatch billed on its own input and output (the continuation re-sends the growing context, which the provider charges as that round's input). The request cost is the sum over every round; each round also appears as its own row in the per-leg ledger.

If a client disconnects mid-stream, the tokens the provider reported as generated before the disconnect are still billed — the provider charged us for them — recorded as a partial leg. You are billed for what the provider reported up to the disconnect, never for output that was never produced.

Tier multiplier

The summed cost is multiplied by a tier factor that reflects the provider's service tier:

Tier Multiplier
standard (default) 1.0
batch 0.5
priority 1.25
priority_on_demand 1.5

An unknown or absent tier is treated as standard (1.0).

External-provider markup

For requests served by an external provider (OpenAI, Anthropic, Mistral API, and every third-party or meta-router such as OpenRouter/Together/Fireworks), the per-leg cost is multiplied by a markup after the tier multiplier and before it is debited from the tenant's budget_usd / wallet:

charged_cost = raw_cost × (1 + markup_pct / 100)
  • The default markup is +10 % for the EU-classified external model mistral-small-latest and +20 % for every other external model (the non-EU default). A platform admin may override the percentage per model (0–100 %) via Settings → Costs → Providers (the admin-only Provider Pricing screen).
  • The markup applies only to real model-of-record legs. Sidecar meter legs (web search, code interpreter, guardrail/PII/document-extract) are never marked up.
  • Myra self-hosted models are never marked up — they are billed at their directly-set internal price (no multiplier).
  • A missing or malformed stored markup degrades to the non-EU default (+20 %) and is logged; it can never zero a cost or crash the cost path (fail-closed).

The markup is a charge-margin applied ONLY to budget/wallet consumption — it is debited from the per-gateway, per-tenant, per-token and per-user (trial) spend ledgers that gate quota, and drains the prepaid wallet. Everything else stays at raw provider cost (COGS): request_log.cost_usd, the per-leg ledger, the cost-attribution analytics, and both the X-AIG-Cost-Micros and X-AIG-LLM-Cost response headers report the raw cost. So a tenant's budget/wallet is drawn down by the marked-up amount, while cost analytics reflect what the request actually cost Myra. The markup does not affect Stripe invoices.


Anthropic prompt caching fields

The prompt caching feature of Anthropic writes frequently-used content (system prompts, document context) to a cache and bills cache reads at a steep discount. Myra AI Workspace tracks and reports these fields separately:

Field Typical price relative to input
cache_write_5m_tokens (5-minute TTL) 1.25× input price
cache_write_1h_tokens (1-hour TTL) 2× input price
cache_read_tokens 0.10× input price
cache_deletion_tokens free (0×)

These fields appear in request logs and are included in the usage chunk emitted for streaming compat endpoint responses (the chunk carries "choices": [] plus usage, the shape OpenAI's stream_options.include_usage chunk uses).


Pricing sources

The gateway resolves model prices using a two-level lookup:

Database prices (highest priority)

Prices stored in the admin UI or via the API take precedence over everything else. Use this to keep prices up-to-date as providers change their rates, or to add custom pricing for fine-tuned models.

# Upsert a model price
curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
  -H 'Content-Type: application/json' \
  -d '{
    "provider": "openai",
    "model": "gpt-4o",
    "input_per_1k": 0.0025,
    "output_per_1k": 0.010,
    "cache_write_per_1k": null,
    "cache_read_per_1k": null
  }'

Built-in fallback prices

Used when no database entry matches the model. These cover the most common models across all supported providers and are maintained by Myra Security.


Built-in fallback prices

The following models have hardcoded fallback prices. All prices are USD per 1,000 tokens. The list below covers the most common providers; additional providers (AWS Bedrock, Perplexity, Azure OpenAI, Cohere, and others) also carry built-in fallback prices, used when a model is not in the Provider Costs table.

OpenAI

Model Input/1K Output/1K
gpt-4o $0.0025 $0.010
gpt-4o-mini $0.00015 $0.0006
gpt-4-turbo $0.010 $0.030
gpt-3.5-turbo $0.0005 $0.0015

Anthropic

Model Input/1K Output/1K Cache write 5m/1K Cache write 1h/1K Cache read/1K
claude-opus-4-6 $0.005 $0.025 $0.00625 $0.01 $0.0005
claude-sonnet-4-6 $0.003 $0.015 $0.00375 $0.006 $0.0003
claude-haiku-4-5-20251001 $0.001 $0.005 $0.00125 $0.002 $0.0001

Gemini / Vertex AI

Model Input/1K Output/1K
gemini-1.5-pro $0.00125 $0.005
gemini-1.5-flash $0.000075 $0.0003
gemini-2.0-flash $0.0001 $0.0004

Mistral

Model Input/1K Output/1K
mistral-large-latest $0.002 $0.006
mistral-small-latest $0.0002 $0.0006
codestral-latest $0.0003 $0.0009

Groq

Model Input/1K Output/1K
llama-3.3-70b-versatile $0.00059 $0.00079
llama-3.1-8b-instant $0.00005 $0.00008

xAI

Model Input/1K Output/1K
grok-3 $0.003 $0.015
grok-3-mini $0.0003 $0.0005

DeepSeek

Model Input/1K Output/1K
deepseek-chat $0.00027 $0.0011
deepseek-reasoner $0.00055 $0.00219

Model pricing data is maintained by Myra Security and updated automatically. View and override individual model prices via the Models and pricing API.


Budget enforcement

Cost attribution feeds directly into the budget enforcement system:

  • Per-gateway budget — set budget_usd in the gateway config; returns 429 quota_exceeded once the accumulated spend reaches the cap.
  • Per-token budget — set budget_usd on an auth token; enforced independently of the gateway budget. Both caps apply: a request is refused when either the token or the gateway budget is exhausted, so a token still under its own cap is blocked once the gateway budget is reached.
  • Reset — DELETE /admin/v1/gateways/{id}/budget resets the gateway counter (platform-admin only and audited — a tenant admin receives 403); DELETE /admin/v1/users/{id}/budget resets all token budgets for a user (token-scoped, unaffected by that gate).

See also