Cost attribution
Myra AI Workspace computes and records the cost of every inference request, accumulates spend against configurable budget caps, and exposes aggregated cost data through the Stats API and admin UI.
💡 Display currency. Costs are stored and enforced in USD, but the UI shows them in the viewer's display currency — EUR by default, or USD if the user set that preference. Only the presentation converts (at the display boundary); budgets, caps, and stored prices stay in USD.
Cost calculation formula
cost_usd = ( (input_tokens / 1000) × input_per_1k
+ (output_tokens / 1000) × output_per_1k
+ (cache_write_5m_tokens / 1000) × cache_write_per_1k
+ (cache_write_1h_tokens / 1000) × cache_write_1h_per_1k
+ (cache_read_tokens / 1000) × cache_read_per_1k
+ (cache_deletion_tokens / 1000) × cache_delete_per_1k )
× tier_multiplier
For Anthropic requests that use prompt caching, the cache-write, cache-read and cache-deletion terms apply; for all other requests those token counts are zero and only the input and output terms contribute.
Token counts come from the provider response. For streaming requests, they are accumulated across all chunks and recorded when the stream ends.
When a single request makes several provider rounds in one turn — a server-side tool loop, or a length-cap auto-continuation — each round is a separate provider dispatch billed on its own input and output (the continuation re-sends the growing context, which the provider charges as that round's input). The request cost is the sum over every round; each round also appears as its own row in the per-leg ledger.
If a client disconnects mid-stream, the tokens the provider reported as generated before the disconnect are still billed — the provider charged us for them — recorded as a partial leg. You are billed for what the provider reported up to the disconnect, never for output that was never produced.
Tier multiplier
The summed cost is multiplied by a tier factor that reflects the provider's service tier:
| Tier | Multiplier |
|---|---|
standard (default) |
1.0 |
batch |
0.5 |
priority |
1.25 |
priority_on_demand |
1.5 |
An unknown or absent tier is treated as standard (1.0).
External-provider markup
For requests served by an external provider (OpenAI, Anthropic, Mistral API, and every
third-party or meta-router such as OpenRouter/Together/Fireworks), the per-leg cost is multiplied by
a markup after the tier multiplier and before it is debited from the tenant's budget_usd /
wallet:
- The default markup is +10 % for the EU-classified external model
mistral-small-latestand +20 % for every other external model (the non-EU default). A platform admin may override the percentage per model (0–100 %) via Settings → Costs → Providers (the admin-only Provider Pricing screen). - The markup applies only to real model-of-record legs. Sidecar meter legs (web search, code interpreter, guardrail/PII/document-extract) are never marked up.
- Myra self-hosted models are never marked up — they are billed at their directly-set internal price (no multiplier).
- A missing or malformed stored markup degrades to the non-EU default (+20 %) and is logged; it can never zero a cost or crash the cost path (fail-closed).
The markup is a charge-margin applied ONLY to budget/wallet consumption — it is debited from the
per-gateway, per-tenant, per-token and per-user (trial) spend ledgers that gate quota, and drains the
prepaid wallet. Everything else stays at raw provider cost (COGS): request_log.cost_usd, the
per-leg ledger, the cost-attribution analytics, and both the X-AIG-Cost-Micros and X-AIG-LLM-Cost
response headers report the raw cost. So a tenant's budget/wallet is drawn down by the marked-up
amount, while cost analytics reflect what the request actually cost Myra. The markup does not
affect Stripe invoices.
Anthropic prompt caching fields
The prompt caching feature of Anthropic writes frequently-used content (system prompts, document context) to a cache and bills cache reads at a steep discount. Myra AI Workspace tracks and reports these fields separately:
| Field | Typical price relative to input |
|---|---|
cache_write_5m_tokens (5-minute TTL) |
1.25× input price |
cache_write_1h_tokens (1-hour TTL) |
2× input price |
cache_read_tokens |
0.10× input price |
cache_deletion_tokens |
free (0×) |
These fields appear in request logs and are included in the usage chunk emitted for streaming compat endpoint responses (the chunk carries "choices": [] plus usage, the shape OpenAI's stream_options.include_usage chunk uses).
Pricing sources
The gateway resolves model prices using a two-level lookup:
Database prices (highest priority)
Prices stored in the admin UI or via the API take precedence over everything else. Use this to keep prices up-to-date as providers change their rates, or to add custom pricing for fine-tuned models.
# Upsert a model price
curl -X PUT https://<your-gateway-host>/admin/v1/model-prices \
-H 'Content-Type: application/json' \
-d '{
"provider": "openai",
"model": "gpt-4o",
"input_per_1k": 0.0025,
"output_per_1k": 0.010,
"cache_write_per_1k": null,
"cache_read_per_1k": null
}'
Built-in fallback prices
Used when no database entry matches the model. These cover the most common models across all supported providers and are maintained by Myra Security.
Built-in fallback prices
The following models have hardcoded fallback prices. All prices are USD per 1,000 tokens. The list below covers the most common providers; additional providers (AWS Bedrock, Perplexity, Azure OpenAI, Cohere, and others) also carry built-in fallback prices, used when a model is not in the Provider Costs table.
OpenAI
| Model | Input/1K | Output/1K |
|---|---|---|
| gpt-4o | $0.0025 | $0.010 |
| gpt-4o-mini | $0.00015 | $0.0006 |
| gpt-4-turbo | $0.010 | $0.030 |
| gpt-3.5-turbo | $0.0005 | $0.0015 |
Anthropic
| Model | Input/1K | Output/1K | Cache write 5m/1K | Cache write 1h/1K | Cache read/1K |
|---|---|---|---|---|---|
| claude-opus-4-6 | $0.005 | $0.025 | $0.00625 | $0.01 | $0.0005 |
| claude-sonnet-4-6 | $0.003 | $0.015 | $0.00375 | $0.006 | $0.0003 |
| claude-haiku-4-5-20251001 | $0.001 | $0.005 | $0.00125 | $0.002 | $0.0001 |
Gemini / Vertex AI
| Model | Input/1K | Output/1K |
|---|---|---|
| gemini-1.5-pro | $0.00125 | $0.005 |
| gemini-1.5-flash | $0.000075 | $0.0003 |
| gemini-2.0-flash | $0.0001 | $0.0004 |
Mistral
| Model | Input/1K | Output/1K |
|---|---|---|
| mistral-large-latest | $0.002 | $0.006 |
| mistral-small-latest | $0.0002 | $0.0006 |
| codestral-latest | $0.0003 | $0.0009 |
Groq
| Model | Input/1K | Output/1K |
|---|---|---|
| llama-3.3-70b-versatile | $0.00059 | $0.00079 |
| llama-3.1-8b-instant | $0.00005 | $0.00008 |
xAI
| Model | Input/1K | Output/1K |
|---|---|---|
| grok-3 | $0.003 | $0.015 |
| grok-3-mini | $0.0003 | $0.0005 |
DeepSeek
| Model | Input/1K | Output/1K |
|---|---|---|
| deepseek-chat | $0.00027 | $0.0011 |
| deepseek-reasoner | $0.00055 | $0.00219 |
Model pricing data is maintained by Myra Security and updated automatically. View and override individual model prices via the Models and pricing API.
Budget enforcement
Cost attribution feeds directly into the budget enforcement system:
- Per-gateway budget — set
budget_usdin the gateway config; returns429 quota_exceededonce the accumulated spend reaches the cap. - Per-token budget — set
budget_usdon an auth token; enforced independently of the gateway budget. Both caps apply: a request is refused when either the token or the gateway budget is exhausted, so a token still under its own cap is blocked once the gateway budget is reached. - Reset —
DELETE /admin/v1/gateways/{id}/budgetresets the gateway counter (platform-admin only and audited — a tenant admin receives403);DELETE /admin/v1/users/{id}/budgetresets all token budgets for a user (token-scoped, unaffected by that gate).
See also
- Request pipeline — where cost calculation fits in the request lifecycle
- Cost analytics — how to view spend breakdowns
- Budgets and quotas — configuring and resetting spend caps
- Response caching —
saved_cost_usdlogged on cache hits - Models API — managing model prices at runtime