Skip to content

Budgets

A budget is a spend ceiling that the gateway enforces before forwarding a request to an upstream provider. When the accumulated spend within the active period reaches the budget, the gateway rejects further requests with a quota_exceeded response until the period resets.

Three-tier model

Budgets are enforced at three independent levels:

  • Token-level budget — bound to a single authentication token. Useful for capping the spend of a specific application or integration.
  • Tenant-level budget — bound to a tenant. Aggregates the spend across every gateway and every token of the tenant.
  • Gateway-level budget — bound to a gateway. Aggregates the spend across every token of the gateway.

The most restrictive applicable budget wins. A request is allowed only when every applicable level has remaining budget.

Reset periods

The configured budget_period determines when the spend counter resets:

  • daily — the period key is the calendar date (YYYY-MM-DD); the counter resets at the start of each day.
  • monthly (default) — the period key is the calendar month (YYYY-MM); the counter resets on the first day of the month.
  • total — the period key is the literal string total; the counter never resets automatically.

Period boundaries follow the server's local date, not a fixed time zone. They coincide with UTC only when the gateway container runs in the UTC time zone — the standard deployment configuration.

Spend ledger

The gateway persists every request's cost in a spend ledger. The ledger is the source of truth for budget enforcement and for the Cost analytics view.

Concurrent requests

A request's exact cost is only known after the model has answered, so the gateway cannot compare a request's real cost to the budget before admitting it. To keep a token's spend cap meaningful when many requests arrive at once on the same token, the gateway places a small reservation against the token budget the moment a request is admitted, and reconciles that reservation to the request's true cost once the answer is complete.

The effect on the token budget is:

  • Requests are admitted only while the committed spend plus the reservations already held by in-flight requests are below the cap. Concurrent requests on the same token can therefore no longer all slip through against the same pre-spend balance — once the cap is committed or reserved, further concurrent requests are rejected with quota_exceeded, exactly as a later sequential request would be.
  • The reservation is settled as soon as the request finishes: for a billed turn the request's real cost is committed to the ledger and the reservation released in its place; for a turn that spends nothing (served from cache, blocked downstream, or aborted) the reservation is simply released. Either way an abandoned request does not permanently tie up budget.
  • Because the reservation is a fixed estimate rather than the request's true cost, a token's cap can still be exceeded by the combined real cost of the requests that were already in flight when the cap was reached. The reservation bounds how many requests are admitted at once — roughly cap ÷ reservation — and each of those admitted requests is then allowed to complete at its real cost. With the default reservation a small cap can still admit a couple of concurrent turns (for example, two on a $0.02 cap), each of which may cost more than the reservation. This is a large improvement over the unbounded pre-reservation behaviour (where every concurrent request slipped through), but it is not a hard single-request ceiling; size a token's cap with the token's expected concurrency in mind.

Enforcement remains exact for billing: the ledger always records each request's true cost. The reservation only governs admission. If the spend check cannot be completed because the database is briefly unavailable, the token gate fails open for that request (it is admitted and still metered at the end) rather than blocking traffic on a transient fault — the tenant- and gateway-level budgets are checked independently and still apply.

The reservation amount is a fixed per-request estimate (a small default) and is not normally tuned; your operator can override it as a deployment setting (an absent, zero, negative, or non-numeric value falls back to the safe default).

Consumption visibility and budget alerts

The same ledger drives every consumption surface, so what a user sees can never drift from what enforcement decides:

  • Per-token consumption — the Profile page's Tokens table shows each token's current-period spend and remaining amount, not only its cap.
  • In-app budget alerts (dashboard banner) — administrators and tenant-administrators see a persistent banner on the dashboard for budgets that are nearing or over their limit, across the three budget scopes: the tenant pool, each gateway, and each auth token — plus, for a self-serve tenant, its plan allowance (folded into the same alert set). Each alert is tiered: approaching (at or above 80 % of the cap), reached (at 100 %), and exceeded (over 100 %). Reached and exceeded mean the entity is already blocked at 429 — the "blocked" tier reuses the same spent ≥ cap comparison as enforcement, so the banner cannot show a softer state than reality. A personal token-scope alert links to the Profile page (Account › Profile) so the owner can see the token's spend. A service / scheduler token-scope alert (no personal owner) instead names the token and states that its scheduled runs and requests-made-with-it are blocked (not the viewing admin's own requests), and links to that token's gateway Auth Tokens card, where its spend cap can be raised.
  • Only administrators and tenant-administrators receive alerts; members never do (enforced server-side). A self-serve tenant's plan allowance replaces the tenant-pool budget scope in the alert set — its own over-cap state (degraded or blocked) maps to the same tiers, so a self-serve admin gets one coherent set (gateway / token budgets plus the allowance) rather than a separate surface. If a spend figure cannot be read momentarily, that entity is simply omitted — never shown as 0 % or as blocked on a stale value.
  • Notifications center — the same alerts also appear in a page-independent notifications center: a bell in the sidebar shows the count of active alerts (marked urgent when any is blocked) and opens a panel listing each one. It reads the same server-computed alert set as the dashboard banner, so the two can never disagree; unlike the dismissible banner, the center always reflects the current active set.
  • Token-scope alerts are owner-scoped. A personal per-token budget only blocks its own token's requests — never anyone else's — so its alert is shown only to the token's owner, never tenant-wide. This prevents a personal token that has run out of budget from raising a global "New requests are blocked" banner for unrelated members of the same tenant whose own requests are unaffected. A service / scheduler token that no person owns (no user_id) still surfaces to administrators and tenant-administrators, since it is the shared class that can otherwise hard-stop scheduled agents silently. Tenant- and gateway-scope alerts remain visible to all administrators, because those caps do block every user routing through them.
  • An approaching banner is dismissible; the dismissal is stored server-side and the banner re-appears automatically when the alert set changes (a budget crosses into a higher tier, a different entity breaches, or a new budget period begins). A blocked (reached / exceeded) banner cannot be dismissed, so an actively hard-stopping budget stays visible until it is resolved. A service / scheduler token block renders in the softer warning tone rather than the red error tone — it is not the viewing admin's own traffic — but it is still always-shown and non-dismissible while active, because a scheduler token silently hitting its 429 is exactly the failure this banner exists to surface; only the red own-block (tenant / gateway / personal token) uses the error tone.
  • Rate-limit alerts (rate_limit scope). A gateway or auth token that is being rate-limited (429s from its per-gateway / per-token request rate limit, distinct from the spend budget) surfaces as a rate_limit-scope alert carrying the count of throttle events that day. These are counted in a per-worker in-memory tally and periodically flushed to a persisted cross-worker aggregate (per local calendar day), which the alert reads — so the count is reliable across workers and process restarts, unlike the ephemeral sliding-window rate-limit counters themselves. Rate-limit alerts appear in the notifications center only (not the dismissible dashboard banner) and are not part of the dismiss signature, so surfacing or dropping a transient throttle never re-arms a dismissed budget warning. Token rate-limit alerts follow the same owner-scoping as token budget alerts (a personal token only to its owner; a service token to administrators). The alert clears when that day's counter period rolls over.

Cost calculation

The cost of a request is computed as (input_tokens ÷ 1000) × input_price + (output_tokens ÷ 1000) × output_price, using the price stored in the Provider Costs view (Settings › Costs) for the provider and model — prices are quoted per 1,000 tokens. Cache-read, cache-write, and priority-tier terms are added on top; see Cost attribution for the full formula. Requests against models without a known price are routed but not counted against budgets and do not appear in spend totals.

See also