Stats API
The Stats API exposes aggregated request metrics. Use GET /stats for the dashboard summary view (today, yesterday, last 7 days, last hour, last minute) and GET /stats/timeseries for time-bucketed data suitable for charts.
Base URL: https://<your-gateway-host>/admin/v1
GET /stats
Returns a snapshot of request counts, token usage, cost, and latency across several time windows. Also includes a per-tenant breakdown and the most recent log entries.
Every caller below admin/tenant_admin (member, viewer, ki_manager, demouser) is scoped to their own data. Non-admin callers are pinned to their own tenant server-side (a spoofed tenant_id is ignored); a caller with no tenant is refused 403 { "error": "forbidden" } (the query never runs). A platform admin may target one tenant with ?tenant_id= or omit it (and an empty value) for the global, all-tenant view.
The calendar-day windows (today, yesterday, last_7d) are computed in the viewer's timezone, supplied by the tz_offset query parameter — so "today" means the viewer's local calendar day, not a UTC day. The rolling windows (hour, last_min) are relative to the current instant and are timezone-independent.
Query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
tenant_id |
string (UUID) | caller's tenant | Platform admin only: target one tenant, or omit / empty for the global view. Ignored (server-pinned) for every other role. |
tz_offset |
integer | 0 (UTC) |
The viewer's UTC offset in minutes, in the JavaScript Date.getTimezoneOffset() convention: UTC − local, so Central European Summer Time (UTC+2) is -120 and US Eastern Standard Time (UTC−5) is 300. Used only to anchor the today / yesterday / last_7d calendar-day boundaries to the viewer's local midnight. Accepted: an integer within [-840, 720] (UTC−14 … UTC+14). Rejected → falls back to UTC (offset 0): any non-numeric, empty, repeated, out-of-range, or missing value (fractional values are floored toward −∞). It never causes an error and never influences authorization — it shifts a display window only. |
The GET /stats/timeseries series that back the dashboard sparkline glyphs accept the same tz_offset parameter and bucket in the viewer's local frame — so the daily (1d) bucket and, for a fractional-hour offset (e.g. India, UTC+5:30), the hourly buckets align with the viewer's local day/hour rather than UTC. For whole-hour offsets (e.g. Central Europe, US) the hourly buckets coincide with local hours either way. Omit tz_offset (or send a rejected value) and buckets fall back to UTC.
Response structure
{
"today": { ...PeriodStats },
"yesterday": { ...PeriodStats },
"last_7d": { ...PeriodStats },
"hour": { ...PeriodStats },
"last_min": { ...PeriodStats },
"by_tenant": [ ...TenantStats ],
"recent": [ ...RecentEntry ],
"recent_blocked": [ ...RecentBlockedEntry ]
}
recent and recent_blocked resemble LogEntry but are not the same shape. recent (up to 10 newest entries) omits id and instead adds the tenant and gateway slugs alongside the IDs. recent_blocked (up to 20 newest blocked/scrubbed/flagged entries) is a narrow subset carrying only the block/guardrail fields plus prompt_scrubbed and response_raw.
PeriodStats fields
| Field | Type | Description |
|---|---|---|
requests |
integer | Total inference requests in the period. |
cached |
integer | Requests served from cache (no provider call made). |
blocked |
integer | Requests blocked by auth, rate limit, quota, or detectors. |
rate_limited |
integer | Requests that hit a rate limit in the period. |
scrubbed |
integer | Requests where a guardrail scrubbed PII (personally identifiable information) from the payload but allowed the request through. |
flagged |
integer | Requests where a guardrail raised a flag but took no blocking or scrubbing action. |
input_tokens |
integer | Total prompt tokens consumed. |
output_tokens |
integer | Total completion tokens generated. |
cost_usd |
number | Total cost in USD (from model pricing table). |
saved_cost_usd |
number | Cost saved by cache hits (would-have-been cost of cached requests). |
avg_latency_ms |
number | Average end-to-end request latency in milliseconds. |
avg_upstream_latency_ms |
number | Average time waiting for the upstream provider, excluding gateway overhead. |
TenantStats fields (GET /stats)
by_tenant in GET /stats is a summary view with the following fields:
| Field | Type | Description |
|---|---|---|
tenant_id |
string | Tenant UUID. |
tenant |
string | Tenant slug. |
requests |
integer | Total requests. |
input_tokens |
integer | Total input tokens. |
output_tokens |
integer | Total output tokens. |
cost_usd |
number | Total cost in USD. |
budget_remaining |
number | Remaining budget in USD for the tenant's current period (the live-monitor "Quota Left" column). Absent when the tenant has no budget. |
budget_remaining_unknown |
boolean | true when the remaining budget could not be computed (a transient spend-read failure) — distinct from "no limit". |
The full TenantStats shape (with blocked, cached, errors, avg_latency_ms, etc.) is returned by GET /stats/analytics. See below.
Example response
{
"today": {
"requests": 1423,
"cached": 82,
"blocked": 14,
"input_tokens": 1840200,
"output_tokens": 312400,
"cost_usd": 9.84,
"saved_cost_usd": 0.41,
"avg_latency_ms": 820,
"avg_upstream_latency_ms": 760
},
"yesterday": {
"requests": 2101,
"cached": 134,
"blocked": 22,
"input_tokens": 2750000,
"output_tokens": 480000,
"cost_usd": 14.30,
"saved_cost_usd": 0.67,
"avg_latency_ms": 905,
"avg_upstream_latency_ms": 850
},
"last_7d": { "requests": 11200, "cost_usd": 68.12, "..." : "..." },
"hour": { "requests": 61, "cost_usd": 0.38, "...": "..." },
"last_min": { "requests": 2, "cost_usd": 0.01, "...": "..." },
"by_tenant": [
{
"tenant_id": "ten_abc123",
"tenant": "myapp",
"requests": 1423,
"input_tokens": 1840200,
"output_tokens": 312400,
"cost_usd": 9.84
}
],
"recent": [ { "...": "LogEntry" } ],
"recent_blocked": [ { "...": "LogEntry" } ]
}
For the related LogEntry field definitions see the Logs API. Note that the recent and recent_blocked rows are not full LogEntry objects (see the field notes above).
GET /stats/analytics
Returns latency percentiles, top-model breakdown, and per-tenant, per-gateway, per-channel, per-agent, and per-user cost summaries, plus activity counters. Used by the analytics dashboard view.
On the global, all-tenant view (a platform admin who omits tenant_id), every grouping and counter excludes synthetic fixture tenants — e.g. the e2e-synthetic monitoring prober — so the platform analytics reflect real usage. A tenant-scoped view (?tenant_id=…) still returns them.
Query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
since |
integer | 24 hours ago | Start of the analysis window as Unix milliseconds. |
until |
integer | now | Optional upper bound as Unix milliseconds. Used for closed windows such as "yesterday". |
tenant_id |
string | — | Filter by tenant. Honoured only for callers with the admin role. Non-admin callers are always scoped to their own tenant. |
anonymize |
boolean | true (anonymized) |
Anonymized by default. Unless the caller explicitly opts out with anonymize=0 (or false), the response omits the per-user grouping (by_user) entirely — no individual identifiers (email/user_id) are returned. Any other value (absent, 1, true, or malformed) keeps the response anonymized (fail-closed). Aggregate groupings (channel, agent, model, gateway, tenant) and counters are always included. The flag is scoped to this endpoint. Note: in a single-user tenant, per-agent/per-model aggregates can still imply one person's usage. |
group_id |
string | — | Restrict the whole response to the members of one user group. For a KI-Manager this must be a group they manage (see below); admin/tenant_admin may narrow to any group in their tenant scope. An unknown or out-of-scope group_id (including a group_id paired with a mismatched tenant_id) returns 404 not found; a KI-Manager requesting a same-tenant group they do not manage returns 403 forbidden. |
tz_offset |
integer | 0 (UTC) |
The viewer's UTC offset in minutes, in the JavaScript Date.getTimezoneOffset() convention: UTC − local (Central European Summer Time = -120, US Eastern Standard Time = 300). Applies only to the per-agent daily series (by_agent_series), whose buckets are floored in the viewer's local frame so each daily bar starts at the viewer's local midnight — aligning it with the calendar-day window (since/until) the caller computes with the same offset. Accepted: an integer within [-840, 720] (UTC−14 … UTC+14). Rejected → falls back to UTC (offset 0): any non-numeric, empty, repeated, out-of-range, or missing value (fractional values are floored toward −∞). It never causes an error and never influences authorization — it shifts the bucket boundaries only. |
💡 Note:
/stats/analyticsrequires theadmin,tenant_admin, orki_managerrole; lower-privileged callers (member,viewer,demouser) receive403 forbidden. Theanonymize=0opt-out is therefore confined to those privileged roles, and only ever reveals users within the caller's authorized scope:admin/tenant_adminsee their own tenant (atenant_adminis pinned to it), while aki_manageris pinned to the members of the groups they manage. Any role that could sendanonymize=0without authorization is already blocked with403before the handler runs; the client value never decides access.
KI-Manager group scoping
A KI-Manager may optionally see the statistics of the groups they manage. Unlike admin/tenant_admin (who see the whole tenant/workspace), a KI-Manager is always scoped to a group's members and can never see tenant-wide or another group's data:
- "Groups they manage" are the groups the KI-Manager created (
user_group.created_by). This is the same boundary as the group-management API (can_manage_group). Limitation: the data model has no separate "group manager" or delegation concept, so a group created by an admin and handed to a KI-Manager to run — or a group whose creator account was later deleted (created_bybecomesNULL) — is not visible to the KI-Manager. Onlyadmin/tenant_adminsee those. - Without
group_id: the response covers the union of the members of all groups the KI-Manager manages. A KI-Manager who manages no (non-empty) groups gets an all-zero response — never a tenant-wide fallback. - With
group_id: the response covers only that group's members, and only if the KI-Manager manages it (else403/404as above). - "A group's statistics" = activity by the group's members. All groupings and counters are filtered to
request_logrows / rows owned by the member set. Becauserequest_logcarries no group dimension, this is the only coherent definition. Consequences to keep in mind: by_agentis filtered on the run-as owner (user_id) — which, because invoke is owner-only, is the real member who invoked. Rows roll up over the clone lineage (see AgentStats), so a KI-Manager sees, for each shared agent, the invocations/conversations/distinct-users contributed by their members' clones of it (the display name is the root agent's, which may be owned outside the group).active_usersis windowed: under a group filter it is the count of the group's members active in the window (distinct members with ≥1 request in the period), not the group's current member count.
Response structure
{
"percentiles": { ...LatencyPercentiles },
"top_models": [ ...TopModelRow ],
"by_tenant": [ ...TenantStats ],
"by_gateway": [ ...GatewayStats ],
"by_channel": [ ...ChannelStats ],
"by_agent": [ ...AgentStats ],
"by_agent_series": [ ...AgentSeries ],
"by_user": [ ...UserStats ],
"counters": { ...ActivityCounters },
"anonymized": true,
"cache_efficiency": { ...CacheEfficiency }
}
ChannelStats fields (UI vs API vs agent)
Each row is the usage attributed to one access channel. The channel is derived: agent for agent invocations, ui for chat-session (playground) tokens, api otherwise.
| Field | Type | Description |
|---|---|---|
channel |
string | ui, api, or agent. |
requests |
integer | Request count for the channel. |
blocked |
integer | Requests blocked or scrubbed by a guardrail. |
cached |
integer | Requests served from the response cache. |
input_tokens |
integer | Total prompt tokens consumed on the channel. |
output_tokens |
integer | Total completion tokens produced on the channel. |
cost_usd |
number | Total cost. |
saved_cost_usd |
number | Cost avoided by cache hits. |
errors |
integer | HTTP ≥ 400 (excluding client-cancel 499). |
client_aborted |
integer | Requests the client cancelled (HTTP 499). |
avg_latency_ms |
integer | Average latency for non-blocked requests. |
AgentStats fields (per agent, rolled up over the clone lineage)
Each row is one lineage root — a shared/original agent plus every clone made from
it. Rows are keyed by agent_id = the root's id; a clone's usage rolls into its root
(the counting model below). Filtering (tenant, KI-Manager group scope) applies before the
roll-up.
| Field | Type | Description |
|---|---|---|
agent_id |
string | The lineage root agent's id (a shared agent + its clones share one row). |
agent |
string | Root agent name (or slug). A hard-deleted root falls back to the raw id. |
invocations |
integer | Invocations across the lineage in the window. |
conversations |
integer | Distinct chat sessions the lineage was used in. Headless invokes (API/schedule) carry no session id and count as invocations, not conversations — so this is 0 for agents only ever invoked headlessly, and becomes non-zero when an agent is used inside a chat thread. Read from request_log_legs.conversation_id; no per-invoke session id is fabricated. |
distinct_users |
integer | Distinct people who used the lineage — see the counting model below. |
cost_usd |
number | Total cost across the lineage. |
Per-agent "distinct users" — the counting model (supersedes the earlier model). Invoke is owner-only (
agent_invokerefuses a non-owner caller), and the catalog "use" action is clone-then-own (a member who wants a shared agent clones it into a new agent they own, then invokes that). So a single agent's distinct-invoker count is structurally ≤1 — the "Scheinwert". The fix is not to drop the metric but to count across the clone lineage: each member who uses a shared agent owns a distinct clone and invokes it as owner (invoker_id= that member), soCOUNT(DISTINCT invoker_id)over the root and its clones = the true number of people using it. A clone records its lineage root inagent.source_agent_id(set by the clone flow, flattened so the tree is one level); every request stampsrequest_log.agent_root_id, so the count survives the clone being deleted. Caveats (bid-relevant): lineage is recorded from the release that introduced it — pre-existing clones are not retro-linked; and because there is no server clone endpoint yet, the lineage pointer is client-declared within the caller's visibility (a member can omit it → undercount, or attribute a clone to any agent they can already see) — sodistinct_usersis a best-effort, self-reported figure, not server-authoritative. Unattended runs (scheduler/service token) log aNULLinvoker and are excluded from the count.
AgentSeries fields (per-agent calls over time)
One entry per top agent (up to 8, by invocation count) carrying a daily-bucketed
invocation time series. The series is dense and zero-filled across the window, so every
day is present even with no calls. Timestamps are milliseconds since epoch (bucket
start); when the tz_offset query parameter is supplied the bucket start is the viewer's
local midnight rather than UTC midnight (so the bars align with the viewer-local
calendar-day windows). Empty ([]) when there are no agent invocations in the window.
| Field | Type | Description |
|---|---|---|
agent_id |
string | The saved agent's id. |
agent |
string | Agent name (or slug). |
series |
array | Daily points: { "ts": <ms>, "invocations": <integer> }. |
ActivityCounters fields
Counts for the window, scoped to the tenant filter. *_created count items created in the window; active_users counts distinct users active in the window — users with at least one logged request (UI or API channel) between since and until. System/unattributed requests (no user_id) are excluded; duplicate requests by one user count once. It is not the provisioned-user headcount, so it can differ from the number of user rows (e.g. a user active in the window then deleted still counts).
Two caveats follow from active_users being derived from request_log:
- No explicit window (as in the tenant export's usage-statistics snapshot) → active_users reflects the trailing 24h, consistent with the sibling *_created counters, which use the same default. The provisioned headcount remains available separately (e.g. the export's users array).
- Request logging disabled for a tenant (request_logging_disabled) → no request_log rows are written, so active_users reads 0 even while conversations/messages (sourced from the chat tables) show activity.
| Field | Type |
|---|---|
conversations |
integer |
messages |
integer |
agents_created |
integer |
projects_created |
integer |
documents_created |
integer |
active_users |
integer (distinct users active in the window) |
LatencyPercentiles fields
| Field | Type | Description |
|---|---|---|
p50 |
number | null | Median end-to-end latency in milliseconds. null if no data. |
p95 |
number | null | 95th-percentile latency in milliseconds. |
p99 |
number | null | 99th-percentile latency in milliseconds. |
Percentiles cover only non-blocked requests.
TopModelRow fields
| Field | Type | Description |
|---|---|---|
model |
string | Model name. |
provider |
string | Provider slug (e.g. openai, anthropic). |
requests |
integer | Request count in the window. |
blocked |
integer | Requests blocked or scrubbed by a guardrail. |
cached |
integer | Requests served from the response cache. |
input_tokens |
integer | Total prompt tokens consumed by the model. |
output_tokens |
integer | Total completion tokens produced by the model. |
cost_usd |
number | Total cost in USD. |
saved_cost_usd |
number | Cost avoided by cache hits. |
errors |
integer | HTTP ≥ 400 (excluding client-cancel 499). |
client_aborted |
integer | Requests the client cancelled (HTTP 499). |
avg_latency_ms |
number | null | Average end-to-end latency (ms) over non-blocked requests; null if every request was blocked. |
Up to 10 models are returned, ordered by request count descending.
GatewayStats fields
| Field | Type | Description |
|---|---|---|
gateway_id |
string | Gateway UUID. |
gateway |
string | Gateway slug. |
tenant |
string | null | Tenant slug. |
requests |
integer | Total requests. |
blocked |
integer | Requests blocked or scrubbed (counts blocked=1 OR scrub_applied=1). |
cached |
integer | Cache hit count. |
input_tokens |
integer | Total input tokens. |
output_tokens |
integer | Total output tokens. |
cost_usd |
number | Total cost in USD. |
saved_cost_usd |
number | Cost saved by cache hits. |
avg_latency_ms |
number | Average end-to-end latency in milliseconds. |
errors |
integer | Requests that returned an HTTP status >= 400, excluding 499 client-aborted. |
client_aborted |
integer | Requests the client aborted before response headers (HTTP 499). A stream cancelled after headers (Stop pressed mid-answer) is a status 200 row with aborted 1 in the request log and is not counted here. |
UserStats fields
| Field | Type | Description |
|---|---|---|
user_id |
string | User UUID (from the user_id field of the auth token). |
tenant_id |
string | Tenant UUID the user's requests belong to. |
email |
string | null | User email, if known. |
requests |
integer | Total requests attributed to this user. |
blocked |
integer | Requests blocked or scrubbed (counts blocked=1 OR scrub_applied=1). |
cached |
integer | Cache hit count. |
input_tokens |
integer | Total input tokens. |
output_tokens |
integer | Total output tokens. |
cost_usd |
number | Total cost in USD. |
saved_cost_usd |
number | Cost saved by cache hits. |
avg_latency_ms |
number | Average end-to-end latency in milliseconds. |
errors |
integer | Requests that returned an HTTP status >= 400, excluding 499 client-aborted. |
client_aborted |
integer | Requests the client aborted before response headers (HTTP 499). A stream cancelled after headers (Stop pressed mid-answer) is a status 200 row with aborted 1 in the request log and is not counted here. |
by_user is empty by default (anonymized) and is populated only when the caller explicitly opts out with anonymize=0. When populated, it only includes requests made with auth tokens that have a user_id set; anonymous token requests are not included. Up to 50 users are returned, ordered by cost descending. The example below shows the opted-out (identifying) response — the default response has "by_user": [] and "anonymized": true.
CacheEfficiency fields
cache_efficiency summarises input-token caching and input cost over the window (successful requests across all providers). Providers with prompt caching (Anthropic) populate the cache-read/write terms; providers without it (e.g. local models) contribute only their standard input cost at a 0 % hit rate.
| Field | Type | Description |
|---|---|---|
cache_write_tokens |
integer | Prompt-cache write tokens — the total across both TTL tiers (the 1-hour part is counted once; it is a share of the total, not an addition to it). |
cache_read_tokens |
integer | Prompt-cache read tokens. |
standard_input_tokens |
integer | Non-cached input (prompt) tokens. |
uncached_cost_usd |
number | Cost that would have applied without the prompt cache. |
cached_cost_usd |
number | Cost actually incurred reading from the cache. |
cache_hit_pct |
number | null | Share of input tokens served from cache, as a percentage. null when there is no data. |
compaction_tokens_saved |
integer | Input tokens saved by conversation compaction. |
compaction_cost_saved |
number | Cost saved by conversation compaction in USD. |
Example response
{
"percentiles": { "p50": 420, "p95": 1840, "p99": 3210 },
"top_models": [
{ "model": "gpt-4o", "provider": "openai", "requests": 842,
"input_tokens": 1120400, "output_tokens": 208900, "cost_usd": 3.14, "avg_latency_ms": 680 }
],
"by_tenant": [
{
"tenant_id": "ten_abc123", "tenant": "myapp",
"requests": 1423, "blocked": 14, "cached": 82,
"input_tokens": 1840200, "output_tokens": 312400,
"cost_usd": 9.84, "saved_cost_usd": 0.41,
"avg_latency_ms": 820, "errors": 3, "client_aborted": 1
}
],
"by_gateway": [
{
"gateway_id": "gw_xyz789", "gateway": "prod", "tenant": "myapp",
"requests": 1423, "blocked": 14, "cached": 82,
"input_tokens": 1840200, "output_tokens": 312400,
"cost_usd": 9.84, "saved_cost_usd": 0.41,
"avg_latency_ms": 820, "errors": 3, "client_aborted": 1
}
],
"by_channel": [
{ "channel": "ui", "requests": 900, "input_tokens": 1200000, "output_tokens": 210000,
"cost_usd": 6.10, "errors": 2, "avg_latency_ms": 610 },
{ "channel": "api", "requests": 400, "input_tokens": 560000, "output_tokens": 92000,
"cost_usd": 3.20, "errors": 1, "avg_latency_ms": 700 },
{ "channel": "agent", "requests": 123, "input_tokens": 80200, "output_tokens": 10400,
"cost_usd": 0.54, "errors": 0, "avg_latency_ms": 540 }
],
"by_agent": [
{ "agent_id": "agt_support", "agent": "Support Bot", "invocations": 90,
"conversations": 12, "distinct_users": 4, "cost_usd": 0.54 }
],
"by_agent_series": [
{ "agent_id": "agt_support", "agent": "Support Bot", "series": [
{ "ts": 1742457600000, "invocations": 12 },
{ "ts": 1742544000000, "invocations": 0 },
{ "ts": 1742630400000, "invocations": 31 }
] }
],
"by_user": [
{
"user_id": "usr_alice", "tenant_id": "ten_abc123", "email": "alice@myapp.example",
"requests": 341, "blocked": 2, "cached": 18,
"input_tokens": 440200, "output_tokens": 74800,
"cost_usd": 2.31, "saved_cost_usd": 0.09,
"avg_latency_ms": 790, "errors": 1, "client_aborted": 0
}
],
"anonymized": false,
"cache_efficiency": {
"cache_write_tokens": 12800, "cache_read_tokens": 90400,
"standard_input_tokens": 1737000,
"uncached_cost_usd": 27.40, "cached_cost_usd": 0.27,
"cache_hit_pct": 4.9,
"compaction_tokens_saved": 0, "compaction_cost_saved": 0
}
}
GET /stats/timeseries
Returns an array of time-bucketed data points, suitable for rendering sparklines or charts. Buckets with no activity are included as zero-filled entries so the array length is always exactly n.
Scope follows GET /stats: non-admin callers are pinned to their own tenant and own data server-side; a caller with no tenant is refused 403. A platform admin may pass ?tenant_id= to target one tenant, or omit it (empty value included) for the global view. The global series excludes synthetic fixture tenants (e.g. the e2e-synthetic monitoring prober); a tenant-scoped series includes them.
Query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
bucket |
string | 1h |
Bucket size. One of: 5m, 15m, 30m, 1h, 6h, 1d. |
n |
integer | 24 |
Number of buckets to return. Range: 1–168. |
until |
integer | now | End of the time range as a Unix timestamp in seconds. Defaults to the current time. |
tz_offset |
integer | 0 (UTC) |
The viewer's UTC offset in minutes, in the JavaScript Date.getTimezoneOffset() convention: UTC − local (Central European Summer Time = -120, US Eastern Standard Time = 300). Buckets are floored in the viewer's local frame, so the 1d bucket starts at local midnight and a fractional-hour offset (e.g. UTC+5:30) shifts the hourly buckets to local boundaries. Accepted: an integer within [-840, 720] (UTC−14 … UTC+14). Rejected → falls back to UTC (offset 0): any non-numeric, empty, repeated, out-of-range, or missing value (fractional values are floored toward −∞). It never causes an error and never influences authorization — it shifts the bucket boundaries only. |
TimeseriesPoint fields
| Field | Type | Description |
|---|---|---|
ts |
integer | Bucket start time as Unix milliseconds. |
requests |
integer | Total requests in the bucket. |
blocked |
integer | Requests blocked in the bucket. |
rate_limited |
integer | Requests rate-limited in the bucket. |
errors |
integer | Requests that returned an HTTP status >= 400 (excluding 499 client-aborted) in the bucket. |
client_aborted |
integer | Requests the client aborted before response headers (HTTP 499) in the bucket. Streams cancelled after headers are status 200 rows with aborted 1 in the request log, not counted here. |
cost_usd |
number | Total cost in USD for the bucket. |
💡 Note:
tsis in Unix milliseconds (not seconds) for direct compatibility with charting libraries that expect millisecond-precision timestamps (e.g. Chart.js, Recharts, Grafana).
Example response
[
{"ts": 1742540400000, "requests": 42, "blocked": 1, "rate_limited": 0, "errors": 1, "client_aborted": 0, "cost_usd": 0.0314},
{"ts": 1742544000000, "requests": 67, "blocked": 0, "rate_limited": 0, "errors": 0, "client_aborted": 0, "cost_usd": 0.0521},
{"ts": 1742547600000, "requests": 0, "blocked": 0, "rate_limited": 0, "errors": 0, "client_aborted": 0, "cost_usd": 0.0},
{"ts": 1742551200000, "requests": 18, "blocked": 2, "rate_limited": 1, "errors": 2, "client_aborted": 1, "cost_usd": 0.0119}
]
GET /health/dashboard
GET /admin/v1/health/dashboard
Returns the aggregated operational-health summary that the Health dashboard renders in a single call — the synthetic-probe results, the last-24-hours non-2xx and client-aborted request counts, the phoned-home client errors, and the other failure surfaces shown there.
Requires the HEALTH_VIEW permission; the Health dashboard itself is available to the admin role only. A backend failure returns 500.
Platform web-search key (AGF-2874). The summary carries web_search_platform_key_missing (boolean): true when the platform Linkup key (the trial_linkup_api_key setting — applied to every keyless Linkup gateway on every plan) is unset while at least one production gateway of a non-deleted tenant carries a keyless Linkup web_search block — i.e. web search is silently dark for it. false when the key is set, or when nothing depends on it yet. The companion object web_search_platform_key gives the inputs: { "key_set": boolean, "keyless_production_gateways": integer }. If the count could not be read, all three are absent and web_search_platform_key_error (string) names the fault — a failed read is never reported as healthy; the dashboard renders "unknown". The same condition drives the [platform_alert] signal web_search_platform_key_missing (checked every 10 minutes, ERR-logged each check, notified at most every 6 hours).
Plan-aware model copy (AGF-2930). The summary carries plan_model_copy_stale (boolean): true when, for at least one offered model set served in the last 7 days, the set's current copy generation was rejected by the validator (or the set's catalog facts could not be read as valid input) — the model picker then shows that set's hand-written seed line, a checked catalog tagline or the model name instead (see Tagline resolution). plan_model_copy_indistinct (boolean) is true when a current set offers two models no fact tells apart (same price rank, context, capabilities, hosting and speed) — a plan-design problem, not a generator fault. The companion object plan_model_copy gives the counts: { "stale_sets": integer, "indistinct_sets": integer }. If the status could not be read, all three are absent and plan_model_copy_error (string) names the fault — the dashboard renders "unknown", never healthy.
Cost-of-goods (COGS)
The COGS endpoints aggregate the per-leg billing ledger (request_log_legs) into cost breakdowns by model, by tenant, by service, or by API key. A single inference request may produce several legs (the primary call, any fallback, plus side services such as guardrails or web search), so these views give a finer-grained cost picture than the request-level Stats summary.
All four endpoints share the same time-window selection:
Query parameters (all COGS endpoints)
| Parameter | Type | Default | Description |
|---|---|---|---|
range |
string | 30d |
Relative window ending now. Format <n><unit> where unit is d (days), h (hours), or w (weeks) — e.g. 24h, 2w, 90d. Ignored when from/to are given. |
from |
integer | — | Explicit window start as a Unix timestamp in seconds. Must be paired with to. |
to |
integer | — | Explicit window end as a Unix timestamp in seconds. Must be paired with from. |
💡 Note: When neither
from/tonorrangeis given, the window defaults to the last 30 days plus one day into the future (the extra day captures in-flight legs still being written).
CSV export (format) — cost-allocation endpoints
The three cost-allocation endpoints — /cogs/per-token, /cogs/per-user, and /cogs/per-agent — accept a format parameter for a machine-readable download:
| Parameter | Type | Default | Description |
|---|---|---|---|
format |
string | json |
Response encoding. json (or absent) returns the JSON array documented per endpoint. csv returns the same rows as text/csv; charset=utf-8 with a fixed header row (the response-field order below), served as a file attachment (Content-Disposition: attachment). Any other value — an unknown string, a repeated ?format= (which arrives as a list), or a bare ?format (no value) — is rejected with 400 { "error": "invalid format (json or csv)" } (fail-closed allowlist), validated before the aggregate query runs. |
Notes on the CSV encoding:
- The column order is fixed and equals the endpoint's response-field order; a missing value renders as an empty cell.
- Cells are RFC-4180 quoted, and any value that begins with
=,+,-,@, TAB, or CR is prefixed with a single quote to neutralize spreadsheet formula injection (CWE-1236) — so a user email or agent name such as=SUM(...)cannot execute when the file is opened in a spreadsheet. - No byte-order mark (BOM) is emitted.
- The download filename is a fixed per-endpoint constant (
cogs-per-token.csv/cogs-per-user.csv/cogs-per-agent.csv); the server setsAccess-Control-Expose-Headers: Content-Dispositionso a browser SPA fetching cross-origin can read it. - The
bucket,tenant_id, scope, andbillable_onlyrules are identical to the JSON response;formatonly changes the encoding.
GET /cogs/per-model
Cost and token totals grouped by (provider, model, tier, byok), ordered by cost descending. Required role: admin or tenant_admin.
Additional query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
tenant_id |
string | — | Restrict to one tenant. Honoured only for callers with the admin role; every other caller (including tenant_admin) is scoped to their own tenant server-side, and a caller with no tenant is refused 403. |
billable_only |
string | 1 |
Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier or internal). |
Response
Array of rows. Returns 500 { "error": ... } on a storage failure.
| Field | Type | Description |
|---|---|---|
provider |
string | Provider that served the leg. |
model |
string | Model used for the leg. |
tier |
string | null | Pricing tier recorded for the leg. |
byok |
integer | 1 if the leg used a bring-your-own-key credential, 0 otherwise. |
call_count |
integer | Number of legs in this group. |
in_t |
integer | Sum of input (prompt) tokens. |
out_t |
integer | Sum of output (completion) tokens. |
cache_read_t |
integer | Sum of prompt-cache read tokens. |
cache_write_t |
integer | Sum of prompt-cache creation (write) tokens. |
cost_usd |
number | Total cost in USD for the group. |
partial_count |
integer | Number of legs flagged partial (interrupted / incomplete). |
[
{
"provider": "anthropic",
"model": "claude-opus-4-6",
"tier": null,
"byok": 0,
"call_count": 412,
"in_t": 1840200,
"out_t": 312400,
"cache_read_t": 90400,
"cache_write_t": 12800,
"cost_usd": 9.84,
"partial_count": 2
}
]
GET /cogs/per-tenant
Cost and token totals grouped by tenant, billable legs only, ordered by cost descending. Required role: platform admin only — lower-privileged callers receive 403 { "error": "forbidden: platform admin only" }.
Response
Array of rows. Returns 500 { "error": ... } on a storage failure.
| Field | Type | Description |
|---|---|---|
tenant_id |
string | Tenant UUID. |
call_count |
integer | Number of billable legs for the tenant. |
cost_usd |
number | Total cost in USD. |
in_t |
integer | Sum of input tokens. |
out_t |
integer | Sum of output tokens. |
[
{ "tenant_id": "ten_abc123", "call_count": 1423, "cost_usd": 14.30, "in_t": 2750000, "out_t": 480000 }
]
GET /cogs/per-service
Per-service usage grouped by (provider, service, unit_kind), ordered by call count descending. Counts only legs that carry a non-null service (e.g. web_search, guardrail). Required role: admin or tenant_admin.
Additional query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
tenant_id |
string | — | Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's — or the global, all-tenant — per-service totals. |
Plus the shared range / from / to window parameters above.
Response
Array of rows. Returns 500 { "error": ... } on a storage failure.
| Field | Type | Description |
|---|---|---|
provider |
string | Provider that served the service leg. |
service |
string | Service name (e.g. web_search, guardrail). |
unit_kind |
string | null | Unit the service is metered in (e.g. requests, tokens). |
call_count |
integer | Number of legs in this group. |
units_total |
number | Sum of metered units. |
avg_latency_ms |
number | Average leg latency in milliseconds. |
in_t |
integer | Sum of input tokens (where applicable). |
out_t |
integer | Sum of output tokens (where applicable). |
[
{ "provider": "anthropic", "service": "web_search", "unit_kind": "requests", "call_count": 88, "units_total": 88, "avg_latency_ms": 640, "in_t": 0, "out_t": 0 }
]
GET /cogs/per-token
Cost and usage aggregated per API key (token) × period (Zeitraum) × KI-Modell, ordered by cost descending. This is the billing-disclosure view: it answers "what did each API key spend, on which model, in which period." Required role: admin or tenant_admin.
Only genuine LLM legs are counted: legs with a NULL model (non-LLM side services such as PII scrubbing or guardrails) are excluded, so the per-model figures are clean. Customer-paid (BYOK) legs are kept on their own rows via the byok dimension and are never summed into platform cost. A single user request that fans out into several same-model legs (e.g. a tool loop) counts as one request.
Additional query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
bucket |
string | month |
Period granularity for the period column. One of day, week, month, total. Any other value is rejected with 400 { "error": "invalid bucket" } (fail-closed allowlist — the value is never interpolated into SQL). total collapses the whole window into a single period labelled "total". |
tenant_id |
string | — | Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's cost rows. |
token_id |
string | — | Restrict to a single API key. Must be a single scalar value — a repeated param (?token_id=a&token_id=b) is rejected with 400 { "error": "token_id must be a single value" } (fail-closed; it is never bound into SQL). Combined with the mandatory tenant scope: a token_id belonging to another tenant returns an empty result. |
billable_only |
string | 1 |
Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier / self-hosted). Non-billable legs still obey the model IS NOT NULL rule. |
Plus the shared range / from / to window parameters above.
Response
Array of rows. Returns 500 { "error": ... } on a storage failure, 400 { "error": "invalid bucket" } on a bad bucket.
| Field | Type | Description |
|---|---|---|
token_id |
string | API-key id the legs were attributed to. Empty string ("") means non-API-key traffic — interactive UI / playground sessions that carry no API key. |
token_label |
string | Human label of the API key (from auth_token.label). Empty string if the key has no label or was deleted. |
period |
string | Period bucket. Format depends on bucket: 2026-07-15 (day), 2026-W29 (ISO year-week), 2026-07 (month), or total. Boundaries are computed in the gateway database session timezone (UTC on the standard deployment). Note: the week bucket uses the ISO week-numbering year, so late-December usage can fall in the following calendar year's W01 (e.g. 2025-12-31 → 2026-W01) — while day/month use the calendar year. |
provider |
string | Provider that served the model. |
model |
string | Model id (never null — non-model side-service legs are excluded). |
byok |
integer | 1 if the legs used a bring-your-own-key credential (customer-paid), 0 for platform-paid. Kept as a separate dimension so customer and platform cost never merge. |
requests |
integer | Distinct end-user requests in this group (COUNT(DISTINCT request_log_id)), not the leg count. This is a per-row count and is not additive across rows: a single request that fans out across models (a fallback) or across a period boundary is counted once in each row it touches, so summing the requests column overstates the true request total. |
input_tokens |
integer | Sum of input (prompt) tokens. |
output_tokens |
integer | Sum of output (completion) tokens. |
total_tokens |
integer | input_tokens + output_tokens. |
cost_usd |
number | Total cost in USD for the group, rounded to 6 decimals. |
[
{
"token_id": "8f3c2a10-6b4e-4d21-9a77-1e5c0b9d4a02",
"token_label": "prod-key",
"period": "2026-07",
"provider": "anthropic",
"model": "claude-opus-4-7",
"byok": 0,
"requests": 412,
"input_tokens": 1840200,
"output_tokens": 312400,
"total_tokens": 2152600,
"cost_usd": 9.84
}
]
GET /cogs/per-user
Cost and usage aggregated per user × period (Zeitraum) × KI-Modell, ordered by cost descending. This is the per-user cost-allocation view: it answers "which user spent how much, on which model, in which period." Required role: admin, tenant_admin, or finance (the COGS_PER_USER_VIEW permission). It is the exact cost twin of /cogs/per-token — same period buckets, same tenant scoping, same row shape — with the grouping key swapped from the API key to the user.
Only genuine LLM legs are counted: legs with a NULL model (non-LLM side services such as PII scrubbing or guardrails) are excluded, so the per-model figures are clean. Customer-paid (BYOK) legs are kept on their own rows via the byok dimension and are never summed into platform cost. A single user request that fans out into several same-model legs (e.g. a tool loop) counts as one request.
Cost allocation is inherently per-identity: this endpoint returns each user's email within the caller's scope. A tenant_admin can only ever see the users of their own tenant (server-side tenant scope, fail-closed); only an admin may scope to another tenant or read across all tenants.
Additional query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
bucket |
string | month |
Period granularity for the period column. One of day, week, month, total. Any other value is rejected with 400 { "error": "invalid bucket" } (fail-closed allowlist — the value is never interpolated into SQL). total collapses the whole window into a single period labelled "total". |
tenant_id |
string | — | Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's cost rows. |
user_id |
string | — | Restrict to a single user. Must be a single scalar value — a repeated param (?user_id=a&user_id=b) is rejected with 400 { "error": "user_id must be a single value" } (fail-closed; never bound into SQL). Combined with the mandatory tenant scope: a user_id belonging to another tenant returns an empty result. |
billable_only |
string | 1 |
Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier / self-hosted). Non-billable legs still obey the model IS NOT NULL rule. |
Plus the shared range / from / to window parameters above.
Response
Array of rows. Returns 500 { "error": ... } on a storage failure, 400 { "error": "invalid bucket" } on a bad bucket.
| Field | Type | Description |
|---|---|---|
user_id |
string | User id the legs were attributed to. Empty string ("") means traffic with no attributed user — e.g. API-key requests that carry no user context. |
user_email |
string | Email of the user (from user.email). Empty string if the user has no email or was deleted. |
period |
string | Period bucket. Format depends on bucket: 2026-07-15 (day), 2026-W29 (ISO year-week), 2026-07 (month), or total. Boundaries are computed in the gateway database session timezone (UTC on the standard deployment). Note: the week bucket uses the ISO week-numbering year, so late-December usage can fall in the following calendar year's W01 (e.g. 2025-12-31 → 2026-W01) — while day/month use the calendar year. |
provider |
string | Provider that served the model. |
model |
string | Model id (never null — non-model side-service legs are excluded). |
byok |
integer | 1 if the legs used a bring-your-own-key credential (customer-paid), 0 for platform-paid. Kept as a separate dimension so customer and platform cost never merge. |
requests |
integer | Distinct end-user requests in this group (COUNT(DISTINCT request_log_id)), not the leg count. This is a per-row count and is not additive across rows: a single request that fans out across models (a fallback) or across a period boundary is counted once in each row it touches, so summing the requests column overstates the true request total. |
input_tokens |
integer | Sum of input (prompt) tokens. |
output_tokens |
integer | Sum of output (completion) tokens. |
total_tokens |
integer | input_tokens + output_tokens. |
cost_usd |
number | Total cost in USD for the group, rounded to 6 decimals. |
[
{
"user_id": "3a1f8c22-9d4e-4b70-8c11-2f6a0e7b5d13",
"user_email": "alex@example.com",
"period": "2026-07",
"provider": "anthropic",
"model": "claude-opus-4-7",
"byok": 0,
"requests": 128,
"input_tokens": 540200,
"output_tokens": 96400,
"total_tokens": 636600,
"cost_usd": 3.11
}
]
GET /cogs/per-agent
Cost and usage aggregated per agent × period (Zeitraum) × KI-Modell, ordered by cost descending. This is the per-agent cost-allocation view: it answers "which agent spent how much, on which model, in which period." Required role: admin or tenant_admin. It is the agent-dimension twin of /cogs/per-user — same period buckets, same tenant scoping, same row shape — with the grouping key swapped from the user to the agent.
Only agent-invoked traffic is included (requests that carry no agent are excluded). Agents are grouped by their clone-lineage root, so a shared agent and all of its clones roll up under one row group — matching the per-agent stats in /stats/analytics. Only genuine LLM legs are counted: legs with a NULL model (non-LLM side services such as PII scrubbing or guardrails) are excluded, so the per-model figures are clean. Customer-paid (BYOK) legs are kept on their own rows via the byok dimension and are never summed into platform cost. A single agent request that fans out into several same-model legs (e.g. a tool loop) counts as one request.
Cost and token figures are tenant-bounded (server-side tenant scope, fail-closed): a tenant_admin can only ever see their own tenant's agent spend; only an admin may scope to another tenant or read across all tenants.
Additional query parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
bucket |
string | month |
Period granularity for the period column. One of day, week, month, total. Any other value is rejected with 400 { "error": "invalid bucket" } (fail-closed allowlist — the value is never interpolated into SQL). total collapses the whole window into a single period labelled "total". |
tenant_id |
string | — | Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's cost rows. |
agent_id |
string | — | Restrict to a single agent, matched against the clone-lineage root id. Must be a single scalar value — a repeated param (?agent_id=a&agent_id=b) is rejected with 400 { "error": "agent_id must be a single value" } (fail-closed; never bound into SQL). Combined with the mandatory tenant scope: an agent_id whose traffic belongs to another tenant returns an empty result. |
billable_only |
string | 1 |
Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier / self-hosted). Non-billable legs still obey the model IS NOT NULL rule. |
Plus the shared range / from / to window parameters above.
Response
Array of rows. Returns 500 { "error": ... } on a storage failure, 400 { "error": "invalid bucket" } on a bad bucket.
| Field | Type | Description |
|---|---|---|
agent_id |
string | Clone-lineage root id the legs were attributed to (the stamped root, falling back to the agent's own id for pre-lineage rows). |
agent |
string | Display name of the root agent (agent.name, falling back to its slug, then the raw root id if the agent was deleted). For a cross-tenant cloned agent the label may resolve to the source agent's name; the cost and token figures always count only the caller's own tenant's legs. |
period |
string | Period bucket. Format depends on bucket: 2026-07-15 (day), 2026-W29 (ISO year-week), 2026-07 (month), or total. Boundaries are computed in the gateway database session timezone (UTC on the standard deployment). Note: the week bucket uses the ISO week-numbering year, so late-December usage can fall in the following calendar year's W01 (e.g. 2025-12-31 → 2026-W01) — while day/month use the calendar year. |
provider |
string | Provider that served the model. |
model |
string | Model id (never null — non-model side-service legs are excluded). |
byok |
integer | 1 if the legs used a bring-your-own-key credential (customer-paid), 0 for platform-paid. Kept as a separate dimension so customer and platform cost never merge. |
requests |
integer | Distinct end-user requests in this group (COUNT(DISTINCT request_log_id)), not the leg count. This is a per-row count and is not additive across rows: a single request that fans out across models (a fallback) or across a period boundary is counted once in each row it touches, so summing the requests column overstates the true request total. |
input_tokens |
integer | Sum of input (prompt) tokens. |
output_tokens |
integer | Sum of output (completion) tokens. |
total_tokens |
integer | input_tokens + output_tokens. |
cost_usd |
number | Total cost in USD for the group, rounded to 6 decimals. |
[
{
"agent_id": "ddc230d8-78ef-11f1-a3ae-6c92cf1260f0",
"agent": "Support Triage",
"period": "2026-07",
"provider": "anthropic",
"model": "claude-opus-4-7",
"byok": 0,
"requests": 64,
"input_tokens": 210400,
"output_tokens": 38800,
"total_tokens": 249200,
"cost_usd": 1.42
}
]
Example curl requests
⭐ Example: The following examples show common query patterns for the timeseries endpoint.
Today's requests by hour (last 24 hours)
Yesterday's requests by hour
# Set 'until' to the end of yesterday (start of today in Unix seconds)
YESTERDAY_END=$(date -d "today 00:00:00" +%s)
curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=1h&n=24&until=${YESTERDAY_END}"
Last 7 days by day
Last hour in 5-minute buckets
Last 30 days in 6-hour buckets
GET /stats/reactivation
The platform-wide roll-up of the member-reactivation e-mails, shown as the Member reactivation card on the platform dashboard.
Authorization: platform admin only (cross-tenant data; 403 otherwise).
Response: mails_sent_7d, mails_sent_30d, reactivated_7d (members who became active within
7 days of a mail sent in the last 30 days), opt_outs_total, opt_outs_30d, and top_tenants —
the ten most-mailed workspaces of the last 30 days as [{ tenant_id, slug, mails_sent_30d }]
(an empty array when nothing was sent). Counts come from the send log only — no per-member
classification runs fleet-wide; the per-tenant states live on
GET /tenants/{id}/reactivation-metrics. A database
fault is 503.
See also
- Finance API — the
financerole and its per-user cost access (GET /cogs/per-user) - Logs API
- Models & Pricing API
- Error Codes
- Budget & Quota Enforcement