Skip to content

Stats API

The Stats API exposes aggregated request metrics. Use GET /stats for the dashboard summary view (today, yesterday, last 7 days, last hour, last minute) and GET /stats/timeseries for time-bucketed data suitable for charts.

Base URL: https://<your-gateway-host>/admin/v1


GET /stats

Returns a snapshot of request counts, token usage, cost, and latency across several time windows. Also includes a per-tenant breakdown and the most recent log entries.

Every caller below admin/tenant_admin (member, viewer, ki_manager, demouser) is scoped to their own data. Non-admin callers are pinned to their own tenant server-side (a spoofed tenant_id is ignored); a caller with no tenant is refused 403 { "error": "forbidden" } (the query never runs). A platform admin may target one tenant with ?tenant_id= or omit it (and an empty value) for the global, all-tenant view.

The calendar-day windows (today, yesterday, last_7d) are computed in the viewer's timezone, supplied by the tz_offset query parameter — so "today" means the viewer's local calendar day, not a UTC day. The rolling windows (hour, last_min) are relative to the current instant and are timezone-independent.

Query parameters

Parameter Type Default Description
tenant_id string (UUID) caller's tenant Platform admin only: target one tenant, or omit / empty for the global view. Ignored (server-pinned) for every other role.
tz_offset integer 0 (UTC) The viewer's UTC offset in minutes, in the JavaScript Date.getTimezoneOffset() convention: UTC − local, so Central European Summer Time (UTC+2) is -120 and US Eastern Standard Time (UTC−5) is 300. Used only to anchor the today / yesterday / last_7d calendar-day boundaries to the viewer's local midnight. Accepted: an integer within [-840, 720] (UTC−14 … UTC+14). Rejected → falls back to UTC (offset 0): any non-numeric, empty, repeated, out-of-range, or missing value (fractional values are floored toward −∞). It never causes an error and never influences authorization — it shifts a display window only.

The GET /stats/timeseries series that back the dashboard sparkline glyphs accept the same tz_offset parameter and bucket in the viewer's local frame — so the daily (1d) bucket and, for a fractional-hour offset (e.g. India, UTC+5:30), the hourly buckets align with the viewer's local day/hour rather than UTC. For whole-hour offsets (e.g. Central Europe, US) the hourly buckets coincide with local hours either way. Omit tz_offset (or send a rejected value) and buckets fall back to UTC.

curl "https://<your-gateway-host>/admin/v1/stats?tz_offset=-120"

Response structure

{
  "today":     { ...PeriodStats },
  "yesterday": { ...PeriodStats },
  "last_7d":   { ...PeriodStats },
  "hour":      { ...PeriodStats },
  "last_min":  { ...PeriodStats },
  "by_tenant": [ ...TenantStats ],
  "recent":    [ ...RecentEntry ],
  "recent_blocked": [ ...RecentBlockedEntry ]
}

recent and recent_blocked resemble LogEntry but are not the same shape. recent (up to 10 newest entries) omits id and instead adds the tenant and gateway slugs alongside the IDs. recent_blocked (up to 20 newest blocked/scrubbed/flagged entries) is a narrow subset carrying only the block/guardrail fields plus prompt_scrubbed and response_raw.

PeriodStats fields

Field Type Description
requests integer Total inference requests in the period.
cached integer Requests served from cache (no provider call made).
blocked integer Requests blocked by auth, rate limit, quota, or detectors.
rate_limited integer Requests that hit a rate limit in the period.
scrubbed integer Requests where a guardrail scrubbed PII (personally identifiable information) from the payload but allowed the request through.
flagged integer Requests where a guardrail raised a flag but took no blocking or scrubbing action.
input_tokens integer Total prompt tokens consumed.
output_tokens integer Total completion tokens generated.
cost_usd number Total cost in USD (from model pricing table).
saved_cost_usd number Cost saved by cache hits (would-have-been cost of cached requests).
avg_latency_ms number Average end-to-end request latency in milliseconds.
avg_upstream_latency_ms number Average time waiting for the upstream provider, excluding gateway overhead.

TenantStats fields (GET /stats)

by_tenant in GET /stats is a summary view with the following fields:

Field Type Description
tenant_id string Tenant UUID.
tenant string Tenant slug.
requests integer Total requests.
input_tokens integer Total input tokens.
output_tokens integer Total output tokens.
cost_usd number Total cost in USD.
budget_remaining number Remaining budget in USD for the tenant's current period (the live-monitor "Quota Left" column). Absent when the tenant has no budget.
budget_remaining_unknown boolean true when the remaining budget could not be computed (a transient spend-read failure) — distinct from "no limit".

The full TenantStats shape (with blocked, cached, errors, avg_latency_ms, etc.) is returned by GET /stats/analytics. See below.

Example response

{
  "today": {
    "requests": 1423,
    "cached": 82,
    "blocked": 14,
    "input_tokens": 1840200,
    "output_tokens": 312400,
    "cost_usd": 9.84,
    "saved_cost_usd": 0.41,
    "avg_latency_ms": 820,
    "avg_upstream_latency_ms": 760
  },
  "yesterday": {
    "requests": 2101,
    "cached": 134,
    "blocked": 22,
    "input_tokens": 2750000,
    "output_tokens": 480000,
    "cost_usd": 14.30,
    "saved_cost_usd": 0.67,
    "avg_latency_ms": 905,
    "avg_upstream_latency_ms": 850
  },
  "last_7d": { "requests": 11200, "cost_usd": 68.12, "..." : "..." },
  "hour":    { "requests": 61, "cost_usd": 0.38, "...": "..." },
  "last_min": { "requests": 2, "cost_usd": 0.01, "...": "..." },
  "by_tenant": [
    {
      "tenant_id": "ten_abc123",
      "tenant": "myapp",
      "requests": 1423,
      "input_tokens": 1840200,
      "output_tokens": 312400,
      "cost_usd": 9.84
    }
  ],

  "recent": [ { "...": "LogEntry" } ],
  "recent_blocked": [ { "...": "LogEntry" } ]
}

For the related LogEntry field definitions see the Logs API. Note that the recent and recent_blocked rows are not full LogEntry objects (see the field notes above).


GET /stats/analytics

Returns latency percentiles, top-model breakdown, and per-tenant, per-gateway, per-channel, per-agent, and per-user cost summaries, plus activity counters. Used by the analytics dashboard view.

On the global, all-tenant view (a platform admin who omits tenant_id), every grouping and counter excludes synthetic fixture tenants — e.g. the e2e-synthetic monitoring prober — so the platform analytics reflect real usage. A tenant-scoped view (?tenant_id=…) still returns them.

curl "https://<your-gateway-host>/admin/v1/stats/analytics?since=1742544000000"

Query parameters

Parameter Type Default Description
since integer 24 hours ago Start of the analysis window as Unix milliseconds.
until integer now Optional upper bound as Unix milliseconds. Used for closed windows such as "yesterday".
tenant_id string — Filter by tenant. Honoured only for callers with the admin role. Non-admin callers are always scoped to their own tenant.
anonymize boolean true (anonymized) Anonymized by default. Unless the caller explicitly opts out with anonymize=0 (or false), the response omits the per-user grouping (by_user) entirely — no individual identifiers (email/user_id) are returned. Any other value (absent, 1, true, or malformed) keeps the response anonymized (fail-closed). Aggregate groupings (channel, agent, model, gateway, tenant) and counters are always included. The flag is scoped to this endpoint. Note: in a single-user tenant, per-agent/per-model aggregates can still imply one person's usage.
group_id string — Restrict the whole response to the members of one user group. For a KI-Manager this must be a group they manage (see below); admin/tenant_admin may narrow to any group in their tenant scope. An unknown or out-of-scope group_id (including a group_id paired with a mismatched tenant_id) returns 404 not found; a KI-Manager requesting a same-tenant group they do not manage returns 403 forbidden.
tz_offset integer 0 (UTC) The viewer's UTC offset in minutes, in the JavaScript Date.getTimezoneOffset() convention: UTC − local (Central European Summer Time = -120, US Eastern Standard Time = 300). Applies only to the per-agent daily series (by_agent_series), whose buckets are floored in the viewer's local frame so each daily bar starts at the viewer's local midnight — aligning it with the calendar-day window (since/until) the caller computes with the same offset. Accepted: an integer within [-840, 720] (UTC−14 … UTC+14). Rejected → falls back to UTC (offset 0): any non-numeric, empty, repeated, out-of-range, or missing value (fractional values are floored toward −∞). It never causes an error and never influences authorization — it shifts the bucket boundaries only.

💡 Note: /stats/analytics requires the admin, tenant_admin, or ki_manager role; lower-privileged callers (member, viewer, demouser) receive 403 forbidden. The anonymize=0 opt-out is therefore confined to those privileged roles, and only ever reveals users within the caller's authorized scope: admin/tenant_admin see their own tenant (a tenant_admin is pinned to it), while a ki_manager is pinned to the members of the groups they manage. Any role that could send anonymize=0 without authorization is already blocked with 403 before the handler runs; the client value never decides access.

KI-Manager group scoping

A KI-Manager may optionally see the statistics of the groups they manage. Unlike admin/tenant_admin (who see the whole tenant/workspace), a KI-Manager is always scoped to a group's members and can never see tenant-wide or another group's data:

  • "Groups they manage" are the groups the KI-Manager created (user_group.created_by). This is the same boundary as the group-management API (can_manage_group). Limitation: the data model has no separate "group manager" or delegation concept, so a group created by an admin and handed to a KI-Manager to run — or a group whose creator account was later deleted (created_by becomes NULL) — is not visible to the KI-Manager. Only admin/tenant_admin see those.
  • Without group_id: the response covers the union of the members of all groups the KI-Manager manages. A KI-Manager who manages no (non-empty) groups gets an all-zero response — never a tenant-wide fallback.
  • With group_id: the response covers only that group's members, and only if the KI-Manager manages it (else 403/404 as above).
  • "A group's statistics" = activity by the group's members. All groupings and counters are filtered to request_log rows / rows owned by the member set. Because request_log carries no group dimension, this is the only coherent definition. Consequences to keep in mind:
  • by_agent is filtered on the run-as owner (user_id) — which, because invoke is owner-only, is the real member who invoked. Rows roll up over the clone lineage (see AgentStats), so a KI-Manager sees, for each shared agent, the invocations/conversations/distinct-users contributed by their members' clones of it (the display name is the root agent's, which may be owned outside the group).
  • active_users is windowed: under a group filter it is the count of the group's members active in the window (distinct members with ≥1 request in the period), not the group's current member count.

Response structure

{
  "percentiles": { ...LatencyPercentiles },
  "top_models":  [ ...TopModelRow ],
  "by_tenant":   [ ...TenantStats ],
  "by_gateway":  [ ...GatewayStats ],
  "by_channel":  [ ...ChannelStats ],
  "by_agent":    [ ...AgentStats ],
  "by_agent_series": [ ...AgentSeries ],
  "by_user":     [ ...UserStats ],
  "counters":    { ...ActivityCounters },
  "anonymized":  true,
  "cache_efficiency": { ...CacheEfficiency }
}

ChannelStats fields (UI vs API vs agent)

Each row is the usage attributed to one access channel. The channel is derived: agent for agent invocations, ui for chat-session (playground) tokens, api otherwise.

Field Type Description
channel string ui, api, or agent.
requests integer Request count for the channel.
blocked integer Requests blocked or scrubbed by a guardrail.
cached integer Requests served from the response cache.
input_tokens integer Total prompt tokens consumed on the channel.
output_tokens integer Total completion tokens produced on the channel.
cost_usd number Total cost.
saved_cost_usd number Cost avoided by cache hits.
errors integer HTTP ≥ 400 (excluding client-cancel 499).
client_aborted integer Requests the client cancelled (HTTP 499).
avg_latency_ms integer Average latency for non-blocked requests.

AgentStats fields (per agent, rolled up over the clone lineage)

Each row is one lineage root — a shared/original agent plus every clone made from it. Rows are keyed by agent_id = the root's id; a clone's usage rolls into its root (the counting model below). Filtering (tenant, KI-Manager group scope) applies before the roll-up.

Field Type Description
agent_id string The lineage root agent's id (a shared agent + its clones share one row).
agent string Root agent name (or slug). A hard-deleted root falls back to the raw id.
invocations integer Invocations across the lineage in the window.
conversations integer Distinct chat sessions the lineage was used in. Headless invokes (API/schedule) carry no session id and count as invocations, not conversations — so this is 0 for agents only ever invoked headlessly, and becomes non-zero when an agent is used inside a chat thread. Read from request_log_legs.conversation_id; no per-invoke session id is fabricated.
distinct_users integer Distinct people who used the lineage — see the counting model below.
cost_usd number Total cost across the lineage.

Per-agent "distinct users" — the counting model (supersedes the earlier model). Invoke is owner-only (agent_invoke refuses a non-owner caller), and the catalog "use" action is clone-then-own (a member who wants a shared agent clones it into a new agent they own, then invokes that). So a single agent's distinct-invoker count is structurally ≤1 — the "Scheinwert". The fix is not to drop the metric but to count across the clone lineage: each member who uses a shared agent owns a distinct clone and invokes it as owner (invoker_id = that member), so COUNT(DISTINCT invoker_id) over the root and its clones = the true number of people using it. A clone records its lineage root in agent.source_agent_id (set by the clone flow, flattened so the tree is one level); every request stamps request_log.agent_root_id, so the count survives the clone being deleted. Caveats (bid-relevant): lineage is recorded from the release that introduced it — pre-existing clones are not retro-linked; and because there is no server clone endpoint yet, the lineage pointer is client-declared within the caller's visibility (a member can omit it → undercount, or attribute a clone to any agent they can already see) — so distinct_users is a best-effort, self-reported figure, not server-authoritative. Unattended runs (scheduler/service token) log a NULL invoker and are excluded from the count.

AgentSeries fields (per-agent calls over time)

One entry per top agent (up to 8, by invocation count) carrying a daily-bucketed invocation time series. The series is dense and zero-filled across the window, so every day is present even with no calls. Timestamps are milliseconds since epoch (bucket start); when the tz_offset query parameter is supplied the bucket start is the viewer's local midnight rather than UTC midnight (so the bars align with the viewer-local calendar-day windows). Empty ([]) when there are no agent invocations in the window.

Field Type Description
agent_id string The saved agent's id.
agent string Agent name (or slug).
series array Daily points: { "ts": <ms>, "invocations": <integer> }.

ActivityCounters fields

Counts for the window, scoped to the tenant filter. *_created count items created in the window; active_users counts distinct users active in the window — users with at least one logged request (UI or API channel) between since and until. System/unattributed requests (no user_id) are excluded; duplicate requests by one user count once. It is not the provisioned-user headcount, so it can differ from the number of user rows (e.g. a user active in the window then deleted still counts).

Two caveats follow from active_users being derived from request_log: - No explicit window (as in the tenant export's usage-statistics snapshot) → active_users reflects the trailing 24h, consistent with the sibling *_created counters, which use the same default. The provisioned headcount remains available separately (e.g. the export's users array). - Request logging disabled for a tenant (request_logging_disabled) → no request_log rows are written, so active_users reads 0 even while conversations/messages (sourced from the chat tables) show activity.

Field Type
conversations integer
messages integer
agents_created integer
projects_created integer
documents_created integer
active_users integer (distinct users active in the window)

LatencyPercentiles fields

Field Type Description
p50 number | null Median end-to-end latency in milliseconds. null if no data.
p95 number | null 95th-percentile latency in milliseconds.
p99 number | null 99th-percentile latency in milliseconds.

Percentiles cover only non-blocked requests.

TopModelRow fields

Field Type Description
model string Model name.
provider string Provider slug (e.g. openai, anthropic).
requests integer Request count in the window.
blocked integer Requests blocked or scrubbed by a guardrail.
cached integer Requests served from the response cache.
input_tokens integer Total prompt tokens consumed by the model.
output_tokens integer Total completion tokens produced by the model.
cost_usd number Total cost in USD.
saved_cost_usd number Cost avoided by cache hits.
errors integer HTTP ≥ 400 (excluding client-cancel 499).
client_aborted integer Requests the client cancelled (HTTP 499).
avg_latency_ms number | null Average end-to-end latency (ms) over non-blocked requests; null if every request was blocked.

Up to 10 models are returned, ordered by request count descending.

GatewayStats fields

Field Type Description
gateway_id string Gateway UUID.
gateway string Gateway slug.
tenant string | null Tenant slug.
requests integer Total requests.
blocked integer Requests blocked or scrubbed (counts blocked=1 OR scrub_applied=1).
cached integer Cache hit count.
input_tokens integer Total input tokens.
output_tokens integer Total output tokens.
cost_usd number Total cost in USD.
saved_cost_usd number Cost saved by cache hits.
avg_latency_ms number Average end-to-end latency in milliseconds.
errors integer Requests that returned an HTTP status >= 400, excluding 499 client-aborted.
client_aborted integer Requests the client aborted before response headers (HTTP 499). A stream cancelled after headers (Stop pressed mid-answer) is a status 200 row with aborted 1 in the request log and is not counted here.

UserStats fields

Field Type Description
user_id string User UUID (from the user_id field of the auth token).
tenant_id string Tenant UUID the user's requests belong to.
email string | null User email, if known.
requests integer Total requests attributed to this user.
blocked integer Requests blocked or scrubbed (counts blocked=1 OR scrub_applied=1).
cached integer Cache hit count.
input_tokens integer Total input tokens.
output_tokens integer Total output tokens.
cost_usd number Total cost in USD.
saved_cost_usd number Cost saved by cache hits.
avg_latency_ms number Average end-to-end latency in milliseconds.
errors integer Requests that returned an HTTP status >= 400, excluding 499 client-aborted.
client_aborted integer Requests the client aborted before response headers (HTTP 499). A stream cancelled after headers (Stop pressed mid-answer) is a status 200 row with aborted 1 in the request log and is not counted here.

by_user is empty by default (anonymized) and is populated only when the caller explicitly opts out with anonymize=0. When populated, it only includes requests made with auth tokens that have a user_id set; anonymous token requests are not included. Up to 50 users are returned, ordered by cost descending. The example below shows the opted-out (identifying) response — the default response has "by_user": [] and "anonymized": true.

CacheEfficiency fields

cache_efficiency summarises input-token caching and input cost over the window (successful requests across all providers). Providers with prompt caching (Anthropic) populate the cache-read/write terms; providers without it (e.g. local models) contribute only their standard input cost at a 0 % hit rate.

Field Type Description
cache_write_tokens integer Prompt-cache write tokens — the total across both TTL tiers (the 1-hour part is counted once; it is a share of the total, not an addition to it).
cache_read_tokens integer Prompt-cache read tokens.
standard_input_tokens integer Non-cached input (prompt) tokens.
uncached_cost_usd number Cost that would have applied without the prompt cache.
cached_cost_usd number Cost actually incurred reading from the cache.
cache_hit_pct number | null Share of input tokens served from cache, as a percentage. null when there is no data.
compaction_tokens_saved integer Input tokens saved by conversation compaction.
compaction_cost_saved number Cost saved by conversation compaction in USD.

Example response

{
  "percentiles": { "p50": 420, "p95": 1840, "p99": 3210 },
  "top_models": [
    { "model": "gpt-4o", "provider": "openai", "requests": 842,
      "input_tokens": 1120400, "output_tokens": 208900, "cost_usd": 3.14, "avg_latency_ms": 680 }
  ],
  "by_tenant": [
    {
      "tenant_id": "ten_abc123", "tenant": "myapp",
      "requests": 1423, "blocked": 14, "cached": 82,
      "input_tokens": 1840200, "output_tokens": 312400,
      "cost_usd": 9.84, "saved_cost_usd": 0.41,
      "avg_latency_ms": 820, "errors": 3, "client_aborted": 1
    }
  ],
  "by_gateway": [
    {
      "gateway_id": "gw_xyz789", "gateway": "prod", "tenant": "myapp",
      "requests": 1423, "blocked": 14, "cached": 82,
      "input_tokens": 1840200, "output_tokens": 312400,
      "cost_usd": 9.84, "saved_cost_usd": 0.41,
      "avg_latency_ms": 820, "errors": 3, "client_aborted": 1
    }
  ],
  "by_channel": [
    { "channel": "ui",    "requests": 900, "input_tokens": 1200000, "output_tokens": 210000,
      "cost_usd": 6.10, "errors": 2, "avg_latency_ms": 610 },
    { "channel": "api",   "requests": 400, "input_tokens": 560000,  "output_tokens": 92000,
      "cost_usd": 3.20, "errors": 1, "avg_latency_ms": 700 },
    { "channel": "agent", "requests": 123, "input_tokens": 80200,   "output_tokens": 10400,
      "cost_usd": 0.54, "errors": 0, "avg_latency_ms": 540 }
  ],
  "by_agent": [
    { "agent_id": "agt_support", "agent": "Support Bot", "invocations": 90,
      "conversations": 12, "distinct_users": 4, "cost_usd": 0.54 }
  ],
  "by_agent_series": [
    { "agent_id": "agt_support", "agent": "Support Bot", "series": [
      { "ts": 1742457600000, "invocations": 12 },
      { "ts": 1742544000000, "invocations": 0 },
      { "ts": 1742630400000, "invocations": 31 }
    ] }
  ],
  "by_user": [
    {
      "user_id": "usr_alice", "tenant_id": "ten_abc123", "email": "alice@myapp.example",
      "requests": 341, "blocked": 2, "cached": 18,
      "input_tokens": 440200, "output_tokens": 74800,
      "cost_usd": 2.31, "saved_cost_usd": 0.09,
      "avg_latency_ms": 790, "errors": 1, "client_aborted": 0
    }
  ],
  "anonymized": false,
  "cache_efficiency": {
    "cache_write_tokens": 12800, "cache_read_tokens": 90400,
    "standard_input_tokens": 1737000,
    "uncached_cost_usd": 27.40, "cached_cost_usd": 0.27,
    "cache_hit_pct": 4.9,
    "compaction_tokens_saved": 0, "compaction_cost_saved": 0
  }
}

GET /stats/timeseries

Returns an array of time-bucketed data points, suitable for rendering sparklines or charts. Buckets with no activity are included as zero-filled entries so the array length is always exactly n.

Scope follows GET /stats: non-admin callers are pinned to their own tenant and own data server-side; a caller with no tenant is refused 403. A platform admin may pass ?tenant_id= to target one tenant, or omit it (empty value included) for the global view. The global series excludes synthetic fixture tenants (e.g. the e2e-synthetic monitoring prober); a tenant-scoped series includes them.

curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=1h&n=24"

Query parameters

Parameter Type Default Description
bucket string 1h Bucket size. One of: 5m, 15m, 30m, 1h, 6h, 1d.
n integer 24 Number of buckets to return. Range: 1–168.
until integer now End of the time range as a Unix timestamp in seconds. Defaults to the current time.
tz_offset integer 0 (UTC) The viewer's UTC offset in minutes, in the JavaScript Date.getTimezoneOffset() convention: UTC − local (Central European Summer Time = -120, US Eastern Standard Time = 300). Buckets are floored in the viewer's local frame, so the 1d bucket starts at local midnight and a fractional-hour offset (e.g. UTC+5:30) shifts the hourly buckets to local boundaries. Accepted: an integer within [-840, 720] (UTC−14 … UTC+14). Rejected → falls back to UTC (offset 0): any non-numeric, empty, repeated, out-of-range, or missing value (fractional values are floored toward −∞). It never causes an error and never influences authorization — it shifts the bucket boundaries only.

TimeseriesPoint fields

Field Type Description
ts integer Bucket start time as Unix milliseconds.
requests integer Total requests in the bucket.
blocked integer Requests blocked in the bucket.
rate_limited integer Requests rate-limited in the bucket.
errors integer Requests that returned an HTTP status >= 400 (excluding 499 client-aborted) in the bucket.
client_aborted integer Requests the client aborted before response headers (HTTP 499) in the bucket. Streams cancelled after headers are status 200 rows with aborted 1 in the request log, not counted here.
cost_usd number Total cost in USD for the bucket.

💡 Note: ts is in Unix milliseconds (not seconds) for direct compatibility with charting libraries that expect millisecond-precision timestamps (e.g. Chart.js, Recharts, Grafana).

Example response

[
  {"ts": 1742540400000, "requests": 42, "blocked": 1, "rate_limited": 0, "errors": 1, "client_aborted": 0, "cost_usd": 0.0314},
  {"ts": 1742544000000, "requests": 67, "blocked": 0, "rate_limited": 0, "errors": 0, "client_aborted": 0, "cost_usd": 0.0521},
  {"ts": 1742547600000, "requests": 0,  "blocked": 0, "rate_limited": 0, "errors": 0, "client_aborted": 0, "cost_usd": 0.0},
  {"ts": 1742551200000, "requests": 18, "blocked": 2, "rate_limited": 1, "errors": 2, "client_aborted": 1, "cost_usd": 0.0119}
]

GET /health/dashboard

GET /admin/v1/health/dashboard

Returns the aggregated operational-health summary that the Health dashboard renders in a single call — the synthetic-probe results, the last-24-hours non-2xx and client-aborted request counts, the phoned-home client errors, and the other failure surfaces shown there.

Requires the HEALTH_VIEW permission; the Health dashboard itself is available to the admin role only. A backend failure returns 500.

Platform web-search key (AGF-2874). The summary carries web_search_platform_key_missing (boolean): true when the platform Linkup key (the trial_linkup_api_key setting — applied to every keyless Linkup gateway on every plan) is unset while at least one production gateway of a non-deleted tenant carries a keyless Linkup web_search block — i.e. web search is silently dark for it. false when the key is set, or when nothing depends on it yet. The companion object web_search_platform_key gives the inputs: { "key_set": boolean, "keyless_production_gateways": integer }. If the count could not be read, all three are absent and web_search_platform_key_error (string) names the fault — a failed read is never reported as healthy; the dashboard renders "unknown". The same condition drives the [platform_alert] signal web_search_platform_key_missing (checked every 10 minutes, ERR-logged each check, notified at most every 6 hours).

Plan-aware model copy (AGF-2930). The summary carries plan_model_copy_stale (boolean): true when, for at least one offered model set served in the last 7 days, the set's current copy generation was rejected by the validator (or the set's catalog facts could not be read as valid input) — the model picker then shows that set's hand-written seed line, a checked catalog tagline or the model name instead (see Tagline resolution). plan_model_copy_indistinct (boolean) is true when a current set offers two models no fact tells apart (same price rank, context, capabilities, hosting and speed) — a plan-design problem, not a generator fault. The companion object plan_model_copy gives the counts: { "stale_sets": integer, "indistinct_sets": integer }. If the status could not be read, all three are absent and plan_model_copy_error (string) names the fault — the dashboard renders "unknown", never healthy.


Cost-of-goods (COGS)

The COGS endpoints aggregate the per-leg billing ledger (request_log_legs) into cost breakdowns by model, by tenant, by service, or by API key. A single inference request may produce several legs (the primary call, any fallback, plus side services such as guardrails or web search), so these views give a finer-grained cost picture than the request-level Stats summary.

All four endpoints share the same time-window selection:

Query parameters (all COGS endpoints)

Parameter Type Default Description
range string 30d Relative window ending now. Format <n><unit> where unit is d (days), h (hours), or w (weeks) — e.g. 24h, 2w, 90d. Ignored when from/to are given.
from integer — Explicit window start as a Unix timestamp in seconds. Must be paired with to.
to integer — Explicit window end as a Unix timestamp in seconds. Must be paired with from.

💡 Note: When neither from/to nor range is given, the window defaults to the last 30 days plus one day into the future (the extra day captures in-flight legs still being written).

CSV export (format) — cost-allocation endpoints

The three cost-allocation endpoints — /cogs/per-token, /cogs/per-user, and /cogs/per-agent — accept a format parameter for a machine-readable download:

Parameter Type Default Description
format string json Response encoding. json (or absent) returns the JSON array documented per endpoint. csv returns the same rows as text/csv; charset=utf-8 with a fixed header row (the response-field order below), served as a file attachment (Content-Disposition: attachment). Any other value — an unknown string, a repeated ?format= (which arrives as a list), or a bare ?format (no value) — is rejected with 400 { "error": "invalid format (json or csv)" } (fail-closed allowlist), validated before the aggregate query runs.

Notes on the CSV encoding:

  • The column order is fixed and equals the endpoint's response-field order; a missing value renders as an empty cell.
  • Cells are RFC-4180 quoted, and any value that begins with =, +, -, @, TAB, or CR is prefixed with a single quote to neutralize spreadsheet formula injection (CWE-1236) — so a user email or agent name such as =SUM(...) cannot execute when the file is opened in a spreadsheet.
  • No byte-order mark (BOM) is emitted.
  • The download filename is a fixed per-endpoint constant (cogs-per-token.csv / cogs-per-user.csv / cogs-per-agent.csv); the server sets Access-Control-Expose-Headers: Content-Disposition so a browser SPA fetching cross-origin can read it.
  • The bucket, tenant_id, scope, and billable_only rules are identical to the JSON response; format only changes the encoding.

GET /cogs/per-model

Cost and token totals grouped by (provider, model, tier, byok), ordered by cost descending. Required role: admin or tenant_admin.

curl "https://<your-gateway-host>/admin/v1/cogs/per-model?range=7d"

Additional query parameters

Parameter Type Default Description
tenant_id string — Restrict to one tenant. Honoured only for callers with the admin role; every other caller (including tenant_admin) is scoped to their own tenant server-side, and a caller with no tenant is refused 403.
billable_only string 1 Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier or internal).

Response

Array of rows. Returns 500 { "error": ... } on a storage failure.

Field Type Description
provider string Provider that served the leg.
model string Model used for the leg.
tier string | null Pricing tier recorded for the leg.
byok integer 1 if the leg used a bring-your-own-key credential, 0 otherwise.
call_count integer Number of legs in this group.
in_t integer Sum of input (prompt) tokens.
out_t integer Sum of output (completion) tokens.
cache_read_t integer Sum of prompt-cache read tokens.
cache_write_t integer Sum of prompt-cache creation (write) tokens.
cost_usd number Total cost in USD for the group.
partial_count integer Number of legs flagged partial (interrupted / incomplete).
[
  {
    "provider": "anthropic",
    "model": "claude-opus-4-6",
    "tier": null,
    "byok": 0,
    "call_count": 412,
    "in_t": 1840200,
    "out_t": 312400,
    "cache_read_t": 90400,
    "cache_write_t": 12800,
    "cost_usd": 9.84,
    "partial_count": 2
  }
]

GET /cogs/per-tenant

Cost and token totals grouped by tenant, billable legs only, ordered by cost descending. Required role: platform admin only — lower-privileged callers receive 403 { "error": "forbidden: platform admin only" }.

curl "https://<your-gateway-host>/admin/v1/cogs/per-tenant?range=30d"

Response

Array of rows. Returns 500 { "error": ... } on a storage failure.

Field Type Description
tenant_id string Tenant UUID.
call_count integer Number of billable legs for the tenant.
cost_usd number Total cost in USD.
in_t integer Sum of input tokens.
out_t integer Sum of output tokens.
[
  { "tenant_id": "ten_abc123", "call_count": 1423, "cost_usd": 14.30, "in_t": 2750000, "out_t": 480000 }
]

GET /cogs/per-service

Per-service usage grouped by (provider, service, unit_kind), ordered by call count descending. Counts only legs that carry a non-null service (e.g. web_search, guardrail). Required role: admin or tenant_admin.

curl "https://<your-gateway-host>/admin/v1/cogs/per-service?range=30d"

Additional query parameters

Parameter Type Default Description
tenant_id string — Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's — or the global, all-tenant — per-service totals.

Plus the shared range / from / to window parameters above.

Response

Array of rows. Returns 500 { "error": ... } on a storage failure.

Field Type Description
provider string Provider that served the service leg.
service string Service name (e.g. web_search, guardrail).
unit_kind string | null Unit the service is metered in (e.g. requests, tokens).
call_count integer Number of legs in this group.
units_total number Sum of metered units.
avg_latency_ms number Average leg latency in milliseconds.
in_t integer Sum of input tokens (where applicable).
out_t integer Sum of output tokens (where applicable).
[
  { "provider": "anthropic", "service": "web_search", "unit_kind": "requests", "call_count": 88, "units_total": 88, "avg_latency_ms": 640, "in_t": 0, "out_t": 0 }
]

GET /cogs/per-token

Cost and usage aggregated per API key (token) × period (Zeitraum) × KI-Modell, ordered by cost descending. This is the billing-disclosure view: it answers "what did each API key spend, on which model, in which period." Required role: admin or tenant_admin.

curl "https://<your-gateway-host>/admin/v1/cogs/per-token?range=30d&bucket=month"

Only genuine LLM legs are counted: legs with a NULL model (non-LLM side services such as PII scrubbing or guardrails) are excluded, so the per-model figures are clean. Customer-paid (BYOK) legs are kept on their own rows via the byok dimension and are never summed into platform cost. A single user request that fans out into several same-model legs (e.g. a tool loop) counts as one request.

Additional query parameters

Parameter Type Default Description
bucket string month Period granularity for the period column. One of day, week, month, total. Any other value is rejected with 400 { "error": "invalid bucket" } (fail-closed allowlist — the value is never interpolated into SQL). total collapses the whole window into a single period labelled "total".
tenant_id string — Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's cost rows.
token_id string — Restrict to a single API key. Must be a single scalar value — a repeated param (?token_id=a&token_id=b) is rejected with 400 { "error": "token_id must be a single value" } (fail-closed; it is never bound into SQL). Combined with the mandatory tenant scope: a token_id belonging to another tenant returns an empty result.
billable_only string 1 Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier / self-hosted). Non-billable legs still obey the model IS NOT NULL rule.

Plus the shared range / from / to window parameters above.

Response

Array of rows. Returns 500 { "error": ... } on a storage failure, 400 { "error": "invalid bucket" } on a bad bucket.

Field Type Description
token_id string API-key id the legs were attributed to. Empty string ("") means non-API-key traffic — interactive UI / playground sessions that carry no API key.
token_label string Human label of the API key (from auth_token.label). Empty string if the key has no label or was deleted.
period string Period bucket. Format depends on bucket: 2026-07-15 (day), 2026-W29 (ISO year-week), 2026-07 (month), or total. Boundaries are computed in the gateway database session timezone (UTC on the standard deployment). Note: the week bucket uses the ISO week-numbering year, so late-December usage can fall in the following calendar year's W01 (e.g. 2025-12-31 → 2026-W01) — while day/month use the calendar year.
provider string Provider that served the model.
model string Model id (never null — non-model side-service legs are excluded).
byok integer 1 if the legs used a bring-your-own-key credential (customer-paid), 0 for platform-paid. Kept as a separate dimension so customer and platform cost never merge.
requests integer Distinct end-user requests in this group (COUNT(DISTINCT request_log_id)), not the leg count. This is a per-row count and is not additive across rows: a single request that fans out across models (a fallback) or across a period boundary is counted once in each row it touches, so summing the requests column overstates the true request total.
input_tokens integer Sum of input (prompt) tokens.
output_tokens integer Sum of output (completion) tokens.
total_tokens integer input_tokens + output_tokens.
cost_usd number Total cost in USD for the group, rounded to 6 decimals.
[
  {
    "token_id": "8f3c2a10-6b4e-4d21-9a77-1e5c0b9d4a02",
    "token_label": "prod-key",
    "period": "2026-07",
    "provider": "anthropic",
    "model": "claude-opus-4-7",
    "byok": 0,
    "requests": 412,
    "input_tokens": 1840200,
    "output_tokens": 312400,
    "total_tokens": 2152600,
    "cost_usd": 9.84
  }
]

GET /cogs/per-user

Cost and usage aggregated per user × period (Zeitraum) × KI-Modell, ordered by cost descending. This is the per-user cost-allocation view: it answers "which user spent how much, on which model, in which period." Required role: admin, tenant_admin, or finance (the COGS_PER_USER_VIEW permission). It is the exact cost twin of /cogs/per-token — same period buckets, same tenant scoping, same row shape — with the grouping key swapped from the API key to the user.

curl "https://<your-gateway-host>/admin/v1/cogs/per-user?range=30d&bucket=month"

Only genuine LLM legs are counted: legs with a NULL model (non-LLM side services such as PII scrubbing or guardrails) are excluded, so the per-model figures are clean. Customer-paid (BYOK) legs are kept on their own rows via the byok dimension and are never summed into platform cost. A single user request that fans out into several same-model legs (e.g. a tool loop) counts as one request.

Cost allocation is inherently per-identity: this endpoint returns each user's email within the caller's scope. A tenant_admin can only ever see the users of their own tenant (server-side tenant scope, fail-closed); only an admin may scope to another tenant or read across all tenants.

Additional query parameters

Parameter Type Default Description
bucket string month Period granularity for the period column. One of day, week, month, total. Any other value is rejected with 400 { "error": "invalid bucket" } (fail-closed allowlist — the value is never interpolated into SQL). total collapses the whole window into a single period labelled "total".
tenant_id string — Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's cost rows.
user_id string — Restrict to a single user. Must be a single scalar value — a repeated param (?user_id=a&user_id=b) is rejected with 400 { "error": "user_id must be a single value" } (fail-closed; never bound into SQL). Combined with the mandatory tenant scope: a user_id belonging to another tenant returns an empty result.
billable_only string 1 Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier / self-hosted). Non-billable legs still obey the model IS NOT NULL rule.

Plus the shared range / from / to window parameters above.

Response

Array of rows. Returns 500 { "error": ... } on a storage failure, 400 { "error": "invalid bucket" } on a bad bucket.

Field Type Description
user_id string User id the legs were attributed to. Empty string ("") means traffic with no attributed user — e.g. API-key requests that carry no user context.
user_email string Email of the user (from user.email). Empty string if the user has no email or was deleted.
period string Period bucket. Format depends on bucket: 2026-07-15 (day), 2026-W29 (ISO year-week), 2026-07 (month), or total. Boundaries are computed in the gateway database session timezone (UTC on the standard deployment). Note: the week bucket uses the ISO week-numbering year, so late-December usage can fall in the following calendar year's W01 (e.g. 2025-12-31 → 2026-W01) — while day/month use the calendar year.
provider string Provider that served the model.
model string Model id (never null — non-model side-service legs are excluded).
byok integer 1 if the legs used a bring-your-own-key credential (customer-paid), 0 for platform-paid. Kept as a separate dimension so customer and platform cost never merge.
requests integer Distinct end-user requests in this group (COUNT(DISTINCT request_log_id)), not the leg count. This is a per-row count and is not additive across rows: a single request that fans out across models (a fallback) or across a period boundary is counted once in each row it touches, so summing the requests column overstates the true request total.
input_tokens integer Sum of input (prompt) tokens.
output_tokens integer Sum of output (completion) tokens.
total_tokens integer input_tokens + output_tokens.
cost_usd number Total cost in USD for the group, rounded to 6 decimals.
[
  {
    "user_id": "3a1f8c22-9d4e-4b70-8c11-2f6a0e7b5d13",
    "user_email": "alex@example.com",
    "period": "2026-07",
    "provider": "anthropic",
    "model": "claude-opus-4-7",
    "byok": 0,
    "requests": 128,
    "input_tokens": 540200,
    "output_tokens": 96400,
    "total_tokens": 636600,
    "cost_usd": 3.11
  }
]

GET /cogs/per-agent

Cost and usage aggregated per agent × period (Zeitraum) × KI-Modell, ordered by cost descending. This is the per-agent cost-allocation view: it answers "which agent spent how much, on which model, in which period." Required role: admin or tenant_admin. It is the agent-dimension twin of /cogs/per-user — same period buckets, same tenant scoping, same row shape — with the grouping key swapped from the user to the agent.

curl "https://<your-gateway-host>/admin/v1/cogs/per-agent?range=30d&bucket=month"

Only agent-invoked traffic is included (requests that carry no agent are excluded). Agents are grouped by their clone-lineage root, so a shared agent and all of its clones roll up under one row group — matching the per-agent stats in /stats/analytics. Only genuine LLM legs are counted: legs with a NULL model (non-LLM side services such as PII scrubbing or guardrails) are excluded, so the per-model figures are clean. Customer-paid (BYOK) legs are kept on their own rows via the byok dimension and are never summed into platform cost. A single agent request that fans out into several same-model legs (e.g. a tool loop) counts as one request.

Cost and token figures are tenant-bounded (server-side tenant scope, fail-closed): a tenant_admin can only ever see their own tenant's agent spend; only an admin may scope to another tenant or read across all tenants.

Additional query parameters

Parameter Type Default Description
bucket string month Period granularity for the period column. One of day, week, month, total. Any other value is rejected with 400 { "error": "invalid bucket" } (fail-closed allowlist — the value is never interpolated into SQL). total collapses the whole window into a single period labelled "total".
tenant_id string — Restrict to one tenant. Honoured only for callers with the admin role. Every other caller (including tenant_admin) is scoped to their own tenant server-side; a spoofed tenant_id is ignored, and a caller with no tenant is refused 403. A caller can never read another tenant's cost rows.
agent_id string — Restrict to a single agent, matched against the clone-lineage root id. Must be a single scalar value — a repeated param (?agent_id=a&agent_id=b) is rejected with 400 { "error": "agent_id must be a single value" } (fail-closed; never bound into SQL). Combined with the mandatory tenant scope: an agent_id whose traffic belongs to another tenant returns an empty result.
billable_only string 1 Include only billable legs. Set to 0 to also include non-billable legs (e.g. free-tier / self-hosted). Non-billable legs still obey the model IS NOT NULL rule.

Plus the shared range / from / to window parameters above.

Response

Array of rows. Returns 500 { "error": ... } on a storage failure, 400 { "error": "invalid bucket" } on a bad bucket.

Field Type Description
agent_id string Clone-lineage root id the legs were attributed to (the stamped root, falling back to the agent's own id for pre-lineage rows).
agent string Display name of the root agent (agent.name, falling back to its slug, then the raw root id if the agent was deleted). For a cross-tenant cloned agent the label may resolve to the source agent's name; the cost and token figures always count only the caller's own tenant's legs.
period string Period bucket. Format depends on bucket: 2026-07-15 (day), 2026-W29 (ISO year-week), 2026-07 (month), or total. Boundaries are computed in the gateway database session timezone (UTC on the standard deployment). Note: the week bucket uses the ISO week-numbering year, so late-December usage can fall in the following calendar year's W01 (e.g. 2025-12-31 → 2026-W01) — while day/month use the calendar year.
provider string Provider that served the model.
model string Model id (never null — non-model side-service legs are excluded).
byok integer 1 if the legs used a bring-your-own-key credential (customer-paid), 0 for platform-paid. Kept as a separate dimension so customer and platform cost never merge.
requests integer Distinct end-user requests in this group (COUNT(DISTINCT request_log_id)), not the leg count. This is a per-row count and is not additive across rows: a single request that fans out across models (a fallback) or across a period boundary is counted once in each row it touches, so summing the requests column overstates the true request total.
input_tokens integer Sum of input (prompt) tokens.
output_tokens integer Sum of output (completion) tokens.
total_tokens integer input_tokens + output_tokens.
cost_usd number Total cost in USD for the group, rounded to 6 decimals.
[
  {
    "agent_id": "ddc230d8-78ef-11f1-a3ae-6c92cf1260f0",
    "agent": "Support Triage",
    "period": "2026-07",
    "provider": "anthropic",
    "model": "claude-opus-4-7",
    "byok": 0,
    "requests": 64,
    "input_tokens": 210400,
    "output_tokens": 38800,
    "total_tokens": 249200,
    "cost_usd": 1.42
  }
]

Example curl requests

⭐ Example: The following examples show common query patterns for the timeseries endpoint.

Today's requests by hour (last 24 hours)

curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=1h&n=24"

Yesterday's requests by hour

# Set 'until' to the end of yesterday (start of today in Unix seconds)
YESTERDAY_END=$(date -d "today 00:00:00" +%s)
curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=1h&n=24&until=${YESTERDAY_END}"

Last 7 days by day

curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=1d&n=7"

Last hour in 5-minute buckets

curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=5m&n=12"

Last 30 days in 6-hour buckets

curl "https://<your-gateway-host>/admin/v1/stats/timeseries?bucket=6h&n=120"

GET /stats/reactivation

The platform-wide roll-up of the member-reactivation e-mails, shown as the Member reactivation card on the platform dashboard.

curl https://<your-gateway-host>/admin/v1/stats/reactivation

Authorization: platform admin only (cross-tenant data; 403 otherwise).

Response: mails_sent_7d, mails_sent_30d, reactivated_7d (members who became active within 7 days of a mail sent in the last 30 days), opt_outs_total, opt_outs_30d, and top_tenants — the ten most-mailed workspaces of the last 30 days as [{ tenant_id, slug, mails_sent_30d }] (an empty array when nothing was sent). Counts come from the send log only — no per-member classification runs fleet-wide; the per-tenant states live on GET /tenants/{id}/reactivation-metrics. A database fault is 503.


See also