Skip to content

Request tracing

Gateway tracing records a step-by-step execution trace for each inference request. Each trace captures what happened at every stage of the request pipeline — model resolution, routing decisions, guardrail results, upstream calls, and the final response delivered to the client. Traces are linked to log entries via the trace_id field.

Tracing is opt-in and has no effect on request latency. All writes are fire-and-forget.


Enabling tracing

Proceed as follows to enable tracing on a gateway:

  1. Open the gateway in the admin UI.
  2. The gateway detail page appears.
  3. Click on the Edit button in the Gateway card header.
  4. The gateway configuration panel opens.
  5. Scroll to the Tracing section.
  6. Toggle the Enable request tracing control to on.
  7. If required, toggle Include message bodies in trace to record raw request content and web-search previews in the trace steps.
  8. Enable this for debugging only. It stores prompt text, response previews, and web-search queries in the trace table.
  9. When it is off (the default), these fields are omitted and the request and web-search steps record only metadata (message counts, body sizes, status).
  10. If required, set the Retention (hours) field. Note: this per-gateway value is stored but does not change how long traces are kept — trace retention is a single deployment-wide setting (default 48 hours); see Trace retention below.
  11. Click on the Save Changes button.
  12. The updated configuration is applied within seconds.

Screenshot: the Tracing section of the Edit Gateway dialog The Tracing section of the Edit Gateway dialog

-> Tracing is active for all subsequent requests through this gateway.

Alternatively, apply the configuration via the API:

curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
  -H "Content-Type: application/json" \
  -d '{"config": {"tracing": {"enabled": true}}}'

The full tracing configuration block:

{
  "tracing": {
    "enabled": true,
    "include_bodies": false
  }
}
Field Type Default Description
enabled boolean false Activates request tracing for this gateway.
include_bodies boolean false Includes raw content in these trace-step fields: the request message array (request_received, leg1_request, leg2_request), leg-1 response previews (leg1_response, leg1_direct_answer), web-search queries (search_result), the fetched URLs and page previews (fetch_attempt, fetch_result), the streamed model answer (leg2_response content/raw_content), tool-call arguments and result filenames (tool_call args, tool_result filename), and the raw model text and parsed filename captured by the dropped-tool guards (silent_write_file_drop raw_content/dropped_file, silent_tool_call_drop raw_content). Enable only for debugging — this stores prompt and response text in the trace table. When off, these fields are omitted and those steps record only metadata (counts, sizes, status). The Playground and agent-invoke trace paths never set this flag, so those fields stay metadata-only there. A private (ghost) conversation is always metadata-only regardless of this flag: its trace row is still written for operations and debugging (timestamps, model id, latency, token counts, status, tool/round metadata), but no prompt, response, tool argument, query, or preview content is persisted in it.

Trace retention

Stored gateway traces are purged automatically, so the trace table does not grow without bound. Retention is deployment-wide, not per-gateway: a background job on the gateway deletes gateway traces older than the retention window once per hour.

The window is a deployment setting configured by your operator (default 48 hours, minimum 1). It is not part of a gateway's configuration.


Pipeline step reference

Steps are recorded in sequence order. Only steps that are reached for a given request appear in the trace. For example, a cached response does not produce upstream_request or upstream_response steps.

💡 Note: The steps below are the common ones a reader meets most often; the list is not exhaustive. The pipeline emits roughly thirty step types in total (retries, guardrail phases, tool-loop rounds, delegation, and more), and new steps are added as the pipeline grows.

request_received

Recorded immediately after the request body is parsed.

Field Description
model Model name as received from the client (before normalisation).
provider Provider resolved from the URL path.
messages_count Number of messages in the request body.
streaming Whether the client requested a streaming response.
size_bytes Raw request body size in bytes.
is_compat Whether the request arrived on the OpenAI-compatible endpoint.
messages Full message array (only present when include_bodies: true).

request_transformed

Recorded only when the model or provider changes during normalisation (e.g. compat endpoint provider inference, provider-prefix stripping).

Field Description
model_before Model name before normalisation.
model_after Model name after normalisation.
provider_before Provider before normalisation.
provider_after Provider after normalisation.

routing_applied

Recorded when a routing rule matches and changes the provider or model.

Field Description
rule_id ID of the matched routing rule.
provider_before Provider before routing.
model_before Model before routing.
provider_after Provider after routing.
model_after Model after routing.

guardrail_result

Recorded after the guardrail pipeline completes.

Field Description
verdict "safe", "unsafe", or the guardrail verdict string.
blocked true if the request was blocked.
detector Name of the detector that fired (if blocked).
latency_ms Time spent in the guardrail pipeline.

upstream_request

Recorded immediately before each upstream provider call.

Field Description
attempt Attempt number (1 = first try).
provider Provider being called.
model Model being called.
url Request URL sent to the provider (auth key stripped).

upstream_response

Recorded after each upstream provider response is received.

Field Description
attempt Attempt number.
provider Provider that responded.
status HTTP status code.
latency_ms Time between upstream request and response.
input_tokens Prompt tokens from the provider usage object.
output_tokens Completion tokens from the provider usage object.

upstream_error

Recorded when an upstream call fails with a network error (not an HTTP error response).

Field Description
attempt Attempt number.
provider Provider that failed.
error Error message string.

response_delivered

Recorded after the response is sent to the client.

Field Description
streaming Whether the response was streamed.
provider_status HTTP status code from the upstream provider.
body_size Response body size in bytes (non-streaming).
compat Whether the response was re-encoded for the compat endpoint.

Web search steps

These steps appear when web search is enabled on the gateway.

Step Description
leg1_request First inference call to extract search queries. Records messages_count/tools_count; the raw messages/tools only when include_bodies is on.
leg1_response Response from the first call. Records status/body_len; the raw body_preview only when include_bodies is on.
leg1_direct_answer Recorded when the model answers directly without triggering a search. Records body_len; the raw body_preview only when include_bodies is on.
search_result Queries sent and snippets returned from the search provider. Records queries_count and per-result sizes/URLs; the raw queries only when include_bodies is on.
fetch_attempt URLs selected for full-page fetching. Records url_count; the raw urls only when include_bodies is on.
fetch_result Fetch outcome for each URL. Records ok/text_len; the raw url and page preview only when include_bodies is on.
leg2_request Second inference call with search context injected. Records messages_count; the raw messages only when include_bodies is on.
leg2_response Final response from the second call. Records content_len (and raw_len when the relay changed the text); the raw content/raw_content only when include_bodies is on.

Tool execution steps

These steps appear when the model calls a tool (for example the code interpreter) during an agentic turn.

Step Description
tool_call One tool invocation the model requested. Records round/tool; the raw args only when include_bodies is on.
input_files_staged Code-interpreter data-file staging outcome. Records requested/staged/notes counts only — filenames and bytes are never stored in the trace. notes counts drop notes, not files: one note can cover several requested files, so requested is not necessarily staged plus notes.
tool_result The result returned to the model. Records round/tool/chars; the produced filename only when include_bodies is on.
loop_terminated The agentic loop ended early. Records the termination reason and its details.

RAG ranking step

This step appears when the model calls the knowledge-search tool.

Step Description
rag_rank Retrieval ranking for one knowledge-search call, including the optional cross-encoder rerank stage. Records mode (dormant when no reranker is configured, reranked when the reranker reordered the results, or fallback when a configured reranker was unavailable and the fused order was kept), configured (whether a reranker model is set), fused_n (candidate passages after hybrid fusion), kept (passages served after any rerank trim), and sent (passages sent to the reranker). The ranked document identifiers (order) only when include_bodies is on.

Linking traces to log entries

When gateway tracing is enabled, each log entry contains a trace_id field. Use it to retrieve the full execution trace.

Proceed as follows to retrieve the trace for a log entry:

  1. Fetch the log entry to obtain the trace_id.
curl "https://<your-gateway-host>/admin/v1/logs/log_abc789"
  • The response contains a trace_id field, for example "trc_abc123".

  • Fetch the full trace using the trace_id.

curl "https://<your-gateway-host>/admin/v1/traces/trc_abc123"

-> The full execution trace for the request is returned.

💡 Note: trace_id is null in log entries when tracing was not active for that request.

Playground requests always produce traces regardless of the gateway tracing config. Playground traces have source: "playground" and are accessible via the same GET /traces/{id} endpoint.


Configuring OTLP export

The gateway exports distributed traces as OpenTelemetry Protocol (OTLP) spans to any OpenTelemetry-compatible backend — Jaeger, Grafana Tempo, Datadog, Honeycomb, and others. OTLP export is independent of the internal pipeline trace described above.

W3C traceparent propagation

The gateway reads an incoming traceparent header from the client (if present) and propagates the trace context to upstream LLM providers. This allows end-to-end trace correlation across your services, the gateway, and the provider. When the gateway initialises tracing for a request — whether from an incoming traceparent or by generating new IDs — it injects a traceparent header into the upstream provider request.

OTLP/HTTP span export

After each request, the gateway emits a root span (SERVER) and, when an upstream call was made, a child span (CLIENT) to your OTel collector. Delivery is fully asynchronous and never adds latency.

Configuration

Add otlp_endpoint to your existing tracing block:

{
  "tracing": {
    "enabled": true,
    "otlp_endpoint": "http://otel-collector:4318",
    "service_name": "ai-gateway",
    "headers": {},
    "sample_rate": 1.0,
    "include_bodies": false
  }
}
Field Type Default Description
enabled boolean false Activates internal pipeline tracing
otlp_endpoint string — Base URL of your OTel collector, e.g. http://otel-collector:4318. Setting this enables OTLP export.
service_name string ai-gateway service.name resource attribute on all emitted spans
headers object {} Extra HTTP headers to include in the OTLP request (e.g. auth tokens for managed collectors)
sample_rate number 1.0 Fraction of requests to export (0.0 = never, 1.0 = always)
include_bodies boolean false When true, adds aig.request_size_bytes to spans

💡 Note: enabled: true alone activates only the internal pipeline trace (playground / Traces API). Setting otlp_endpoint enables OTLP export and implies tracing is active for that gateway — export runs on otlp_endpoint independently of enabled, so a gateway can export spans with enabled: false.

💡 Note: The four OTLP-export fields (otlp_endpoint, service_name, headers, sample_rate) are configured via the Admin API and have no control in the gateway Edit dialog. The dialog exposes only enabled / include_bodies / retention, and preserves the OTLP fields untouched on save — editing an unrelated gateway setting never wipes your export configuration, and toggling "enable tracing" off in the dialog does not stop OTLP export (remove otlp_endpoint via the API to do that).

Span model

Each exported trace contains up to two spans:

Span Kind When emitted Name
Root span SERVER (2) Every request inference
Upstream span CLIENT (3) Only when an upstream LLM call was made upstream.<provider>

The parentSpanId of the upstream span is the spanId of the root span, forming a parent-child relationship.

Root span attributes (GenAI semantic conventions)

Attribute Type Description
gen_ai.system string LLM provider (openai, anthropic, …)
gen_ai.request.model string Model name
gen_ai.usage.input_tokens integer Prompt token count
gen_ai.usage.output_tokens integer Completion token count
gen_ai.request.cost_usd double Estimated inference cost
http.status_code integer HTTP status returned to the client
aig.tenant_id string Tenant identifier
aig.gateway_id string Gateway identifier
aig.cached boolean Whether the response came from cache
aig.blocked boolean Whether the request was blocked
aig.blocked_by string Block reason category (when blocked)
aig.upstream_attempts integer Provider dials attempted (emitted when > 0)

Upstream span attributes

Attribute Type Description
gen_ai.system string Provider
gen_ai.request.model string Model
http.status_code integer Provider HTTP status
aig.upstream_latency_ms integer Provider round-trip latency
aig.upstream_attempts integer Retry count
aig.fallback_provider string Fallback provider (when failover occurred)
aig.fallback_model string Fallback model (when failover occurred)

Sampling

Use sample_rate to reduce export volume for high-traffic gateways:

{ "tracing": { "otlp_endpoint": "http://otel:4318", "sample_rate": 0.1 } }

This exports roughly 10 % of requests. Sampling is applied at export time, after the request completes: when the random draw exceeds sample_rate, the export is skipped and no span is built at all — nothing is assembled and then discarded. There is no head-based sampling; every request processes normally.

Configuration examples

The following examples show how to configure OTLP export for common backends.

Jaeger (all-in-one)

{ "tracing": { "otlp_endpoint": "http://jaeger:4318" } }

Grafana Tempo

{
  "tracing": {
    "otlp_endpoint": "http://tempo:4318",
    "service_name": "ai-gateway-prod"
  }
}

Datadog OTLP endpoint

{
  "tracing": {
    "otlp_endpoint": "http://datadog-agent:4318",
    "headers": { "DD-API-KEY": "your-api-key" },
    "service_name": "ai-gateway"
  }
}

Honeycomb

{
  "tracing": {
    "otlp_endpoint": "https://api.honeycomb.io",
    "headers": { "x-honeycomb-team": "your-api-key" },
    "service_name": "ai-gateway"
  }
}

See also