Skip to content

Guardrails

💡 Note: This page defines the guardrail concept — the two tiers, the verdict lifecycle, and targets. For the guardrails overview and the comparison table that helps you pick one, see Data protection and compliance. To create, order, and manage guardrails, see Building a guardrail pipeline.

A guardrail is a check applied to a request or a response that produces a verdict. Guardrails enforce data-protection, output-policy, and prompt-injection rules before a request reaches the upstream provider or before a response reaches the client.

Tiers

The gateway runs guardrails in two tiers, based on where each guardrail executes and how fast it runs.

Tier Guardrail types Execution Latency
Tier 1 regex, keyword, jailbreak, json_schema, contains_code, gibberish, language, custom_pii In-process within the gateway Sub-millisecond
Tier 2 presidio, prompt_guard, prompt_injection, pii_protector Sidecar HTTP call within the Myra perimeter Milliseconds

Tier 1 always runs before Tier 2. Within the same tier, guardrails run in the order they appear in the gateway configuration. A Tier 2 guardrail is always dispatched when it is configured; if its sidecar is not deployed or is unreachable, the guardrail returns an error verdict, which is then resolved by its fail_open setting (see Verdict lifecycle) rather than being silently skipped.

Verdict lifecycle

Each guardrail returns one verdict. The verdict controls how the gateway handles the traffic.

Verdict Effect
pass The content does not match. The pipeline continues unchanged.
scrubbed The content matches and the guardrail rewrites it before the pipeline continues. An external provider never sees the original value; downstream guardrails see the rewritten payload. (Scrub is skipped on a wholly-local Myra/EU model leg — see PII Protector — Local model legs are not masked.)
flagged The content matches, but the request is allowed to continue. The match is recorded for observability without modifying the traffic. A flag on a top-level turn also queues a moderation review item on the Tool reviews inbox.
block The content matches and the request is rejected. The gateway returns a synthetic HTTP 200 response whose assistant message states the request was blocked (so OpenAI-compatible clients render it inline). The pipeline aborts on the first block verdict.
error The guardrail failed to run. Behaviour depends on the fail_open setting: when fail_open is true, the verdict is treated as pass; when false, it is treated as block.

A block verdict from any single guardrail stops the entire pipeline. The scrubbed and flagged verdicts are non-terminal — the remaining guardrails still run.

💡 Note: The verdict is the outcome the guardrail returns (scrubbed, flagged). The action field in the guardrail configuration uses the imperative form (scrub, flag, block) and selects which outcome a match produces. The verdict and the action are related but not the same field.

Targets

Each guardrail declares which traffic direction it inspects.

Target Description
request Inspect the outbound request body (default).
response Inspect the inbound model response body.
both Inspect both the request and the response.

Logging

Every verdict is recorded in the request log. Guardrails that fired (returned block, scrub, or flag) are listed under detectors_fired. The blocking guardrail and the matched pattern or category are recorded under blocked_by and block_reason.

See also