Skip to content

Guardrails API

The guardrails API exposes the configured detector list, recent block events, aggregated statistics, and the per-provider circuit breaker state for a gateway. All endpoints require an authenticated session with access to the gateway — for most of them any user in the gateway's tenant qualifies; they are not restricted to administrators. On recent guardrail events the per-request meta protection markers are additionally withheld from callers without REQUEST_LOGS_VIEW; the events themselves are returned to any user with gateway access.


Listing configured detectors

GET /admin/v1/gateways/<ID>/detectors

Returns the guardrail detectors configured on the gateway, in the order they appear in the gateway configuration (no pipeline-order sort). Each entry has these fields:

Field Type Description
type string Detector type.
name string Detector name.
action string Action taken when the detector fires. For a prompt_injection detector: block (the default for an absent, null or unrecognised value — this type is the one whose default runs the other way from the content detectors) or flag; scrub is not supported and is treated as block. A flag prompt-injection detector is also always fail-open — see Prompt Injection. For a presidio detector: scrub masks (anonymizes) the matched PII, block rejects the request, and flag only records the match without masking. An absent or null action defaults to scrub (the detector masks) — both at runtime and in the "PII protected" indicators; set action: "flag" explicitly for observe-only. For a regex or keyword detector, scrub masks every matched span in place with the detector's scrub_placeholder (default [REDACTED]) and records the outcome as scrubbed; their default action is flag (NOT scrub), so masking applies only when action: "scrub" is set explicitly. A flag on a top-level turn also queues a moderation review_item (flagged_prompt / flagged_output) that a security reviewer allows/denies/annotates — the queued row carries only the interaction's trace reference and the detector name(s), never the matched content (see Agents → moderation queue). Note that a broad flag pattern generates one moderation item per matching interaction, so scope it deliberately.

Accepted shape, and what an absent or malformed value means. action is not validated at the write boundary for any type except pii_protector / custom_pii, which accept only an omitted, null or "scrub" value because they always tokenise. Every other type resolves it at request time and then dispatches on a literal. A value that is not a string is treated exactly like an absent one: an absent action, a JSON null, a JSON false, a number, a boolean and an object all take the type's own default. An unrecognised string is different — it is never performed, and takes the type's else branch (flag for the content detectors, block for prompt_injection). The one exception is presidio, whose own rule normalises only null / false / absent to scrub — so a presidio detector whose action is a number, boolean or object reports flag, not scrub. (prompt_injection is not an exception to this rule: an absent and a non-string action both mean block there. It differs only on unrecognised strings, which it also treats as block.)

Previously a JSON null did not take the default: it survived the fallback (it is truthy in the runtime's semantics), matched no literal, and landed on the else branch — so on a type whose default is block a stored null, 7, true or {} bought less enforcement than storing nothing, silently turning a detector that refuses into one that only logs. That is fixed: no malformed value is weaker than an omitted one. If you have a stored non-string action on a prompt_guard, json_schema, contains_code, gibberish or language detector, it now REFUSES where it previously only flagged — the gateway logs a [detector_action] warning naming the detector so the config can be corrected. GET /admin/v1/gateways/{id}/detectors reports the action the runtime will actually apply, never the raw stored value, so for every detector type the gateway recognises its action is one of block, scrub or flag. A detector whose type the gateway does not recognise is the one exception, because nothing here knows what such a detector would do: its stored action is echoed unchanged when it is a string — so this is also the only case where the field can hold something else, including "" — and when it holds no string at all the endpoint reports flag.

The defaults, by type: block for prompt_injection, prompt_guard, json_schema, contains_code, gibberish and language; scrub for presidio; flag for regex, keyword and jailbreak. Only presidio, regex and keyword read action and perform scrub — pii_protector and custom_pii always tokenise and never read the field at all, and on every remaining type a stored scrub is not a mask, it is that type's fallback.
enabled boolean Whether the detector is switched on. A detector with enabled: false does not run — no masking, no block, no flag. Any value other than a literal boolean false (absent, null, true, or a non-boolean) counts as enabled (fail-safe: the detector runs, so a malformed value never silently disables protection). Disabling a PII-masking detector lets that PII flow to the model UNMASKED — the previous behavior always masked regardless of enabled. In the admin console this is gated behind an explicit confirmation, and every deactivation (toggling enabled: false or removing the detector) is recorded in the audit log under the action guardrail_masker.deactivated (actor, gateway, detector name + type — never the detector's keywords/entities). Disabling a masker unmasks text only: inline base64 media is still blocked with pii_media_unmaskable, and a custom_pii keyword list still guards tool / web-search egress, while any PII-protective detector remains present on the gateway (these gateway-level fail-safes over-protect, they never leak). A detector whose only PII masker is disabled is also no longer selected as a PII-preferred route by Auto.
target string Stage the detector applies to — request (the model input), response (the model output), or both. An omitted target defaults to request.

Accepted shape, and what is rejected. A write that introduces a target which is not one of those three strings is rejected with 400 on POST /tenants/{id}/gateways, PATCH /gateways/{id} and the guardrail-config import — including an explicit JSON null. Refusing a new null keeps a single canonical spelling of the default (omit the field); it is a canonical-spelling rule, not a safety one. Read-time semantics: a JSON null target resolves to request. A null, a false and an omitted target all fall back to request, so a detector carrying any of them runs on the request phase and is reported active: true with target: "request" (the report emits the resolved value, never raw null). This matches how a per-role (role_policies) detector has always read a null target, so one stored shape now resolves identically on both carriers. A non-string target that is not false (a number) or an unrecognised string (a typo such as resposne) is not the default — it matches no phase, the detector never runs, and it is reported active: false with inactive_reason: "wrong_phase". A value already stored on a gateway is grandfathered (carried through unchanged) so an existing configuration never becomes unsaveable; a stored null is normalised to request on the next save. A response/both detector inspects the model's answer, so on a streaming request the gateway first collects the full answer (it cannot un-send tokens already streamed) and then applies the detector before delivering it — a response-stage block/mask therefore adds the model's generation time as buffering latency before the first visible byte.
active boolean Whether the detector can ever match anything. false means it holds no content to match on, or is pointed at a phase it never runs in — the gateway returns pass for it on every request. Distinct from enabled: enabled: false is a deliberate switch-off, active: false is a detector that looks configured but is not.
inactive_reason string Present only when active is false. One of no_patterns, no_keywords, no_schema, no_languages, wrong_phase.

Detectors that can never match

A guardrail with nothing to match on is stored, listed and counted like any other, but the gateway returns pass for it on every request. Reading such a detector as protection is the failure mode this field exists to prevent — a regex with action: "block" and no patterns at all presents as the strictest setting available while blocking nothing.

inactive_reason Detector types Condition
no_patterns regex Neither patterns nor custom_patterns yields anything usable. A patterns entry that is not a known pattern or set name — a typo such as credit_card for cc — is silently ignored, so a non-empty patterns array can still leave the detector unable to fire.
no_keywords keyword, custom_pii keywords is absent, null, not an array, or holds no non-empty string. An empty-string keyword is skipped at match time, so [""] matches nothing.
no_schema json_schema schema is absent (or false). A schema that is present but unusable (JSON null, a scalar, an array, a malformed required / properties / constraint) is not reported inactive — the detector still runs and refuses (or, with action: "flag", flags) every response with invalid_schema_config until the configuration is corrected; POST / PATCH /gateways and the guardrail-config import reject such a schema with 400 so it cannot be stored.
no_languages language allowed is absent or an empty array. A null or non-array allowed is not reported inactive — the detector errors rather than passing, and a non-empty list of unusable entries blocks every request.
wrong_phase any type The target matches no phase, so the detector never runs — a typo (resposne) or a non-string that is not false (a number). An omitted target, a JSON null, and false all resolve to request and so are not wrong_phase — the detector runs. Also reported for json_schema and gibberish, which run only on the response phase, when target is left at the request default.

Only a detector the gateway provably treats as a no-op is reported inactive — the report mirrors each detector's own early return, including its edge cases, so a detector that errors or that blocks everything is never described as harmless. Detector types that ship working defaults are never reported inactive: a jailbreak with an empty keyword list keeps its built-in phrases; a prompt_guard with no categories enforces every category; a presidio or pii_protector with no usable entities analyses every entity type. A malformed detector, or a type the gateway does not recognise, is likewise reported as active — reporting a working detector as dead would lead an operator to delete protection that does run.

active is always present; inactive_reason only when active is false. The de_recognizers flag and the scrub_placeholder string are configuration inputs (on a pii_protector, resp. a regex / keyword scrub detector) — not returned fields. scrub_placeholder (default [REDACTED]) is the text substituted for each matched span when action: "scrub". On a keyword detector it is inserted literally (a % is not interpreted); on a regex detector it is used as a pattern replacement string in which % is a special character, so a literal % must be escaped as %%.

When a scrub cannot be applied. A keyword / regex scrub replaces text over the whole JSON document, so a keyword or pattern that matches a JSON key (content, role, text, model, …), or a placeholder that contains " or \ (which break JSON validity), would rewrite the document's structure rather than a value. The gateway compares the masked document's shape at every depth with the original — the same keys, the same array lengths, and only string values may differ (a number, boolean or null must be byte-identical: they are unquoted in the encoded body, so a digit pattern with a numeric placeholder would otherwise rewrite max_tokens or created, and a keyword true would flip stream) — and refuses to commit a mask that changed it. What happens next differs by phase, deliberately: on the request phase nothing is masked, the request is marked scrub_uncommittable and the detector reports pass — the body stays intact so the PII layer (tier 2) still sees and masks it; on the response phase the detector blocks (block_reason: scrub_uncommittable, no gap marker — the layer did protect) — there is no layer behind a response scrub, and a pass would deliver the matched term to the user. Previously the response phase had no check at all and shipped the corrupted body while recording scrubbed. Because a scrub on a key can never be applied — and on the response phase would block every response — POST / PATCH /gateways and the guardrail-config import reject a keyword scrub detector whose keyword is a JSON envelope field name (content, model, role, type, name, id, messages, choices, …) with 400; a regex pattern is not judged statically. Note in particular the unquoted numeric envelope fields of an OpenAI-shaped response (created, index, usage.*): the built-in phone and npi patterns (and the bundles that include them — pii_basic, gdpr_structured, hipaa_structured) match a ten-digit created timestamp, so a response-phase regex scrub with one of those patterns is refused on every response and blocks all of them; use such patterns on the request phase, or a pattern anchored to the value's real shape. The durable fix — scrubbing the string values of the decoded document instead of the encoded body — removes the key-, number- and boolean-hit classes at the root. See German structured-PII recognizers below for de_recognizers.

Required role: any user with access to the gateway's tenant.

Configuration shape and malformed input

The detector pipeline is the guardrails array in the gateway configuration. Each element must be a detector object. The legacy detectors key is accepted as a fallback, but only when guardrails is absent (the key is missing) or the boolean false; a present-but-null (or scalar, or object) guardrails value takes precedence and resolves to zero detectors — it does not fall back to the legacy detectors key. In other words, an explicit "guardrails": null behaves identically to an empty array "guardrails": [] (no detectors configured), while a config carrying only the legacy "detectors" key and no guardrails key uses those legacy detectors. The gateway never fails a request with a server error over a malformed value: a null array, a scalar, or an object in the guardrails position all resolve to zero detectors, and any non-object elements in the array (null, numbers, strings) are ignored. This is fail-safe, not a bypass: a gateway with no active PII detector still cannot serve a request under a project whose permission tier mandates PII masking — that request is blocked before anything leaves the gateway. When updating a gateway with PATCH /admin/v1/gateways/<ID>, a config of JSON null is treated as "no change"; a non-object config is rejected with 400.

The same fail-safe rule applies inside a detector object. The request body a detector inspects is raw client JSON, so no member of it is trusted: a messages value that is not an array, an element of messages that is not an object, a content that is a number, a text block whose text is JSON null, and a system / prompt / input / instructions of any non-string type are each skipped rather than indexed. Nothing in a request body can make a detector throw, because a throwing detector resolves through fail_open and — on the default fail_open: true — would let the request through unscanned. Likewise an entities value that is not an array of strings means "every supported entity type", never "no entity types".


German structured-PII recognizers (de_recognizers)

The hosted PII analyzer is US-centric and misses German structured identifiers — a Steuer-ID, for example, is scored as a weak phone-number match and would pass through to the model unmasked. Setting de_recognizers: true on a pii_protector detector adds gateway-side recognizers that catch these by format + checksum and tokenize them the same reversible way as every other PII value (the user sees the original in the final answer; the model and any logs see a token). They run on both the request body and tool results (read_file / knowledge / web-search output), so an ID arriving via a document is masked too.

Each recognizer validates a checksum, so an ordinary number of the same length is not masked (no over-redaction). Detected types and accepted shapes:

Entity Accepts Rejected
DE_STEUER_ID 11 digits, ISO 7064 MOD 11,10 check digit (separators tolerated: 12 345 678 901) any 11-digit number failing the checksum
DE_UST_IDNR DE + 9 digits, MOD 11,10 check wrong check digit; missing DE; leading-zero block
DE_KV_NR 1 uppercase letter + 9 digits (eGK Versichertennummer), Luhn check lowercase; wrong length; failed check
DE_SV_NR AA TTMMJJ B NN P Rentenversicherungsnummer, weighted check digit missing the birth-name letter; failed check
DE_PERSONALAUSWEIS 9 alphanumeric + ICAO-9303 (7-3-1) check digit digits-only without the check; failed check

Scope. This closes the leak of structured German identifiers. German names and addresses are free-text NER, handled separately by setting the detector's language to German — see below.

Survives a PII-analyzer outage (degrades closed). The German structured-ID recognizers are pure gateway-side format + checksum checks with no dependency on the hosted PII analyzer. If the analyzer is unreachable when a request is scanned, these recognizers still run and still mask the structured identifiers they recognize — so a Steuer-ID / USt-IdNr / KV-Nr / SV-Nr / Personalausweis is masked even during an analyzer outage on a fail_open gateway (where the free-text NER would otherwise degrade). The analyzer-outage signal is preserved independently, so a non-fail_open gateway still blocks on the NER outage (fail-closed) as before. This strengthens the "nicht abschaltbare Maskierung deutscher Personendaten" guarantee for the structured-ID family.

Accepted input, and what is rejected. The scanned text is caller-controlled — a message, an uploaded document, a fetched page, or a tool result — so it is treated as untrusted. Two classes of text are exempt from the scan, and no others:

  1. the regions this request has already masked, skipped so a digit run inside a redaction token is not masked twice (which would make the original unrecoverable) — described below;
  2. cells the table classifier reads as a numeric metric column rather than an ID column, a precision trade-off documented under German structured IDs in spreadsheets and tables. The tool, web-search and MCP egress seams re-scan suppression-free, so that exemption does not extend to them.

That exemption is keyed on the request's own mask record, never on the appearance of the text. Concretely:

Input Scanned
A redaction token produced by this request ([MYRA-REDACT-…], [MYRA-CUSTOM:…]) No — already masked
The same shape typed by the caller, e.g. [MYRA- … ] or [MYRA-REDACT-X …] Yes
A token from another request, or a truncated / re-cased one Yes — it cannot be restored, so there is nothing to protect
Malformed, empty, or absent mask record Yes — the whole text is scanned (absent and malformed share the protective path)

A shape-based exemption would be caller-controlled and would therefore let anyone put a checksum-valid German ID beyond the recognizers' reach by typing a prefix.

The same mask record governs the NER spans the pii_protector detector emits: an analyzer span — untrusted third-party output — is clamped in-bounds and then clipped to the prose around the request's own tokens before it is masked, so a span over a token this request produced masks nothing, a straddling span masks the text on either side, and a caller-typed token shape is masked like any other text. A span clipped away entirely is not reported as a masked entity.

The presidio detector's anonymizer request is canonical: per span exactly { entity_type, start, end, score } — a string label (entity_type, else the type alias, else PII), integer offsets clamped inside the field (a fractional offset widened, an out-of-range one clamped — a superset is masked, never less; only a span that cannot be placed at all is dropped), a finite score in [0, 1] — duplicates collapsed to one span with the higher score, whitespace-only spans dropped. Measured (replicated) against the sidecar: a null / empty / alias-only label is refused 422, an object label 500, a fractional offset 422, an infinite score cannot be encoded; a whitespace-only span is accepted but inserts a placeholder between two words; an exact duplicate is accepted with a rare sporadic 502. An anonymizer failure on a field with detected personal data is a fail-closed refusal, so the untrusted analyzer response is normalised before it reaches that boundary. The same label rule (presidio_chunk.span_label) names the span in every verdict and token, so a malformed label can no longer throw the detector into an error verdict that a fail_open detector would forward.


When the masker is down, under a PII mandate

A detector that cannot reach its sidecar produces no verdict. By default that degrades and proceeds (fail_open), which is a deliberate availability choice: a Presidio blip should not take a gateway down.

Under an active PII mandate it does not. On a pii_mandatory project — or a tenant with masking enforced — the detector that satisfies that mandate refuses the request instead of letting it through unprotected, and this is the one place the per-detector fail_open: true flag does not win. A mandate is a choice of compliance over availability; serving the request unmasked because a sidecar returned 502 would invert that choice silently, while the gateway went on being reported as compliant.

Scope, precisely:

situation outcome
mandate active, masker outage refused — guardrail_unavailable, marker pii_mandate_masker_unavailable, nothing egresses
mandate active, fail_open: true set explicitly refused — the mandate outranks the flag
no mandate, masker outage proceeds (fail_open honoured, guardrail_degraded recorded) — unchanged
mandate active, a NON-mandate-satisfying detector fails that detector's own fail_open applies — unchanged
wholly-local model leg unchanged: nothing leaves the estate, so there is nothing to refuse — and a tool result bound for that leg is likewise not a mandate case (a tool result on that leg; never persisted content — see the last row)
mandate active, the web-search query scan fails withheld — the search is skipped (the reply says web search is unavailable), whatever the detector's fail_open says; marker pii_mandate_masker_unavailable. The query's bytes leave for the search provider, so the mandate binds here even on a local-only gateway whose model leg is exempt.
mandate active, a tool-result scan fails (a pii_protector or a presidio scrub detector) withheld — the result is replaced by the [tool result withheld: PII scan unavailable] placeholder and the turn continues; marker pii_mandate_masker_unavailable. Not on a wholly-local leg (see above). A withhold is not a gap: none of guardrail_degraded / guardrail_error / pii_scan_degraded is written for it, so the row reads guardrail_gap: false.
no mandate, a tool-seam scan fails on a fail-open detector forwarded with at most the offline German-ID masking (unchanged) — and, recorded: guardrail_error (name + class), guardrail_degraded, pii_scan_degraded and guardrail_verdict = error, exactly as a request-phase fail-open is — where the forwarded bytes actually leave the estate: not on a wholly-local model leg, not for a detector the agentic-fetch inner leg's own request phase runs again over the document, and not when a later detector (or a later query of the same search turn) withholds the content after all (see Protection markers). Before, such a row read guardrail_gap: false.
mandate active, persisted content — a workflow form's uploaded file or an inbound email body on its way into the run state the persistence fence is a strict leg: the mandate binds whatever the detector's fail_open says and never as a wholly-local case (a local-only gateway's model-leg exemption does not reach persisted content). A masker outage withholds the content behind [tool result withheld: PII scan unavailable]; a mandate with no pii_protector on the gateway withholds it behind [content withheld: personal-data masking is mandatory …] instead of persisting it raw. The fence's throwaway context never reaches a request row, so the durable records are the [tool_result_ner_withheld] … gateway= error line and the model-error triage row (phase = persistence, classifier persisted_content_pii_withheld — its own row, distinct from the chat tool-result one). Ladder: how a run consumes an uploaded file.

The refusal is retryable and says so, and it names the privacy filter rather than claiming the gateway is "configured to fail closed" — on a fail_open: true gateway that would be untrue and would point the operator at a setting that is already the other way.

German names and addresses (language: "de")

The PII analyzer's NER model detects German PERSON and LOCATION entities, but its default confidence floor for those two high-false-positive types is 0.9 — tuned for English, where it is well separated from real names. German names and addresses routinely score 0.6–0.9 (e.g. a plain "Klaus Müller" ≈ 0.79, a city ≈ 0.71), so at the 0.9 default roughly a quarter of real German names and most addresses pass through to the model unmasked.

A request under a PII mandate gets this floor regardless of language — a pii_mandatory project, or a tenant with masking enforced. A mandate is the statement that masking is not optional, so it is applied rather than merely certified: a detector configured language: "en" on such a tenant still analyses and masks PERSON and LOCATION. Which legs: the model leg, unless the gateway is local-only (every configured provider first-party — that leg is exempt — for a chat turn; an agent-shaped run on such a gateway is NOT exempt, its output leaves the estate), and the model-generated web-search query always — its bytes leave for the search provider, so the local-only exemption does not reach it. Nothing is refused by this — on a compliance path the fail-safe direction is to mask more, never to serve less. The cost is false positives, since both types drop to the German-calibrated 0.6; a model leg whose floor came from the mandate records the pii_floor_forced_by_mandate marker, and a web-search query whose floor did records pii_egress_floor_forced_by_mandate, so the extra masking can be explained without reading source. Two seams get no NER floor at all and are governed by a separate refuse-on-structured-PII gate that deliberately lets free-text names through: URLs handed to fetch_url and MCP tool arguments.

Outside a mandate, the German PERSON / LOCATION calibration is fail-safe by default. Whether it applies is decided by the detector's language field as a tri-state — a request never fails over a malformed value:

language value German floor
absent, "", "auto", or a non-string (number/bool/object) ON (fail-safe — a PII gateway that omits language sends "auto" to the analyzer, which auto-detects German, so the 0.6 floor must apply)
"de" or a de-/de_-prefixed locale ("de-DE", "de_AT", "de-CH") ON, and the analyzer is forced to the German model (rather than auto-detect)
any other explicit locale — "en", "fr", and "deu" / "german" OFF (English 0.9 calibration)

A request that runs without the floor records the pii_floor_off_by_language marker in its request log — the mirror of pii_floor_forced_by_mandate above, so an operator can always tell from the record alone which calibration a given request actually got. It covers a masking detector, which loses the PERSON/LOCATION union and the 0.6 threshold, and equally a presidio detector that blocks or flags with one of those types ticked, which keeps its own list but loses the threshold and so stops acting on the German name band. It is not written for a bound detector that ticked neither type, where the floor would have changed nothing.

Accepted / rejected shape. Only "de" or a de-/de_-prefixed string is treated as German. "deu", "ger", and "german" are not — they read as non-German and turn the German floor off. Use "de".

When the German floor is ON:

  • the PERSON / LOCATION confidence floor drops to a German-calibrated 0.6;
  • PERSON and LOCATION are added to the analyzed set when the detector masks — that is, a presidio detector whose action is scrub (the default), and any pii_protector detector. (custom_pii is a keyword matcher — it has no entity list and no NLP analysis, so the floor does not apply to it at all.) Masking more types than were ticked never refuses a request, so the floor is applied there without asking;
  • but a presidio detector that refuses or flags — any action other than scrub — analyzes exactly the entity types you ticked, and nothing else — unless the request is under a PII mandate, where the floor still wins. A gateway certified as satisfying a mandate must actually analyze the floor types, so the selection cannot bind there. A detector that blocks a request on an entity the administrator deliberately excluded is a control that does not control, so the selection binds wherever it changes what the caller experiences. Tick PERSON if you want a blocking detector to act on names;
  • all other entity types keep their normal thresholds.

fail_open — accepted shape

fail_open decides what a detector does when its sidecar cannot answer: proceed (true) or refuse / withhold (false). Each phase has its own default when the field is absent — the request/response phases proceed (except a prompt_injection detector, which refuses unless it is set to flag), the tool-result and web-search-query seams withhold — and both true and false are real operator choices.

What is rejected. A write that introduces a fail_open that is not a literal boolean — an explicit JSON null, a string such as "false", a number — is rejected with 400 on POST /tenants/{id}/gateways, PATCH /gateways/{id} and the guardrail-config import, naming the field and the detector. A non-boolean is never an operator choice: every reader treats it exactly as an absent field (the phase's own default, as above), while the Guardrail Builder's checkbox renders it as one of the two real values and an export/import round-trips it, so the stored value and what the operator sees would never agree. A value already stored on a detector is grandfathered (carried through unchanged) so an existing configuration never becomes unsaveable; correcting it to a boolean repairs it.

timeout_ms — accepted shape

timeout_ms is the read timeout for a detector's Tier-2 sidecar call. It applies to the four detector types that make one — prompt_guard, prompt_injection, presidio and pii_protector; on any other type the field is ignored and not validated.

value result
absent the detector type's default (2000 for the classifiers, 15000 for the PII types)
a whole number in [1000, 120000] used as given
below 1000 400. The call cannot finish, so the detector becomes a guaranteed timeout — and on a fail_open detector the request is then forwarded unchecked while the gateway still reports an active protection
above 120000 400. It would hold the caller open on a guardrail that is not answering
fractional, NaN, infinite, a string, a boolean, an object 400
JSON null 400. It is not the same as omitting the field: null is truthy in the gateway's decoder, so it used to reach the socket call and throw

A value already stored is carried through unchanged, so an existing gateway never becomes unsaveable over a value it already had — only a write that introduces or changes one is rejected. Precisely, a write may keep an out-of-range timeout_ms when the same value is already stored on a detector of the same type in the same carrier, and only as many times as it is stored. That scope is deliberately name-blind: the detector name is editable on every card while the timeout field is editable on only one detector type, so a name-scoped rule would have made renaming such a detector permanently impossible. guardrails and detectors count as one carrier (they are two spellings of one list, and the UI migrates the legacy key on save); role_policies is separate. The one shape never carried through is a number that cannot round-trip through JSON (NaN, infinity) — a config holding one cannot be read back out of the API at all, so carrying it forward rescues nothing. Independently, the runtime clamps whatever it reads into [1000, 120000] and falls back to the type default for an unusable value, so a configuration that predates this rule is corrected rather than obeyed. "Unusable" there is narrower than what the write boundary refuses: the runtime coerces anything numeric, so a carried-through string like "5000" (or "0x1388", or one padded with spaces) is read as that number and then clamped, while a boolean, an object, null and NaN all fall back to the type default. Either way an unusable stored value can never reach the socket call — which is the half a write-time check cannot deliver on its own, because it never sees a configuration that was written before the rule existed.

An entities value that is not an array — including JSON null and a scalar — means "analyze every supported entity type". It is rejected as a selection, never partially applied: a malformed value must not silently narrow what a detector inspects.

{ "type": "pii_protector", "action": "scrub", "language": "de",
  "entities": ["PERSON", "EMAIL_ADDRESS", "IBAN_CODE"], "de_recognizers": true }

Operator-visible controls. Both axes are exposed in the gateway Guardrails editor (no raw JSON needed). The Language calibration select (Automatic → clears language; German → "de"; English / other → "en") is on the pii_protector card and, on the presidio card — one shared control, because both detector types read the field identically and the floor it decides was previously unreachable on the Presidio card altogether. The select writes the exact value and English keeps the English 0.9 calibration — it does not clear language. The German ID recognizers checkbox (de_recognizers, off by default) stays on the pii_protector card alone: guardrails/presidio never reads it. de_recognizers runs independently of language (structured-ID masking does not require German name calibration, and vice-versa).

Switching a masking detector from a German-calibrated language to English / other now asks for confirmation before it is applied, and the requests that then run without the floor are marked with pii_floor_off_by_language in the request log — the control used to reduce protection in one click, silently, and leave nothing behind that could explain it afterwards.

Notes:

  • score_threshold and every entity_score_thresholds value are confidences in [0,1]. A value outside that range (or a non-numeric one) is rejected and the detector falls back to the protective default 0.7 — it is not clamped to 1.0. This is fail-safe: a threshold of e.g. 5 (a typo for 0.5) would otherwise be impossible for any candidate to meet, so Presidio would return no entities and PII would reach the model unmasked. The gateway editor also blocks out-of-range input at the source and warns when the threshold is set above 0.85 (the highest built-in entity floor, ORG — above it the global starts filtering entities server-side, including structured high-value PII like email/IBAN/card). An out-of-range per-entity override is ignored (the entity uses the global / high-FP value instead).
  • The German PERSON / LOCATION floor is a hard floor: a per-entity entity_score_thresholds override may only make it stricter (a lower value), never weaker — so a misconfigured "PERSON": 0.95 cannot silently re-open the leak.
  • The 0.6 floor is a deliberate, false-positive-safe calibration and is not lowered further (analysis 2026-07-17): rare real cities scoring 0.5–0.6 (e.g. "Bremen" 0.564) are not masked — the accepted, bounded trade-off; most real places score ≥ 0.7 in context, and lowering the floor admits measurable spurious LOCATION spans.
  • A German-configured gateway is calibrated for German text. English text processed through it may over-redact common words that the NER model scores as names — this is intentional (a PII boundary errs toward masking). Use a separate English-configured gateway (language: "en") for English workloads.
  • Combine with de_recognizers: true to also catch structured German identifiers (Steuer-ID, KV-Nr, …) in one detector.
  • A presidio detector honors the same language tri-state (identical floor), and the same Language calibration select is on its card too. (The German ID recognizers checkbox is still pii_protector-only, because guardrails/presidio does not read de_recognizers.)
  • language must be a non-empty string, or absent. A write that introduces null, a number, a boolean, an object, a string that is empty or whitespace-only, or one carrying a control character is rejected with 400; a value already stored is carried through so the configuration stays saveable, and the Guardrail Builder drops it on the next save of the Guardrails section. A role_policies detector has no editor, so a malformed value there must be repaired through the API or a guardrail-config import — the same known limitation the legacy url field has. At runtime any non-string resolves to "auto", which is calibrated as German — so a malformed value masks more, never less. Previously a stored "language": null threw while the analyzer request was built, the orchestrator recorded a detector error, and a fail_open: true detector then forwarded the request completely unmasked.

Name-recall propagation

The NER model's PERSON recall is case- and repetition-sensitive: it reliably catches a capitalised first mention ("Anna") but can miss a later lowercase or repeated occurrence ("anna") of the same name in the same text. To close that under-redaction gap, a pii_protector detector automatically propagates every name the model already confirmed as a PERSON to its other whole-word, case-insensitive occurrences in the same message, masking them with their own (case-exact, reversible) token.

  • allow_list_match must be "exact" or "regex", or absent. A write that introduces or changes it to anything else — "partial" in particular, which the Guardrail Builder offered and this page once documented — is rejected with 400 on POST, PATCH and the guardrail-config import; a value already stored is carried through. The analyzer answers HTTP 500 to any other mode, which the detector reports as an error and a fail_open: true detector then turns into an unmasked pass — so at runtime an unsupported value is coerced to absent, which is the analyzer's exact default (narrower allow-listing, i.e. more masking). Only pii_protector sends the field; presidio is allow-list-blind and a value stored on one is inert.
  • allow_list on a presidio detector is refused. That detector builds its own analyzer request — text, language, entities and score_threshold, nothing else — so an allow list on it was never sent and never honoured, and on an action: "block" detector the request was still refused on the value the operator had explicitly exempted. A write that introduces the field on a presidio detector is rejected with 400 on POST, PATCH and the guardrail-config import; a value already stored is carried through, dropped on the next save from the Guardrail Builder, and stripped from an exported guardrail-config file so the file stays importable onto another gateway. The field remains valid and honoured on pii_protector and custom_pii.
  • This is always on — it needs no configuration and changes no score threshold.
  • It introduces no new false-positive class: only a string the model already classified as a name is propagated.
  • It is bounded to whole-word matches, so a name embedded in a larger word (e.g. "anna" inside the German word "Annahme") is left untouched.
  • Like the German calibration, it errs toward masking: if a confirmed name also appears as a common word in the same message, that occurrence is over-redacted (and restored on the response) rather than leaked.

Limitation — this is propagation, not new detection. A name the model never scores above the confidence floor in any form (an unusual/non-Western name, a typo, or text where the only occurrences are lowercase) is still missed; that is bound by the NER model and its threshold, not by this gateway logic.

Deterministic remedy for known names. For names you already know (employees, frequent contacts) that the NER model scores below the floor, add them to a custom_pii keyword detector on the gateway (or, per end-user, the My-PII keyword list). A custom_pii keyword is matched literally, whole-word, and case-insensitively by default — so it masks a lowercase-only name (anna) or an unusual/typo'd spelling the NER model would miss, at any confidence and before the request leaves the gateway, regardless of the PERSON floor. This is the intended lever for the residual above; it does not depend on Presidio detecting the name.

Multi-word terms also match common separator and case variants. A protected term of two or more words (e.g. Project Orion) is matched not only as the exact phrase but also in its snake_case (project_orion), kebab-case (project-orion), concatenated / PascalCase / camelCase (ProjectOrion, projectorion) and re-spaced forms — case-insensitively unless the detector sets case_sensitive. This closes the exact-match bypass where a user (or a tool result) wrote the term in a different casing or with a different separator. The generated variants are always matched whole-word (so ProjectOrion is masked but myProjectOrionX is not) and all variants of one term share a single reversible [MYRA-CUSTOM:…] token, restored to the canonical configured term. A term is expanded only when it has ≥2 words, each ≥2 characters, ≥6 characters total — single-word and trivially short terms (Go To) keep exact-match behaviour to avoid noise. Accepted input is any keyword string, including non-ASCII; a keyword that is not valid UTF-8 degrades safely to exact-match only (never a malformed or over-broad variant).

Documented limits of the keyword matcher. The matcher operates on text, so:

  • Concatenated/hyphenated dictionary words. A length threshold cannot tell an identifier from a real word, so a protected term whose words concatenate or hyphenate into a common word (Data Base → database, Log In → login) will also mask that word wherever it appears — including inside a tool / RAG result the model needs for grounding. Choose distinctive multi-word terms, or expect the common form to be masked too.
  • Splitting across messages. If the words of a protected term are sent in separate messages (never adjacent in one field), no single field contains the term, so it is not masked at the request phase — masking individual generic words would over-mask. If the model then reconstructs and emits the term, the response sweep tokenises it.
  • Context inference. Masking replaces the term, not the surrounding facts. A model can still reconstruct a masked term from uniquely identifying clues left in the user's own text (a proprietary language + database + version fingerprints a product; a unique penalty structure fingerprints a regulation). Generic descriptions and arbitrary identifiers stay safe. This is inherent to a text matcher over un-maskable context; when masking a well-known term, sanitise the identifying context clues too.

Custom keywords are also masked on tool-result egress. A configured keyword that arrives back through a tool result — a read_file of an uploaded document, knowledge / RAG retrieval, a web_search result, a fetch_url page, or an MCP tool result — is re-masked with the same reversible [MYRA-CUSTOM:…] token before that result is dispatched to the model, exactly as the request phase masks it (the two paths share one keyword tokenizer and one per-turn token map, so a keyword seen on both maps to one token and is restored once in the final answer). This closes a leak where a custom keyword returning through any tool result previously egressed to the provider in clear. The masking is fail-closed: because a tool result is untrusted content, a fault while masking it withholds the whole result behind the [tool result withheld: PII scan unavailable] placeholder rather than forwarding the raw keyword. (Accepted input: any tool-result string; a malformed / empty / non-array keyword list is treated as "no keywords" — the same inert result an absent list gives, consistent with the request phase.)

Why the English PERSON floor is not simply lowered to catch these. A live GLiNER sweep of 30 must-mask names against 130 benign German/English business prompts (tests/false_positives/person_floor_analysis.py) found no confidence floor that is false-positive-free: at every gate the model masks capitalised common-noun role terms (Der Kläger 0.99, General Manager 0.93) as PERSON, and lowering the English floor from 0.9 to 0.6 multiplies English false positives ~7× while recovering only two of the sub-floor names. So the floor stays calibrated per language (German 0.6, English 0.9), and the keyword lever above — not a blind floor drop — is the deterministic answer for known low-confidence names. For German deployments, keep language unset / "auto" / "de" (an explicit "en" opts into the 0.9 English floor and drops sub-0.9 German names).


Web-search query egress (fail-closed)

When a gateway runs web search, the search query the model generates is itself an outbound string sent to a third-party search provider (Brave, operated in the US). A model can put personal data into that query, which would otherwise leave the gateway unmasked — bypassing the request-phase PII layer entirely.

If the gateway has a pii_protector detector, the gateway runs the same reversible pseudonymization over the model-generated query, at the trust boundary, before it egresses:

  • No PII detected → the query is sent unchanged.
  • PII detected → each value is replaced with a [MYRA-REDACT-…] token (the same token the request phase would assign); the search provider receives only the tokenized query. This applies to every third-party egress of the query — the search request and the optional weather lookup — and to the query as it appears in the operator logs, the X-Web-Search-Query response header, and the search-result trace.
  • PII scan unavailable (the analyzer is down, the time budget is exhausted, or the scan errors) on a detector that is not fail_open — or, under a PII mandate, on any detector (the mandate outranks the flag here; see When the masker is down) → the gateway fails closed: the raw query is not sent to the provider, and the model is told the search was withheld. Data is never leaked to satisfy a search.

Accepted / rejected shape. The query is accepted as a non-empty string. A query carrying detectable PII is accepted but rewritten to tokens before egress; a query that cannot be scanned (analyzer outage) is rejected for egress (the search is withheld) unless no PII mandate is in force and the detector explicitly opts into fail_open — that forward is recorded on the request as a guardrail gap (guardrail_degraded / pii_scan_degraded / guardrail_error).

Scope. The fail-closed guarantee is provided by a non-fail_open pii_protector detector — the only detector type that reliably fails closed. A gateway without such a detector does not pseudonymize the query (a presidio scrub detector fails open on outage and cannot give a third-country guarantee). Under a PII mandate — a pii_mandatory project, a forced agent invoke, or a tenant with masking enforced (an unresolvable project tier blocks every egress tool outright, one step earlier) — a gateway with no pii_protector therefore withholds the search rather than send the raw query (keyed on the mandate from any source, not only the tenant flag, and on both search engines with one wording: "this workspace requires personal-data masking, but no reversible PII masker is configured"). The local-only model-leg exemption does not reach this gate: the query's bytes leave for a third party. The data flow leaving the gateway to Brave is therefore: model query → pii_protector tokenization → provider sees only tokens, or the search is not sent at all. (This section covers the two model-generated web-search paths; the admin playground search box sends operator-typed text and is out of scope.)


Injection scanning of tool results

Tool results — the content a server-side tool leg fetches and fuses back into the next model call (read_file, knowledge/RAG passages, web-search results, MCP tool output, agentic fetch_url, code-interpreter stdout) — are untrusted third-party / model content. An indirect prompt injection embedded in such content (a poisoned document, a hostile web page) would otherwise reach the model unscanned, because the request-phase detectors run on the user's prompt, not on tool output.

This includes text extracted from a fetched document (PDF / DOCX / XLSX pulled by fetch_url / agentic_fetch; see Web search → Document fetch). The raw document bytes are untrusted: they are magic-byte-sniffed to an allow-list, never executed, and parsed only by the sandboxed extractor; the resulting text then passes through this same tool-result trust boundary (injection scan + PII masking) before it reaches the model. On the agentic_fetch inner-model leg the extracted document text is masked/scanned before the inner model, and under a PII mandate (a pii_mandatory project, a forced agent invoke, a tenant with masking enforced) on a gateway with no reversible masker the document is withheld outright (fail-closed). The document goes to the inner model, so this is a model-leg decision: a wholly-local inner model is not a mandate case (nothing leaves the estate, and the inner request phase skips masking there too), and on a local-only gateway of an enforcing tenant the exemption applies — the same leg rule the tool-result seam uses. Inside an agent-shaped run the inner leg is never treated as local and the run is not exempt, so the document is withheld there under a mandate. A scanned/no-text-layer PDF read via OCR is additionally lower-confidence and reaches the model labelled as such (may contain transcription errors — verify figures/dates against the source), so it is never presented as authoritative.

When the gateway web search tool is offered, fetch_url / agentic_fetch are now also offered so the model can open a link it found in the search results (co-arm). The URL is then an untrusted model choice driven by untrusted search-result content, so every such fetch stays gated by the SSRF/redirect guard, the outbound PII/secret exfiltration check (fail-closed), the refuse-on-structured-PII URL gate (armed on a PII-active gateway or under a PII mandate from any source), and this tool-result injection scan on what comes back; a bounded per-turn fetch cap limits fan-out.

If the gateway has a jailbreak detector, the gateway now runs that detector's matcher over each tool result as well, at the same trust boundary where tool-result PII masking happens, before the result re-enters the model. Which jailbreak detectors apply is decided exactly as on the request phase: a detector with target request, both, or absent applies (a tool result is the next leg's model input); a response-only detector does not.

Evasion handling. Before matching, each result is normalized and decoded so the common lightweight obfuscations do not slip an injection past the matcher: whitespace runs are collapsed, zero-width / BOM characters are handled both as intra-word noise (stripped) and as word separators (mapped to a space), and the text is expanded through the same evasion-decode used by the exfiltration egress guard (percent-decoding incl. double-encoding, ROT13, and per-token Base64 / hex). This decode expansion has fixed bounds: it is single-layer (one Base64/hex decode, not nested base64(base64(…))) and skips any single encoded token longer than 4 KB; the scan itself covers the first 16 KB of each result (a defensive bound shared with the egress guard). An injection that is nested-encoded, hidden in one very large Base64 blob, or placed entirely beyond the 16 KB offset is therefore not decoded/scanned.

Action. The detector's own action decides the outcome, per gateway configuration:

  • action: "block" → the offending result is withheld: its content is replaced with a neutral placeholder before fusion, so the injected text never reaches the model. The rest of the turn proceeds.
  • action: "flag" (the default) → observe-only: the match is recorded but the result content is passed through unchanged. An injection in a flagged gateway therefore still reaches the model — set action: "block" for enforcement. On a top-level turn the flag also queues a moderation review_item (trace reference + detector name only — never the matched content; see Agents → moderation queue).

Accepted / rejected shape. A tool result is accepted as a string. A result that matches an applicable block jailbreak detector (in its raw, normalized, or decoded form, within the first 16 KB) is rejected for fusion (withheld behind the placeholder); all other results pass through. A withheld result is never additionally PII-processed.

Scope and limitations. The matcher is the jailbreak keyword list (the built-in set, or the detector's custom keywords when configured) — a literal-phrase tripwire. It catches known jailbreak / injection phrases and their encoded/whitespace-obfuscated variants; it does not catch paraphrased, non-English, or homoglyph-substituted injection. Because it is a keyword match, benign content that legitimately quotes such a phrase (for example security documentation, or a user's own note containing "ignore previous instructions") will also be withheld under a block detector — the accepted cost of the tripwire. Operators tune this by choosing flag vs block and by supplying a custom keywords list; each match is logged with the tool index and the matched keyword.

Classifier scanning (prompt_injection)

If the gateway has a prompt_injection detector, the Llama Prompt Guard 2 classifier also scans each tool result at the same boundary — a trained, multilingual model that catches paraphrased and non-English injections the keyword tripwire cannot. It runs in addition to any keyword scan; a result already withheld by the keyword scan is not re-classified. Applicability follows the same target rule (request/both/absent).

Each result is classified over its first ~4 KB (injection instructions sit at the head; the model reads roughly the first 512 tokens). The work is bounded so a hostile source cannot amplify it into a denial of service: at most max_texts results per tool leg (default 64), batched in small POSTs, under a per-leg wall-clock budget (tool_budget_ms, default 5000 ms).

Classifier configuration (system-global, DB). The classifier endpoint, its basic-auth credential, and the two resource bounds above are system-global configuration stored in the database (not deployment configuration, not per-gateway) — set once by a platform admin via the classifier-config endpoint below. Until a platform admin configures it, every gateway with a block-action prompt_injection detector treats the classifier as unconfigured and fails closed (blocks, unless the detector is fail_open). A flag-action detector passes the request through instead: observe-only never enforces, on a detection or on an outage, and fail_open is inert for it — including an explicit fail_open: false. The credential is stored encrypted and is never returned by any read. A configuration change propagates to all serving hosts within a short cache TTL. Only the per-detector injection_threshold cutoff is per-gateway (in the gateway's guardrail config).

Action and fail-closed. A result whose p_injection >= injection_threshold under a block-action detector is withheld behind the placeholder; a flag detector logs it and passes it through. On a classifier outage — or when a result is left unscanned because a bound was hit — a block-action detector that is not fail_open withholds the affected results (fail-closed), so injected content never reaches the model unscanned; a flag-only or fail_open detector passes them through — and that pass-through is recorded on the request when its cause is an outage: an unconfigured classifier (guardrail_error.error_class config), a failed classifier call (http_NNN / transport), or a spent tool_budget_ms (budget_exhausted) write guardrail_error, guardrail_degraded and guardrail_verdict = error, so the row counts as a guardrail gap and in the guardrail_unavailable alert — for a flag-only detector too, exactly as the request phase records its outage. A result left unscanned by the max_texts cap is a designed bound, not an outage, and is not recorded. This detector defaults to fail_open: false.

Accepted / rejected classifier response. The classifier's answer is untrusted. Per submitted text the gateway accepts only an object with label ∈ {"benign", "injection"} and p_injection a finite number in [0, 1]; the numeric probability is authoritative for the verdict (p_injection >= injection_threshold) and the label is never used to override it. A non-2xx status, transport failure, non-JSON/scalar/array body, a results array whose length does not match the submitted texts, a missing/non-string/out-of-enum label, or a missing/null/non-number/NaN/out-of-[0,1] p_injection is rejected and treated as a classifier outage (governed by fail_open). See Prompt Injection → Classifier output validation.


Prompt-injection classifier config (system-global)

Platform-admin endpoints to configure the shared Prompt Guard 2 classifier. Platform-admin only (permission PROMPT_INJECTION_CONFIG_MANAGE); a caller without it receives 403. Every write is audited (hash-chained), recording the endpoint and the resulting configured flag — the credential is never written to the audit log, returned by a read, or logged.

GET /admin/v1/system/prompt-injection-classifier

Returns the current config without the credential:

{ "endpoint": "https://…/classify", "configured": true, "tool_budget_ms": 5000, "max_texts": 64 }

configured is true only when the stored config would resolve into a usable classifier (a present, decryptable credential and a valid https endpoint). An unset config returns { "endpoint": "", "configured": false, "tool_budget_ms": 5000, "max_texts": 64 }.

PUT /admin/v1/system/prompt-injection-classifier

Body:

Field Type Notes
endpoint string, required The classifier URL. Must be https:// and must not embed userinfo (user:pass@host) — the credential must never ride the logged/stored URL.
basic_auth string, required The user:password basic-auth credential. Stored encrypted; never returned or logged.
tool_budget_ms integer, optional Per-tool-leg wall-clock budget. Clamped to [100, 60000]. Omitted ⇒ the existing value is preserved (default 5000 on a first write).
max_texts integer, optional Per-leg result cap. Clamped to [1, 1000]. Omitted ⇒ preserved (default 64).

Rejected (400): a missing/empty endpoint or basic_auth, a non-https scheme, a userinfo-bearing URL, or a non-numeric tool_budget_ms/max_texts. On success (200) the change takes effect on the next classification, propagating to all serving hosts within a short cache TTL. There is no SPA surface for this — it is a platform operations setting, set via the API.


Per-role policy (role_policies)

A gateway can bind a stricter guardrail profile and/or a narrower model allowlist to a specific user role within the single gateway, via an optional role_policies block on the gateway configuration:

"role_policies": {
  "member":  { "guardrails": [ /* extra detector objects */ ], "models": [ "qwen-eu", "gpt-eu" ] },
  "default": { "models": [ "qwen-eu" ] }        // applies to any role without its own entry
}

The role is the server-derived authenticated role — a caller cannot assert its own role. The dimension is tighten-only: it can only make policy stricter than the gateway base, never looser, so the gateway base configuration remains a hard floor that a role override can never weaken or bypass.

  • guardrails — EXTRA detectors, unioned onto the gateway's base detectors for that role (a role can add strictness; it never removes a base detector, which always still runs). A role detector runs in the phase(s) its target names: request (the default), response, or both — response-phase role detectors are honored (a response/both role detector on a streamed request forces the answer to be internally buffered so the detector actually runs, exactly as a base response detector does). Role detectors run after the base set (they are not re-tier-sorted), so a role detector observes the text after any base masking — it cannot block on content a base PII detector has already scrubbed. A role detector's target reads the same way as a base detector's: an omitted, null or false target resolves to request — a null-target role detector runs on the request phase, the same outcome the base runtime now produces, so one stored shape resolves identically on both carriers.
  • models — an intersection mask: a member of that role may use only a model that is also in this list (and still within any gateway/plan allowlist — the gates compose, so a role can never grant a model the base disallows). An empty models list denies every model for that role (a deliberate "no model here" configuration; note this is the opposite of an absent models key, which imposes no mask). The mask is enforced on the final resolved model, at the guardrail stage — after routing, per-turn Auto model upgrades, and self-serve failover have run — so a system-initiated model swap cannot carry a role onto a model its mask excludes (the residency / data-sensitivity guarantee). A request whose final model is outside the role mask is blocked with the guardrail_role_model class. The mask governs the primary (user-chosen) model only — inner/system legs (an agentic fetch_url sub-request, a tool-loop continuation, a web-search probe) run on a gateway-chosen system model and are not role-mask-gated. The mask is also re-asserted on the primary turn's upstream swaps — the retry-time alternate-model swap and the provider-fallback-chain swap — at the per-attempt dispatch point before any network call, so a fallback can never egress the user's turn to a model the role excludes; an all-denied chain returns the role_model_not_allowed (403) error code.

Role matching. role_policies keys must be one of the system roles (admin, tenant_admin, ki_manager, member, viewer, demouser, finance) or the literal default. A role with no matching entry uses the default entry if present, otherwise the gateway base — i.e. unknown/unlisted roles default to base (allow); enumerate every role (or set default) if the intent is default-deny. A custom (non-system) RBAC role never matches and is treated as a config typo: a role_policies block with an unknown role key — or a malformed shape (a non-object block, or a non-object role entry) — is rejected 400 at every gateway-config write boundary (create, PATCH, and guardrail-config import), never silently stored to no-match at runtime.

Residency deployments — set a fail-closed default. A request with no role (an anonymous / user-less token, or an auth_required=false gateway) resolves to a nil role, which lands on the default entry if present, otherwise the gateway base (allow). If the mask exists to enforce residency, a nil-role request must not float to the looser base — set a restrictive default (or make the gateway base itself the strict floor) so an unauthenticated or user-less request cannot obtain a broader model set than any named role.

Accepted / rejected shape. role_policies and every nested field are accepted as objects/ arrays; any malformed shape (null, a scalar, a wrong type at any level) is treated as no override → the gateway base applies — it can never weaken the base or error a request. An out-of-mask model is rejected (403); an in-mask model and an unmatched role pass through.

Agent invokes. A saved-agent / agent-as-tool invoke runs under the agent owner's role (the existing runs-as-owner model), so its turn is processed with the owner's role policy, not the calling user's.


Reading guardrail statistics

GET /admin/v1/gateways/<ID>/guardrail-stats

Returns a single aggregate object covering the last 24 hours of traffic for the gateway. The figures are gateway-wide totals, not a per-detector breakdown:

Field Type Description
blocked integer Requests blocked by a guardrail.
scrubbed integer Requests scrubbed (but not blocked).
flagged integer Requests where a detector fired without blocking or scrubbing.
degraded integer Requests this gateway did not fully protect — those that reached a provider without the protection they were configured to get. Same definition as guardrail_outcome=degraded on the Logs API. Omitted for a caller without the REQUEST_LOGS_VIEW permission, who also receives only the acted events below.
avg_guardrail_ms integer Average guardrail processing time, in milliseconds.

Required role: any user with access to the gateway's tenant.

The endpoint does not accept time-range query parameters; the window is fixed at the trailing 24 hours.


Listing recent guardrail events

GET /admin/v1/gateways/<ID>/guardrail-events

Returns recent guardrail activity, newest first — both directions of it:

  • the guardrail acted: it blocked, scrubbed, or fired a detector; and
  • the guardrail did not protect the request: the request proceeded despite a detector failing open, a scrub that could not be committed, a body with no scannable text, unmasked PII egressing under an active mandate, a guardrail_verdict of error or indeterminate, or a meta column too large to store in full. A request refused before dispatch is not this — nothing left — even when the refusal happened because the detector errored and its fail mode is closed. One blocked on the response phase, after the prompt already went to the provider, still is.

The second kind usually sets none of blocked, scrub_applied or detectors_fired, so it used to have no row here at all — a fail-open bypass was invisible in the panel that exists to report guardrail activity. It is now reported, and meta names which detector lapsed.

Required role: any user with access to the gateway's tenant, as before. A caller without the REQUEST_LOGS_VIEW permission (held by admin, tenant_admin, and any custom role granted it) receives the narrower acted-only event set this endpoint has always returned, with no meta — the protection markers are per-request facts that the Logs API gates the same way, and one of them records that personal data actually left the estate. Such a caller is never handed a gap event it has no way to explain.

Each event contains:

Field Type Description
ts integer Event time as Unix seconds.
blocked integer 1 if the request was blocked.
scrub_applied integer 1 if scrubbing was applied.
detectors_fired array Names of the detectors that fired.
blocked_by string Detector that caused the block, if any.
block_reason string Reason for the block, if any.
guardrail_latency_ms integer Guardrail processing time, in milliseconds.
guardrail_verdict string Overall guardrail verdict. One of safe, unsafe, error, or indeterminate. safe and unsafe are content verdicts from Prompt Guard. error means a guardrail backend could not be reached or failed. indeterminate means a guardrail backend answered, and the answer could not be read as a verdict — for example a classifier replica returning repetition garbage behind HTTP 200. When several guardrails run in one request, the health verdicts (indeterminate, then error) take precedence over the content verdicts, so a degraded scan is never masked by a later clean one.
provider string Provider the request was routed to.
model string Model used.
latency_ms integer End-to-end request latency, in milliseconds.
meta object The guardrail protection markers for this request, including guardrail_gap and — when a detector lapsed — guardrail_error.name. Same shape and same curation as on the Logs API. Absent for a caller without REQUEST_LOGS_VIEW.
upstream_attempts integer Provider dials attempted (0 = never dispatched). Present because the protection-gap verdict depends on whether the request reached a provider at all.

There is no request or log identifier on these events.

Required role: any user with access to the gateway's tenant.

Optional query parameter:

Parameter Type Description
limit integer Maximum events to return.

Reading the circuit breaker state

GET /admin/v1/gateways/<ID>/circuit-breaker

Returns the circuit breaker state for each provider that currently has a live breaker entry on the gateway, keyed by provider name. Providers with no breaker state are omitted. Each entry has:

Field Type Description
state string One of closed, open, or half_open.
failures integer Current consecutive failure count.
opened_at integer Unix seconds when the breaker opened. Present only while the breaker is open or half-open.

Required role: any user with access to the gateway's tenant.


Masking preview (scrub-preview)

POST /v1/<tenant>/<gateway>/scrub-preview

A billing-free dry run of the request-phase PII masking. It never calls the model; it returns what the guardrails would mask so a user can review — and selectively adjust — the masking before sending. Authenticated with a short-lived gateway playground token (x-aig-token).

Accepted request shape. A JSON body { "text": string }. One optional header is consumed: x-project-id scopes policy resolution to a project exactly as a real send does, so the preview reflects that project's PII mandate (a pii_mandatory project → values withheld, below). It is untrusted input, validated and fail-closed: a malformed or repeated x-project-id is treated as a mandate in force (values withheld), never as absent. Rejected inputs (each fails closed, the message is never scanned or sent):

Condition Response
Method is not POST 405
Body is empty / unreadable 400 { "error": "empty body" }
Body is not JSON, or text is not a string 400 { "error": "expected JSON body { text: string }" }
text is the empty string 200 with an empty, non-degraded result (nothing to scan)
The request phase returned a body whose first message no longer carries a string content 500 { "error": "guardrail post-processing failed", "code": "configuration_error" } — the same refusal the conversation summarisation route gives, and it takes precedence over any detector verdict (even a block): a corrupted pipeline is the truthful answer. The preview never answers 200 with a fabricated empty masked (that was a legacy behaviour: masked: "", degraded: false, "everything was masked"). This should not occur — a keyword/regex scrub that would change the document's shape is refused before it is committed (see when a scrub cannot be applied) — so if it fires, a masking layer rewrote the body outside that guard; the gateway error log names the gateway, the modal shows the error, and the failure counts in the health dashboard's scrub-preview failures. Report it rather than working around it.

Response shape.

Field Type Description
masked string The draft after masking — what the model would receive. Unmasked when nothing matched.
detected array One entry per detected entity type: { type, count, detector, action }. Empty when nothing was detected.
spans array Per reversible-redaction token { token, type, value }, so the UI can offer selective unmask. Only [MYRA-REDACT-*] tokens carry a span. Withheld under a PII mandate: when masking is mandatory for the caller — a pii_mandatory project (resolved from x-project-id), a tenant with pii_masking_enforced, an agent-area mandate, or a fail-closed indeterminate tier (the same pii_mandate_in_force condition that denies a send-time unmask) — spans is [] and no raw value appears anywhere in the response, while detected still reports what was masked. The client is never the authz boundary, so the preview does not hand a caller values it could never unmask.
would_block boolean True when a detector's action is block and the request would be refused.
blocked_by string Name of the blocking detector, present only when would_block is true.
degraded boolean True when a PII masking layer (presidio / pii_protector / custom_pii) could not run on a fail-open gateway (analyzer outage — 401/5xx/timeout/connection-refused, or a detector crash). In that case the draft was not fully scanned, so masked / detected are byte-identical to a genuine "no PII found" result and must not be presented as clean. Always present. A fail-closed PII outage instead returns would_block: true (the request is refused), not degraded. Only PII detector types set this — a non-PII detector (e.g. jailbreak) degrading does not.

Why degraded exists. On a fail-open gateway an unavailable analyzer makes the gateway proceed without masking rather than block. Without this field the preview would show zero detected entities — indistinguishable from "no personal data found" — and mislead the user into believing their draft was scanned and is clean. The masking-preview UI renders a distinct "PII detection incomplete" warning when degraded is true. This surfaces the degraded state honestly; it does not change the fail-open policy (whether a gateway fails open is a tenant/operator decision).


Guardrail configuration as code (export / import)

A gateway's guardrail configuration can be exported to, and imported from, a versioned JSON file for GitOps / infrastructure-as-code workflows. Export is read-only; import applies a file to a gateway and is audited + hash-chained like every admin mutation. Both endpoints require the GATEWAYS_MANAGE permission (enforced server-side — the client is never the authz boundary).

The file describes only the guardrail subset of the gateway configuration, never the whole config. The exported/accepted keys are exactly: guardrails, detectors (legacy detector array), role_policies, pii_thresholds, egress_guard, eu_region_routing, provider_allowlist_enforced, provider_allowlist. No other key is exported (so no BYOK key, webhook secret, web_search.api_key, budget, or internal field can leak) and any other key in an imported file is rejected.

File shape

{
  "kind": "myra.gateway.guardrail-config",
  "schema_version": 1,
  "gateway_slug": "acme-prod",
  "exported_at": 1691000000,
  "config": {
    "guardrails": [ /* detector objects */ ],
    "role_policies": { "default": { "guardrails": [ /* ... */ ] } },
    "eu_region_routing": true,
    "provider_allowlist_enforced": true,
    "provider_allowlist": ["anthropic", "myra"]
  }
}

kind and schema_version are load-bearing (validated on import). gateway_slug and exported_at are informational — import targets the gateway in the URL, so a file exported from one gateway can be promoted to another (e.g. staging → prod). A detector's url field is stripped on export and rejected on import (in every carrier — see below), so no internal detector endpoint is ever written to the file.

GET /admin/v1/gateways/<ID>/guardrail-config/export

Returns the file above as a downloadable application/json attachment (<slug>-guardrails.json). A gateway with no guardrail keys exports "config": {}.

POST /admin/v1/gateways/<ID>/guardrail-config/import

Body: the file above, optionally with "confirm": true. The import declaratively replaces the guardrail subset — every guardrail key present in the file is set, every guardrail key absent from the file is removed, and all non-guardrail keys (BYOK, webhooks, budgets, …) are preserved untouched. (This differs from PATCH /admin/v1/gateways/<ID>, which merges the keys you send; import is a full declarative replace of the guardrail subset.)

Accepted / rejected input (each rejection fails closed — nothing is persisted):

Condition Response
Body not JSON / not an object 400
kind ≠ myra.gateway.guardrail-config 400
schema_version ≠ integer 1 400
config not a JSON object 400 { "error": "config must be an object" }
a config key outside the allowlist above 400 { "error": "unknown guardrail-config key: …" }
a detector carrying a url field — in guardrails, legacy detectors, or role_policies[*].guardrails 400
a detector that is not an object, or lacks a non-empty string type 400
more than 64 detectors total across all carriers 400
the guardrail subset would change and confirm is not true 409 with a diff (below)
valid + confirm: true (or an unchanged/no-op re-apply) 200 { "ok": true, "changed": <bool> }

The confirm gate is fail-closed: any change to the guardrail subset requires "confirm": true. An unchanged re-import (the file equals the current guardrail config) is an idempotent no-op and applies without confirm ("changed": false) — so a GitOps pipeline can re-apply the same file safely. When a change needs confirmation the 409 body carries a diff:

{ "error": "guardrail configuration would change; re-send with \"confirm\": true to apply",
  "diff": { "detectors_added": ["…"], "detectors_removed": ["…"],
            "weakens_residency": false, "introduces_deny_all": false } }

weakens_residency is true when the import would lower the effective EU-residency or provider-allowlist enforcement; introduces_deny_all is true when it would enforce a provider allowlist with no usable providers (bricking all inference). These are advisory: the residency floor is non-downgradable regardless — a tenant/deployment residency floor is re-applied on every config load, so a stored value below the floor never takes effect (the same guarantee PATCH has).


Policy dry-run / diff

POST /admin/v1/gateways/<ID>/guardrail-config/dry-run

Evaluate a candidate guardrail policy against sample text and return the behavioral diff versus the gateway's current policy — what would newly block / flag / scrub / pass — WITHOUT applying anything. Requires GATEWAYS_MANAGE. It reuses the real detector pipeline read-only (no upstream model call, no billing), so presidio / prompt_guard / pii_protector detectors in the candidate do execute against the live analyzer.

Request. { "candidate": <guardrail-config object>, "source": "provided" | "last_n", "samples": [ "text", … ] }. candidate is validated with the same rules as an import config (key allowlist, detector url/shape, 64-detector cap). source selects the sample set (default provided; any other value is rejected 400):

  • provided (default) — evaluate the caller-supplied samples: at least one non-empty string, at most 10 evaluated (blank entries skipped).
  • last_n — evaluate the last-N real pre-mask prompts captured for THIS gateway (samples ignored). This requires the tenant's pre-mask-capture opt-in (default OFF — see below); until a tenant opts in, nothing is captured and last_n returns an empty summary with "note": "no_captured_traffic". The read is GATEWAYS_MANAGE-gated to the owning gateway, and the response is verdict-only (the stored prompt text is never returned).

Pre-mask prompt store (last_n source). Stored request_log.prompt is already PII-masked, so re-scanning it is blind to the very PII a candidate policy targets. To dry-run over real traffic, a pii_protector gateway can — only when its tenant enables pre_mask_capture_enabled (per-tenant DB config, default OFF) — persist the raw PRE-mask user prompt on turns where PII was actually masked. This is a tightly-fenced raw-PII store: per-(tenant, gateway) last-N (default 50, pre_mask_capture_cap), a short per-tenant TTL (default 48h, pre_mask_capture_ttl_hours) reaped hourly, residency-tagged, wiped on tenant erasure. Captured content is user-role prompt text only (system/assistant/tool-call args excluded); it is never logged and never leaves the verdict-only dry-run of the owning gateway. It is INCOMPLETE BY DESIGN — only pii_protector request-phase turns with masked user PII are captured.

Response. Verdicts only — the sample text is never echoed back (nor any matched value / pattern / entity), so a sample carrying PII cannot leak through the diff:

{ "summary": { "newly_blocked": 1, "newly_passed": 0, "changed": 0, "unchanged": 1 },
  "samples": [ { "index": 0, "source": "provided", "before": "pass", "after": "block" },
               { "index": 1, "source": "provided", "before": "pass", "after": "pass" } ] }

before / after are each one of block, scrub, flag, pass. A transition to block is counted newly_blocked; block → anything looser is newly_passed; any other change is changed. An empty samples array (or one with no non-empty string) is rejected 400.