Skip to content

NLP PII detector

The NLP PII detector is a Tier 2 (sidecar HTTP call, milliseconds) guardrail that uses a named entity recognition (NER) engine to detect and optionally anonymise personally identifiable information (PII) in request and response bodies. The detection engine runs as a locally hosted sidecar within Myra's certified infrastructure — data is never transmitted outside the Myra perimeter.

Screenshot: Presidio (NLP) guardrail editor in the Guardrail Builder NLP PII detector editor

When to use the NLP PII detector

Use the NLP PII detector when you need contextual, NLP-based detection of PII — including unstructured forms such as names, locations, and dates that regex cannot reliably match. For structured data formats with known patterns (card numbers, SSNs), the Regex guardrail (Tier 1) is faster and requires no sidecar call.

How it works

The guardrail extracts the message content from the request, sends each content field to a locally hosted Presidio sidecar, and applies the configured action to any span that meets or exceeds the score_threshold. The sidecar analyses each field using NER models and returns detected entity spans with confidence scores.

What is inspected

On the request phase the detector reads the content fields of the decoded request body, never the JSON request envelope. The following carry text and are inspected:

Carrier Notes
messages[].content Plain string content, on every role; a content array of plain strings is also read
messages[].content[] text blocks
messages[].content[] tool_result blocks String content and nested text blocks
messages[].content[] tool_use blocks Every string leaf of input, to a nesting depth of 64
system The native top-level system prompt — a string or a text-block array
prompt Legacy completions-style
input A plain string, an array of plain strings (batch embeddings), or the Responses-API array of message objects — including a function_call_output's output and a top-level input_text item
instructions

The request envelope — role, model, message and tool-call identifiers, block type discriminators, and every other structural value — is never inspected and can never be rewritten. A detector that rewrote a structural value produced a malformed request that the provider rejected with HTTP 400, and analysing the envelope caused the NER model to report PERSON for tokens such as model ids and tenant names on requests containing no personal data at all.

⚠️ Caution — not inspected. The following carry text but are deliberately excluded, so personal data in them is not detected and not masked:

  • thinking and redacted_thinking blocks. Extended-thinking blocks are replayed on later turns together with a cryptographic signature that covers the thinking text. Masking the text would invalidate that signature and the provider would reject the next turn of the conversation. redacted_thinking is opaque ciphertext.
  • messages[].tool_calls[].function.arguments (the OpenAI-compatible form of a tool call). The value is a JSON document encoded inside a string; masking it as plain text would produce invalid inner JSON, and the tool call would then be forwarded with empty arguments. The native equivalent, tool_use.input, is inspected.
  • messages[].name, user, metadata, tools[] definitions, and the string leaves of block types not listed above.

If the only text in a request rides one of the excluded carriers above, nothing is scanned and the request log records presidio_no_scannable_fields so such requests remain countable. On a tenant that mandates PII masking such a request is blocked (pii_unscannable_body) rather than forwarded unscanned; on any other gateway it passes. A request carrying no text at all — an image-only turn, a token count — is ordinary and is never blocked by this rule.

On the response phase the detector walks the same content fields per provider shape — the OpenAI-compatible choices[].message.content (a string or text blocks) and reasoning_content, the native Anthropic content[] text blocks, and an error envelope's message. Structural values (role, finish_reason, tool-call ids, the model name), the model-authored tool-call arguments (tool_calls[].function.arguments / tool_use.input, which the caller re-parses), and thinking signatures are never rewritten, so the provider's own answer keeps its shape. A response body that is not JSON is scanned as one document (there is no structure to corrupt). A response whose JSON shape is not recognised is passed unscanned rather than rewritten, and the request log records presidio_response_shape_unrecognized so it stays countable; a recognised but prose-free body (embeddings, a token count) passes silently. Because the answer is re-encoded from the decoded body when a mask is applied, numeric fields in a masked response may be reformatted (e.g. 1.0 → 1); an unmasked response is delivered byte-for-byte.

💡 Note: Because the masking is now applied to the decoded body, the prompt recorded in the request log for a scrubbing detector is the masked text. The original is not retained anywhere for this detector type.

⚠️ Caution: Running a presidio scrub detector and a PII Protector on the same gateway is still not recommended. A presidio scrub now preserves an already-applied PII Protector / custom-PII token — it never anonymises the token bytes, so the value can still be restored — which removes the most common way the two collided. But the combination remains fragile in the reverse order: once PII Protector has restored a value on egress, a presidio scrub running after it anonymises the real value irreversibly, with no token left to map back. Choose one masking detector per gateway.

Block reasons a scrub can produce

A scrubbing detector fails closed rather than reporting a mask that did not happen. These appear as the block reason in the request log:

Reason Meaning
presidio_partial_scan_failed Personal data was detected, and the sidecar then failed (or the scan budget ran out) before every field could be masked. The request is refused rather than forwarded with data the detector had already identified. Retried first — a transient sidecar error does not produce this.
presidio_mask_not_applied The sidecar returned a field unchanged although it had reported maskable spans in it, so the mask did not land. Re-sent, already-masked history is recognised and does not trigger this.
presidio_reencode_failed The masked body could not be re-serialised, so the two representations the gateway forwards from would disagree.
pii_unscannable_body The tenant mandates masking and the request's only text rides an excluded carrier.

What the gateway hands the anonymizer is canonical, whatever the analyzer returned: exactly four fields per span — entity_type (a string; a null, empty, numeric or object label becomes PII, a type alias is honoured), integer start / end clamped inside the field (a fractional offset is widened to the enclosing integers, so a superset is masked, never less), and a finite score in [0, 1] — with duplicate spans collapsed (same label and range → one span, the higher score) and whitespace-only spans dropped. Measured against the sidecar (2026-09-13, replicated): a null, empty or alias-only label answers 422, an object label 500, a fractional offset 422, an infinite score cannot even be encoded (400), and a whitespace-only span is accepted but inserts a placeholder between two words (John<PERSON>Smith); an exact duplicate is accepted in 24 of 25 calls with one sporadic 502 (the transient class the retry already covers). An anonymizer failure on a field with detected personal data is a refusal, so each deterministic shape used to turn an untrusted analyzer response into a refused request. The analyzer boundary itself never yields a duplicate span on any path (single-chunk texts included).

These are request-phase only. On the response phase the model has already generated, so the same conditions report an error and are resolved by fail_open instead of withholding the answer.

When a scrub cannot be applied

With action: "scrub", masking is all-or-nothing across the whole request. If the sidecar fails on any field, nothing is masked and the detector reports an error, which is then resolved by fail_open; a partially masked request is never forwarded. If the masked body cannot be re-serialised, the request is blocked rather than forwarded, because the two representations the gateway forwards from would otherwise disagree.

The detection engine identifies the language of each request and applies the appropriate NLP model, so no language field is required for detection to work. English and German are fully supported; other Latin-script languages are handled on a best-effort basis. language is nonetheless not inert: it is the calibration control, and it decides the PERSON / LOCATION confidence floor. Leaving it absent is the fail-safe choice — it applies the German 0.6 floor for names and addresses. Setting it to "en" (or any other non-German locale) raises that floor to 0.9, which is correct for English text but means German names and cities scoring 0.6–0.9 — a plain "Klaus Müller" ≈ 0.79, a city ≈ 0.71 — are no longer masked. A request that ran without the floor for this reason records the pii_floor_off_by_language marker in its request log, and the Guardrail Builder asks for confirmation before applying the change. Under a PII mandate the floor is applied whatever language says, so on a mandated request the choice affects nothing but the analyzer's own model selection.

⚠️ action: "scrub" here is DESTRUCTIVE, and it is not the same operation as the PII Protector's. Matched values are replaced with a static label such as <PERSON> or <EMAIL_ADDRESS>, and nothing ever puts them back — not in the model's answer, not in the stored conversation, and the pre-send masking preview cannot recover them (it can only restore values that were tokenised). The model therefore answers about <PERSON>, and that answer is kept. If the answer needs to reference the real values, use the PII Protector guardrail, which tokenises reversibly and restores on the response phase. Both controls are labelled "scrub"; the Guardrail Builder marks this one destructive in its execution plan and says so on the card.


Configuration reference

Field Type Default Description
type string — Must be "presidio"
name string — Human-readable label for this guardrail instance
action string "scrub" What to do when PII is detected: block, scrub, or flag
target string "request" Which phase to inspect: request, response, or both
entities array | null null Entity types to detect; null, a non-array, or an array with no string elements all mean "detect all supported entity types". On a German-calibrated detector with action: "scrub", PERSON and LOCATION are analysed in addition to this list — and on any request under a PII mandate they are analysed whatever language says (a mandate applies the floor at the enforcement layer, so "masking enforced" covers addresses and names even on an English-calibrated detector; such a request records the pii_floor_forced_by_mandate marker so the extra masking is explainable) — see German names and addresses. With action: "block" or "flag" the list binds exactly: only the types you tick are analysed, so unticking PERSON does stop a blocking detector acting on names — except on a request under a PII mandate, where the floor still applies because the gateway is certified as satisfying that mandate. A request is under a mandate when its project is on the pii_mandatory tier, or its tenant has masking enforced, or the project's tier cannot be resolved at all — an unknown tier fails closed. A local-only gateway (every configured provider first-party) exempts the model leg of an enforcing tenant's requests from that mandate — not the model-generated web-search query, which is masked by the pii_protector detector, never by this one
score_threshold number 0.7 Minimum confidence score for a detection to count (0.0–1.0)
entity_score_thresholds object | null null Optional per-entity confidence overrides — a map of entity type to its own threshold (for example { "PERSON": 0.85, "EMAIL_ADDRESS": 0.5 }). An entity listed here uses its own value; any entity not listed falls back to score_threshold. Each value must be a confidence in 0.0–1.0 (an out-of-range or non-numeric value is rejected, exactly as for score_threshold). Lowering the analyzer floor here also lowers the threshold sent to the sidecar so the entity can actually be returned.
language string | null absent The NLP calibration for this detector, and the language sent to the analyzer. It decides the PERSON / LOCATION confidence floor: absent, "", "auto", "de" or a de-/de_-prefixed locale ⇒ the German-calibrated 0.6; any other explicit locale ("en", "fr") ⇒ the English 0.9. See German names and addresses. Exposed in the Guardrail Builder as Language calibration on this card. The "" / non-string readings in this row describe what the runtime does with a value already stored; a write that introduces one is refused — see the accepted-shape note below.
timeout_ms integer 15000 Read timeout for the anonymizer sidecar call in milliseconds. Minimum 1000 ms, and a whole number. A lower value cannot complete a real anonymize call, so the detector becomes a guaranteed timeout and, on a fail_open detector, the request is forwarded unchecked. Values outside [1000, 120000], non-numeric values and an explicit null are rejected with 400; a value already stored on a detector is carried through unchanged. The runtime also clamps to [1000, 120000]. Connect timeout is fixed at 500 ms; send timeout at 2000 ms. The analyzer (detection) call does not use this value — large inputs are chunked and each chunk carries a budget-clamped read timeout instead (see Large-document handling).
fail_open boolean true When true, sidecar errors allow the request to pass through; when false, they block it

💡 Note: score_threshold and every entity_score_thresholds value are confidences between 0.0 and 1.0. A value outside that range, or a non-numeric one, is rejected and the detector falls back to the protective default 0.7 — it is not clamped to 1.0. A too-high threshold (for example 5, a typo for 0.5) would otherwise match no entity and let personal data reach the model unmasked. The Guardrail Builder blocks out-of-range input and warns when the threshold is set above 0.85.

⚠️ Accepted / rejected shape — language. A write that introduces or changes language to anything other than a usable locale string — null, a number, a boolean, an object, whitespace only, or a string carrying a control character — is rejected with 400, naming the detector; a value already stored is carried through unchanged so an existing configuration stays saveable. At runtime any non-string value — and any string carrying a control character, which is unusable rather than merely unrecognised — is treated exactly as an absent one and resolves to "auto", which is calibrated as German, so a malformed value masks more, never less. (A real but unknown locale such as "fr" is a deliberate choice and keeps its English calibration; it is passed to the analyzer verbatim.) language is the only field with a write rule; the other fields the analyzer request is built from (entities, score_threshold) are coerced at runtime only. Note this detector type does not send allow_list / allow_list_match at all — that is the PII Protector's. The Guardrail Builder no longer offers an Allow list control on this card, a write that introduces allow_list on a presidio detector is rejected with 400 (on create, on update, and on guardrail-config import), and a stored value is dropped the next time the gateway's guardrails are saved from the builder. A value already stored is grandfathered — carried through unchanged — so an existing configuration never becomes unsaveable, and it is stripped from an exported guardrail-config file so the file stays importable onto another gateway. To exempt specific values from masking, use a PII Protector detector, which sends the allow list to the analyzer and honours it.

This is not cosmetic. Previously a stored "language": null — JSON null decodes to a truthy value in the gateway runtime, so it survived the "or use the default" fallback — threw while the analyzer request was being built. The guardrail orchestrator recorded that as a detector error, and a detector shipping the default fail_open: true then forwarded the request completely unmasked, with only a log line. A single null disabled the PII layer for every request on the gateway. On a gateway under a PII mandate the same value produced an outage instead of a leak — the turn was refused with pii_mandate_masker_unavailable — which is the correct direction, and the reason the leak went unnoticed on the tenants most likely to report it.


Actions

Action Behaviour
block Request is denied if any entity is detected above the confidence threshold. The caller receives a synthetic assistant message.
scrub Detected spans are replaced with <ENTITY_TYPE> placeholders (e.g. <EMAIL_ADDRESS>). The pipeline continues with the redacted body.
flag Entity detections are recorded in the request log. The pipeline continues without modifying the body.

💡 Note: The entity offsets recorded in the request log are field-local — they index the content field the entity was found in, not the raw request or response body. Rows written before this change indexed the raw body.

💡 Note: With action: "block" the detector stops at the first content field that produces a detection, so the recorded entity list names what was found in that field rather than everything in the request.


Supported entity types

The NLP PII detector supports 50+ entity types. The following are commonly configured. FP rates are benchmarked at score_threshold: 0.7 across representative general-purpose and business text corpora.

Entity type Description FP risk at 0.7
EMAIL_ADDRESS Email addresses Low
PHONE_NUMBER Phone numbers Low
US_SSN US Social Security numbers Low
CREDIT_CARD Credit card numbers (Luhn-validated) Low
US_BANK_NUMBER US bank account numbers Low
IBAN_CODE IBAN bank account codes Low
US_PASSPORT US passport numbers (regex, US format only) Low
PASSPORT Passport numbers in any format — detected via NER (multilingual) Low
US_DRIVER_LICENSE US driver's licence numbers Low
US_ITIN Individual Taxpayer Identification Numbers Low
CRYPTO Cryptocurrency wallet addresses Low
IP_ADDRESS IPv4 and IPv6 addresses Low
MEDICAL_LICENSE Medical licence numbers Low
URL Web URLs Low
ORG Company and organisation names — detected via NER (multilingual) Medium — threshold auto-raised to 0.85
PERSON Full or partial person names High — ~20% FP; threshold 0.6 by default (German), 0.9 for an explicit non-German language. A request that ran at 0.9 for that reason records pii_floor_off_by_language; one that ran at 0.6 because of a PII mandate records pii_floor_forced_by_mandate.
LOCATION Location names High — ~18% FP; threshold 0.6 by default (German), 0.9 for an explicit non-German language. Same two markers as PERSON.
DATE_TIME Dates and times High — ~7–14% FP; threshold auto-raised to 0.9

💡 Note: For action: scrub, restrict entities to the 14 low-FP types and omit ORG, PERSON, LOCATION, and DATE_TIME. (For action: block or flag the same list is a good starting point and now binds exactly — PERSON and LOCATION are no longer added behind it, so a blocking detector will not refuse a request on a name. On a request under a PII mandate — a pii_mandatory project, a masking-enforced tenant (on the model leg, unless the gateway is local-only), or a project whose tier cannot be resolved — the floor still applies, so such a detector will refuse a name there; tick PERSON if you want that behaviour everywhere.) This set produces 0% false positives across benchmarks. The gateway automatically raises score_threshold to 0.85 for ORG and to 0.9 for DATE_TIME when they are included. PERSON and LOCATION use a 0.6 floor by default (the analyzer language defaults to German); the 0.9 threshold applies to them only when a non-German language is configured explicitly.


fail_open behaviour

fail_open Sidecar unavailable
true (default) Request passes through as if no PII was found
false Request is blocked

⚠️ Caution: Set fail_open: false when PII detection is a hard compliance requirement. With fail_open: true, a sidecar outage allows all traffic through uninspected.

Tool results (action: "scrub"). A scrub detector also anonymises tool results before they reach the model. At that seam it does not read fail_open: on an analyzer or anonymizer outage the result is forwarded unmasked (logged, and recorded on the request as guardrail_degraded / pii_scan_degraded / guardrail_error) — unless a PII mandate is in force (a pii_mandatory project, or a tenant with masking enforced), where the result is withheld behind the [tool result withheld: PII scan unavailable] placeholder and the request records pii_mandate_masker_unavailable, even when a pii_protector detector follows it. A result bound for a wholly-local model leg is not a mandate case.


Large-document handling

The NLP detector's latency scales with input length, so a large document (for example a multi-hundred-page case file pasted inline into the prompt) cannot be analysed in a single sidecar call without exceeding its read timeout. To keep a legitimate large document from being fail-closed-rejected, the gateway analyses it in pieces:

  • Chunking. Each inspected field is split into ~16 KB chunks. Chunks are analysed independently and their entity offsets are merged back with global positions, with a small overlap window so an entity straddling a chunk boundary is still detected.
  • Per-chunk retry under one budget. A transient sidecar error on a chunk is retried; the whole scan is bounded by a single wall-clock budget (default 90 s). If the budget is exhausted the request fails per the fail_open setting — for very large documents that exceed the budget, ingest the document via the knowledge/RAG upload path instead of pasting it inline.
  • Ill-formed UTF-8 is normalized first. Input containing stray non-UTF-8 bytes (common in text copied out of PDFs or OCR output) is normalized to valid UTF-8 — each ill-formed byte becomes the Unicode replacement character U+FFFD — before chunking. This lets a large document with a few bad bytes be chunked and scanned rather than diverted to a single non-chunked call that would time out, and it keeps detection offsets aligned so redactions land on the right characters. The normalization is internal to PII analysis: a field is only rewritten in the forwarded request when PII is actually redacted in it; a field with no detected PII is forwarded with its original bytes unchanged.

Analyze-result cache (tenant-scoped)

A multi-turn conversation re-sends its whole history every turn, so the same document text would otherwise be analysed once per turn. The gateway memoizes each analyzer result by a collision-resistant hash of the exact fields sent to the sidecar (text, language, entities, score threshold, allow-list) for one hour, so an unchanged document is analysed once and the repeat is served from cache. The cached value is the detected spans only — never the original PII value, which is read from the live text at redaction time.

The cache is per tenant. The hash includes the requesting tenant, so a cached entry is only ever reused for the tenant that produced it — one tenant's document, template or variable can never produce a cache hit for another tenant. A tenant's own gateways and conversations still share (the multi-turn optimisation above is unchanged). This deliberately gives up cross-tenant reuse of an identical (text, config) (and the small Presidio load it saved): a shared entry let a hit/miss timing difference confirm that another tenant had processed a specific document, which a tenant-isolated, EU-sovereign deployment must not expose.


Example configurations

Block any request containing detected PII

{
  "type": "presidio",
  "name": "block-pii",
  "action": "block",
  "target": "request",
  "fail_open": false
}

Scrub specific financial entity types from both request and response

{
  "type": "presidio",
  "name": "scrub-financial",
  "action": "scrub",
  "target": "both",
  "entities": ["CREDIT_CARD", "IBAN_CODE", "US_BANK_NUMBER"],
  "score_threshold": 0.85
}

Flag PII in responses for audit purposes only

{
  "type": "presidio",
  "name": "flag-pii-responses",
  "action": "flag",
  "target": "response"
}

Using the NLP PII detector after a regex pre-filter

Running a regex guardrail first (Tier 1) reduces the volume of content reaching the Presidio sidecar (Tier 2). The example below blocks obvious PCI patterns in-process, then sends everything else to the NLP PII detector for broader PII detection.

[
  {
    "type": "regex",
    "name": "block-pci",
    "action": "block",
    "target": "request",
    "patterns": ["pci_pan"]
  },
  {
    "type": "presidio",
    "name": "scrub-remaining-pii",
    "action": "scrub",
    "target": "request",
    "entities": ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "US_SSN"]
  }
]

Streaming limitation

⚠️ Caution: When target is "response" or "both", response-phase inspection applies only to non-streaming responses. Streamed responses are not buffered by the gateway, so response-phase scrubbing and flagging are skipped for them. Request-phase inspection is unaffected.


Configuring the NLP PII detector

Screenshot: NLP PII detector card in the Guardrail Builder NLP PII detector card — expanded view

Proceed as follows to configure the NLP PII detector in the Guardrail Builder:

  1. Open the gateway detail page and scroll down to the Guardrails card.
  2. Click on the + Presidio (NLP) button.
  3. A collapsed NLP PII detector card appears at the bottom of the list.
  4. Click on the card to expand it.
  5. Enter a name in the Name text field.
  6. Select the action from the Action drop-down list: block, scrub, or flag.
  7. Select the target from the Target drop-down list: request, response, or both.
  8. If required, adjust the Score threshold field.
  9. If required, select specific entity types from the Entity types list. Leave empty to detect all supported types.
  10. Toggle the Fail Open switch to false if the sidecar must be a hard dependency.
  11. Click on the Save Guardrails button.

ℹ️ There is deliberately no Allow list on this card. This detector type never sends an allow list to the analyzer, so every value entered there was silently ignored — and on an action: "block" detector the request was still refused on a value the operator had explicitly exempted. Use a PII Protector detector to exempt values from masking.

-> The NLP PII detector is saved and appears in the execution plan.


Pipeline position

The NLP PII detector is Tier 2 — it makes an HTTP call to a sidecar service. All Tier 1 guardrails (regex, keyword) run before any Tier 2 guardrail. Within Tier 2, guardrails run in the order they appear in the guardrails array.


See also