NLP PII detector
The NLP PII detector is a Tier 2 (sidecar HTTP call, milliseconds) guardrail that uses a named entity recognition (NER) engine to detect and optionally anonymise personally identifiable information (PII) in request and response bodies. The detection engine runs as a locally hosted sidecar within Myra's certified infrastructure — data is never transmitted outside the Myra perimeter.
NLP PII detector editor
When to use the NLP PII detector
Use the NLP PII detector when you need contextual, NLP-based detection of PII — including unstructured forms such as names, locations, and dates that regex cannot reliably match. For structured data formats with known patterns (card numbers, SSNs), the Regex guardrail (Tier 1) is faster and requires no sidecar call.
How it works
The guardrail extracts the message content from the request, sends each content field to a locally hosted Presidio sidecar, and applies the configured action to any span that meets or exceeds the score_threshold. The sidecar analyses each field using NER models and returns detected entity spans with confidence scores.
What is inspected
On the request phase the detector reads the content fields of the decoded request body, never the JSON request envelope. The following carry text and are inspected:
| Carrier | Notes |
|---|---|
messages[].content |
Plain string content, on every role; a content array of plain strings is also read |
messages[].content[] text blocks |
|
messages[].content[] tool_result blocks |
String content and nested text blocks |
messages[].content[] tool_use blocks |
Every string leaf of input, to a nesting depth of 64 |
system |
The native top-level system prompt — a string or a text-block array |
prompt |
Legacy completions-style |
input |
A plain string, an array of plain strings (batch embeddings), or the Responses-API array of message objects — including a function_call_output's output and a top-level input_text item |
instructions |
The request envelope — role, model, message and tool-call identifiers, block type discriminators, and every other structural value — is never inspected and can never be rewritten. A detector that rewrote a structural value produced a malformed request that the provider rejected with HTTP 400, and analysing the envelope caused the NER model to report PERSON for tokens such as model ids and tenant names on requests containing no personal data at all.
⚠️ Caution — not inspected. The following carry text but are deliberately excluded, so personal data in them is not detected and not masked:
thinkingandredacted_thinkingblocks. Extended-thinking blocks are replayed on later turns together with a cryptographic signature that covers the thinking text. Masking the text would invalidate that signature and the provider would reject the next turn of the conversation.redacted_thinkingis opaque ciphertext.messages[].tool_calls[].function.arguments(the OpenAI-compatible form of a tool call). The value is a JSON document encoded inside a string; masking it as plain text would produce invalid inner JSON, and the tool call would then be forwarded with empty arguments. The native equivalent,tool_use.input, is inspected.messages[].name,user,metadata,tools[]definitions, and the string leaves of block types not listed above.If the only text in a request rides one of the excluded carriers above, nothing is scanned and the request log records
presidio_no_scannable_fieldsso such requests remain countable. On a tenant that mandates PII masking such a request is blocked (pii_unscannable_body) rather than forwarded unscanned; on any other gateway it passes. A request carrying no text at all — an image-only turn, a token count — is ordinary and is never blocked by this rule.
On the response phase the detector walks the same content fields per provider shape — the OpenAI-compatible choices[].message.content (a string or text blocks) and reasoning_content, the native Anthropic content[] text blocks, and an error envelope's message. Structural values (role, finish_reason, tool-call ids, the model name), the model-authored tool-call arguments (tool_calls[].function.arguments / tool_use.input, which the caller re-parses), and thinking signatures are never rewritten, so the provider's own answer keeps its shape. A response body that is not JSON is scanned as one document (there is no structure to corrupt). A response whose JSON shape is not recognised is passed unscanned rather than rewritten, and the request log records presidio_response_shape_unrecognized so it stays countable; a recognised but prose-free body (embeddings, a token count) passes silently. Because the answer is re-encoded from the decoded body when a mask is applied, numeric fields in a masked response may be reformatted (e.g. 1.0 → 1); an unmasked response is delivered byte-for-byte.
💡 Note: Because the masking is now applied to the decoded body, the
promptrecorded in the request log for a scrubbing detector is the masked text. The original is not retained anywhere for this detector type.⚠️ Caution: Running a
presidioscrub detector and a PII Protector on the same gateway is still not recommended. Apresidioscrub now preserves an already-applied PII Protector / custom-PII token — it never anonymises the token bytes, so the value can still be restored — which removes the most common way the two collided. But the combination remains fragile in the reverse order: once PII Protector has restored a value on egress, apresidioscrub running after it anonymises the real value irreversibly, with no token left to map back. Choose one masking detector per gateway.
Block reasons a scrub can produce
A scrubbing detector fails closed rather than reporting a mask that did not happen. These appear as the block reason in the request log:
| Reason | Meaning |
|---|---|
presidio_partial_scan_failed |
Personal data was detected, and the sidecar then failed (or the scan budget ran out) before every field could be masked. The request is refused rather than forwarded with data the detector had already identified. Retried first — a transient sidecar error does not produce this. |
presidio_mask_not_applied |
The sidecar returned a field unchanged although it had reported maskable spans in it, so the mask did not land. Re-sent, already-masked history is recognised and does not trigger this. |
presidio_reencode_failed |
The masked body could not be re-serialised, so the two representations the gateway forwards from would disagree. |
pii_unscannable_body |
The tenant mandates masking and the request's only text rides an excluded carrier. |
What the gateway hands the anonymizer is canonical, whatever the analyzer returned:
exactly four fields per span — entity_type (a string; a null, empty, numeric or
object label becomes PII, a type alias is honoured), integer start / end clamped inside
the field (a fractional offset is widened to the enclosing integers, so a superset is masked,
never less), and a finite score in [0, 1] — with duplicate spans collapsed (same label and
range → one span, the higher score) and whitespace-only spans dropped. Measured against the
sidecar (2026-09-13, replicated): a null, empty or alias-only label answers 422, an object
label 500, a fractional offset 422, an infinite score cannot even be encoded (400), and a
whitespace-only span is accepted but inserts a placeholder between two words
(John<PERSON>Smith); an exact duplicate is accepted in 24 of 25 calls with one sporadic 502
(the transient class the retry already covers). An anonymizer failure on a field with detected
personal data is a refusal, so each deterministic shape used to turn an untrusted analyzer
response into a refused request. The analyzer boundary itself never yields a duplicate span on
any path (single-chunk texts included).
These are request-phase only. On the response phase the model has already generated, so the
same conditions report an error and are resolved by fail_open instead of withholding the answer.
When a scrub cannot be applied
With action: "scrub", masking is all-or-nothing across the whole request. If the sidecar fails on any field, nothing is masked and the detector reports an error, which is then resolved by fail_open; a partially masked request is never forwarded. If the masked body cannot be re-serialised, the request is blocked rather than forwarded, because the two representations the gateway forwards from would otherwise disagree.
The detection engine identifies the language of each request and applies the appropriate NLP model, so no language field is required for detection to work. English and German are fully supported; other Latin-script languages are handled on a best-effort basis. language is nonetheless not inert: it is the calibration control, and it decides the PERSON / LOCATION confidence floor. Leaving it absent is the fail-safe choice — it applies the German 0.6 floor for names and addresses. Setting it to "en" (or any other non-German locale) raises that floor to 0.9, which is correct for English text but means German names and cities scoring 0.6–0.9 — a plain "Klaus Müller" ≈ 0.79, a city ≈ 0.71 — are no longer masked. A request that ran without the floor for this reason records the pii_floor_off_by_language marker in its request log, and the Guardrail Builder asks for confirmation before applying the change. Under a PII mandate the floor is applied whatever language says, so on a mandated request the choice affects nothing but the analyzer's own model selection.
⚠️
action: "scrub"here is DESTRUCTIVE, and it is not the same operation as the PII Protector's. Matched values are replaced with a static label such as<PERSON>or<EMAIL_ADDRESS>, and nothing ever puts them back — not in the model's answer, not in the stored conversation, and the pre-send masking preview cannot recover them (it can only restore values that were tokenised). The model therefore answers about<PERSON>, and that answer is kept. If the answer needs to reference the real values, use the PII Protector guardrail, which tokenises reversibly and restores on the response phase. Both controls are labelled "scrub"; the Guardrail Builder marks this one destructive in its execution plan and says so on the card.
Configuration reference
| Field | Type | Default | Description |
|---|---|---|---|
type |
string | — | Must be "presidio" |
name |
string | — | Human-readable label for this guardrail instance |
action |
string | "scrub" |
What to do when PII is detected: block, scrub, or flag |
target |
string | "request" |
Which phase to inspect: request, response, or both |
entities |
array | null | null |
Entity types to detect; null, a non-array, or an array with no string elements all mean "detect all supported entity types". On a German-calibrated detector with action: "scrub", PERSON and LOCATION are analysed in addition to this list — and on any request under a PII mandate they are analysed whatever language says (a mandate applies the floor at the enforcement layer, so "masking enforced" covers addresses and names even on an English-calibrated detector; such a request records the pii_floor_forced_by_mandate marker so the extra masking is explainable) — see German names and addresses. With action: "block" or "flag" the list binds exactly: only the types you tick are analysed, so unticking PERSON does stop a blocking detector acting on names — except on a request under a PII mandate, where the floor still applies because the gateway is certified as satisfying that mandate. A request is under a mandate when its project is on the pii_mandatory tier, or its tenant has masking enforced, or the project's tier cannot be resolved at all — an unknown tier fails closed. A local-only gateway (every configured provider first-party) exempts the model leg of an enforcing tenant's requests from that mandate — not the model-generated web-search query, which is masked by the pii_protector detector, never by this one |
score_threshold |
number | 0.7 |
Minimum confidence score for a detection to count (0.0–1.0) |
entity_score_thresholds |
object | null | null |
Optional per-entity confidence overrides — a map of entity type to its own threshold (for example { "PERSON": 0.85, "EMAIL_ADDRESS": 0.5 }). An entity listed here uses its own value; any entity not listed falls back to score_threshold. Each value must be a confidence in 0.0–1.0 (an out-of-range or non-numeric value is rejected, exactly as for score_threshold). Lowering the analyzer floor here also lowers the threshold sent to the sidecar so the entity can actually be returned. |
language |
string | null | absent | The NLP calibration for this detector, and the language sent to the analyzer. It decides the PERSON / LOCATION confidence floor: absent, "", "auto", "de" or a de-/de_-prefixed locale ⇒ the German-calibrated 0.6; any other explicit locale ("en", "fr") ⇒ the English 0.9. See German names and addresses. Exposed in the Guardrail Builder as Language calibration on this card. The "" / non-string readings in this row describe what the runtime does with a value already stored; a write that introduces one is refused — see the accepted-shape note below. |
timeout_ms |
integer | 15000 |
Read timeout for the anonymizer sidecar call in milliseconds. Minimum 1000 ms, and a whole number. A lower value cannot complete a real anonymize call, so the detector becomes a guaranteed timeout and, on a fail_open detector, the request is forwarded unchecked. Values outside [1000, 120000], non-numeric values and an explicit null are rejected with 400; a value already stored on a detector is carried through unchanged. The runtime also clamps to [1000, 120000]. Connect timeout is fixed at 500 ms; send timeout at 2000 ms. The analyzer (detection) call does not use this value — large inputs are chunked and each chunk carries a budget-clamped read timeout instead (see Large-document handling). |
fail_open |
boolean | true |
When true, sidecar errors allow the request to pass through; when false, they block it |
💡 Note:
score_thresholdand everyentity_score_thresholdsvalue are confidences between0.0and1.0. A value outside that range, or a non-numeric one, is rejected and the detector falls back to the protective default0.7— it is not clamped to1.0. A too-high threshold (for example5, a typo for0.5) would otherwise match no entity and let personal data reach the model unmasked. The Guardrail Builder blocks out-of-range input and warns when the threshold is set above0.85.⚠️ Accepted / rejected shape —
language. A write that introduces or changeslanguageto anything other than a usable locale string —null, a number, a boolean, an object, whitespace only, or a string carrying a control character — is rejected with400, naming the detector; a value already stored is carried through unchanged so an existing configuration stays saveable. At runtime any non-string value — and any string carrying a control character, which is unusable rather than merely unrecognised — is treated exactly as an absent one and resolves to"auto", which is calibrated as German, so a malformed value masks more, never less. (A real but unknown locale such as"fr"is a deliberate choice and keeps its English calibration; it is passed to the analyzer verbatim.)languageis the only field with a write rule; the other fields the analyzer request is built from (entities,score_threshold) are coerced at runtime only. Note this detector type does not sendallow_list/allow_list_matchat all — that is the PII Protector's. The Guardrail Builder no longer offers an Allow list control on this card, a write that introducesallow_liston apresidiodetector is rejected with400(on create, on update, and on guardrail-config import), and a stored value is dropped the next time the gateway's guardrails are saved from the builder. A value already stored is grandfathered — carried through unchanged — so an existing configuration never becomes unsaveable, and it is stripped from an exported guardrail-config file so the file stays importable onto another gateway. To exempt specific values from masking, use a PII Protector detector, which sends the allow list to the analyzer and honours it.This is not cosmetic. Previously a stored
"language": null— JSONnulldecodes to a truthy value in the gateway runtime, so it survived the "or use the default" fallback — threw while the analyzer request was being built. The guardrail orchestrator recorded that as a detector error, and a detector shipping the defaultfail_open: truethen forwarded the request completely unmasked, with only a log line. A single null disabled the PII layer for every request on the gateway. On a gateway under a PII mandate the same value produced an outage instead of a leak — the turn was refused withpii_mandate_masker_unavailable— which is the correct direction, and the reason the leak went unnoticed on the tenants most likely to report it.
Actions
| Action | Behaviour |
|---|---|
block |
Request is denied if any entity is detected above the confidence threshold. The caller receives a synthetic assistant message. |
scrub |
Detected spans are replaced with <ENTITY_TYPE> placeholders (e.g. <EMAIL_ADDRESS>). The pipeline continues with the redacted body. |
flag |
Entity detections are recorded in the request log. The pipeline continues without modifying the body. |
💡 Note: The entity offsets recorded in the request log are field-local — they index the content field the entity was found in, not the raw request or response body. Rows written before this change indexed the raw body.
💡 Note: With
action: "block"the detector stops at the first content field that produces a detection, so the recorded entity list names what was found in that field rather than everything in the request.
Supported entity types
The NLP PII detector supports 50+ entity types. The following are commonly configured. FP rates are benchmarked at score_threshold: 0.7 across representative general-purpose and business text corpora.
| Entity type | Description | FP risk at 0.7 |
|---|---|---|
EMAIL_ADDRESS |
Email addresses | Low |
PHONE_NUMBER |
Phone numbers | Low |
US_SSN |
US Social Security numbers | Low |
CREDIT_CARD |
Credit card numbers (Luhn-validated) | Low |
US_BANK_NUMBER |
US bank account numbers | Low |
IBAN_CODE |
IBAN bank account codes | Low |
US_PASSPORT |
US passport numbers (regex, US format only) | Low |
PASSPORT |
Passport numbers in any format — detected via NER (multilingual) | Low |
US_DRIVER_LICENSE |
US driver's licence numbers | Low |
US_ITIN |
Individual Taxpayer Identification Numbers | Low |
CRYPTO |
Cryptocurrency wallet addresses | Low |
IP_ADDRESS |
IPv4 and IPv6 addresses | Low |
MEDICAL_LICENSE |
Medical licence numbers | Low |
URL |
Web URLs | Low |
ORG |
Company and organisation names — detected via NER (multilingual) | Medium — threshold auto-raised to 0.85 |
PERSON |
Full or partial person names | High — ~20% FP; threshold 0.6 by default (German), 0.9 for an explicit non-German language. A request that ran at 0.9 for that reason records pii_floor_off_by_language; one that ran at 0.6 because of a PII mandate records pii_floor_forced_by_mandate. |
LOCATION |
Location names | High — ~18% FP; threshold 0.6 by default (German), 0.9 for an explicit non-German language. Same two markers as PERSON. |
DATE_TIME |
Dates and times | High — ~7–14% FP; threshold auto-raised to 0.9 |
💡 Note: For
action: scrub, restrictentitiesto the 14 low-FP types and omitORG,PERSON,LOCATION, andDATE_TIME. (Foraction: blockorflagthe same list is a good starting point and now binds exactly —PERSONandLOCATIONare no longer added behind it, so a blocking detector will not refuse a request on a name. On a request under a PII mandate — apii_mandatoryproject, a masking-enforced tenant (on the model leg, unless the gateway is local-only), or a project whose tier cannot be resolved — the floor still applies, so such a detector will refuse a name there; tickPERSONif you want that behaviour everywhere.) This set produces 0% false positives across benchmarks. The gateway automatically raisesscore_thresholdto 0.85 forORGand to 0.9 forDATE_TIMEwhen they are included.PERSONandLOCATIONuse a 0.6 floor by default (the analyzer language defaults to German); the 0.9 threshold applies to them only when a non-Germanlanguageis configured explicitly.
fail_open behaviour
fail_open |
Sidecar unavailable |
|---|---|
true (default) |
Request passes through as if no PII was found |
false |
Request is blocked |
⚠️ Caution: Set
fail_open: falsewhen PII detection is a hard compliance requirement. Withfail_open: true, a sidecar outage allows all traffic through uninspected.
Tool results (action: "scrub"). A scrub detector also anonymises tool results before they reach the model. At that seam it does not read fail_open: on an analyzer or anonymizer outage the result is forwarded unmasked (logged, and recorded on the request as guardrail_degraded / pii_scan_degraded / guardrail_error) — unless a PII mandate is in force (a pii_mandatory project, or a tenant with masking enforced), where the result is withheld behind the [tool result withheld: PII scan unavailable] placeholder and the request records pii_mandate_masker_unavailable, even when a pii_protector detector follows it. A result bound for a wholly-local model leg is not a mandate case.
Large-document handling
The NLP detector's latency scales with input length, so a large document (for example a multi-hundred-page case file pasted inline into the prompt) cannot be analysed in a single sidecar call without exceeding its read timeout. To keep a legitimate large document from being fail-closed-rejected, the gateway analyses it in pieces:
- Chunking. Each inspected field is split into ~16 KB chunks. Chunks are analysed independently and their entity offsets are merged back with global positions, with a small overlap window so an entity straddling a chunk boundary is still detected.
- Per-chunk retry under one budget. A transient sidecar error on a chunk is retried;
the whole scan is bounded by a single wall-clock budget (default 90 s). If the budget is
exhausted the request fails per the
fail_opensetting — for very large documents that exceed the budget, ingest the document via the knowledge/RAG upload path instead of pasting it inline. - Ill-formed UTF-8 is normalized first. Input containing stray non-UTF-8 bytes (common
in text copied out of PDFs or OCR output) is normalized to valid UTF-8 — each ill-formed
byte becomes the Unicode replacement character
U+FFFD— before chunking. This lets a large document with a few bad bytes be chunked and scanned rather than diverted to a single non-chunked call that would time out, and it keeps detection offsets aligned so redactions land on the right characters. The normalization is internal to PII analysis: a field is only rewritten in the forwarded request when PII is actually redacted in it; a field with no detected PII is forwarded with its original bytes unchanged.
Analyze-result cache (tenant-scoped)
A multi-turn conversation re-sends its whole history every turn, so the same document text would otherwise be analysed once per turn. The gateway memoizes each analyzer result by a collision-resistant hash of the exact fields sent to the sidecar (text, language, entities, score threshold, allow-list) for one hour, so an unchanged document is analysed once and the repeat is served from cache. The cached value is the detected spans only — never the original PII value, which is read from the live text at redaction time.
The cache is per tenant. The hash includes the requesting tenant, so a cached entry is
only ever reused for the tenant that produced it — one tenant's document, template or variable
can never produce a cache hit for another tenant. A tenant's own gateways and conversations
still share (the multi-turn optimisation above is unchanged). This deliberately gives up
cross-tenant reuse of an identical (text, config) (and the small Presidio load it saved): a
shared entry let a hit/miss timing difference confirm that another tenant had processed a
specific document, which a tenant-isolated, EU-sovereign deployment must not expose.
Example configurations
Block any request containing detected PII
{
"type": "presidio",
"name": "block-pii",
"action": "block",
"target": "request",
"fail_open": false
}
Scrub specific financial entity types from both request and response
{
"type": "presidio",
"name": "scrub-financial",
"action": "scrub",
"target": "both",
"entities": ["CREDIT_CARD", "IBAN_CODE", "US_BANK_NUMBER"],
"score_threshold": 0.85
}
Flag PII in responses for audit purposes only
Using the NLP PII detector after a regex pre-filter
Running a regex guardrail first (Tier 1) reduces the volume of content reaching the Presidio sidecar (Tier 2). The example below blocks obvious PCI patterns in-process, then sends everything else to the NLP PII detector for broader PII detection.
[
{
"type": "regex",
"name": "block-pci",
"action": "block",
"target": "request",
"patterns": ["pci_pan"]
},
{
"type": "presidio",
"name": "scrub-remaining-pii",
"action": "scrub",
"target": "request",
"entities": ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "US_SSN"]
}
]
Streaming limitation
⚠️ Caution: When
targetis"response"or"both", response-phase inspection applies only to non-streaming responses. Streamed responses are not buffered by the gateway, so response-phase scrubbing and flagging are skipped for them. Request-phase inspection is unaffected.
Configuring the NLP PII detector
NLP PII detector card — expanded view
Proceed as follows to configure the NLP PII detector in the Guardrail Builder:
- Open the gateway detail page and scroll down to the Guardrails card.
- Click on the + Presidio (NLP) button.
- A collapsed NLP PII detector card appears at the bottom of the list.
- Click on the card to expand it.
- Enter a name in the Name text field.
- Select the action from the Action drop-down list:
block,scrub, orflag. - Select the target from the Target drop-down list:
request,response, orboth. - If required, adjust the Score threshold field.
- If required, select specific entity types from the Entity types list. Leave empty to detect all supported types.
- Toggle the Fail Open switch to
falseif the sidecar must be a hard dependency. - Click on the Save Guardrails button.
ℹ️ There is deliberately no Allow list on this card. This detector type never sends an allow list to the analyzer, so every value entered there was silently ignored — and on an
action: "block"detector the request was still refused on a value the operator had explicitly exempted. Use a PII Protector detector to exempt values from masking.
-> The NLP PII detector is saved and appears in the execution plan.
Pipeline position
The NLP PII detector is Tier 2 — it makes an HTTP call to a sidecar service. All Tier 1 guardrails (regex, keyword) run before any Tier 2 guardrail. Within Tier 2, guardrails run in the order they appear in the guardrails array.
See also
- Guardrail pipeline overview
- Regex guardrail — in-process pattern matching, no sidecar required
- PII Protector — reversible tokenisation that restores original values in the response