Data protection and compliance
💡 Note: This is the guardrails overview — the place to pick the right guardrail for your use case. For the concept (tiers, verdict lifecycle, targets), see Guardrails. To create and manage a pipeline, see Building a guardrail pipeline.
A guardrail is a check applied to every request and response that produces a verdict — block, scrub, or flag. Guardrails run in a guardrail pipeline that inspects content before it reaches the AI provider and before it reaches the user. For the tier model and the full verdict lifecycle, see Guardrails.
This chapter lists the available guardrails, maps common compliance requirements to the right one, and points you to the configuration reference.
Why data protection matters for AI workloads
When users interact with an AI model through a chat interface or API, they frequently include sensitive information in their messages — personal details, financial data, internal reference numbers, or confidential business context. Without controls in place, this content is forwarded verbatim to the AI provider, stored in provider logs, and may be used in model training depending on the provider agreement.
Myra AI Workspace intercepts every request at the gateway level, before it leaves the Myra perimeter, and applies the configured protection rules. No custom code is required in the client application.
Available guardrails
The table below lists every guardrail. Each runs in Tier 1 (in-process, sub-millisecond) or Tier 2 (sidecar, milliseconds), and is either a Detector (it blocks or flags a match) or a Protector (it reversibly masks or tokenises a match and restores the original value later). For the tier model and the verdict lifecycle, see Guardrails.
| Guardrail | Tier | Kind | What it does |
|---|---|---|---|
| Regex / Pattern | 1 | Detector | In-process regex and named pattern matching, with checksum-validated patterns such as pci_pan and iban. |
| Keyword | 1 | Detector | In-process exact keyword matching. |
| Jailbreak | 1 | Detector | Pre-configured jailbreak and prompt-injection detector — zero configuration required. |
| JSON schema | 1 | Detector | Validates model responses against a declared JSON schema — enforces structured output. |
| Contains code | 1 | Detector | Detects source code in requests or responses. |
| Gibberish | 1 | Detector | Detects low-quality or incoherent model responses using entropy and vocabulary heuristics. |
| Language | 1 | Detector | Detects the dominant writing system of request or response text — permits or blocks by script. |
| Custom PII blacklist | 1 | Protector | Reversible masking of admin-defined keywords — client names, codenames, and surnames masked before reaching the model. |
| NLP PII detector (Presidio) | 2 | Detector | NLP-based PII detection — locally hosted within Myra's certified infrastructure. |
| Prompt Guard | 2 | Detector | Safety classification via Llama Guard 3 — locally hosted within Myra's certified infrastructure. |
| PII Protector | 2 | Protector | Reversible PII tokenisation — real values restored in the response. |
Choosing the right protection approach
The table below maps common compliance and security requirements to the recommended guardrails.
| Requirement | Recommended guardrail(s) |
|---|---|
| Prevent personal data (names, emails, phone numbers) from reaching the provider | PII Protector or NLP PII detector with action: "scrub" |
| Protect PII while preserving contextual continuity in the response | PII Protector |
| Block credit card numbers, SSNs, IBANs, and other structured data | Regex guardrail with named patterns (pci_pan, iban, etc.) |
| Block specific terms, product codes, or confidential strings | Keyword guardrail |
| Prevent prompt injection and jailbreak attacks | Jailbreak detector and/or Prompt Guard |
| Enforce structured output format (e.g. JSON) | JSON schema guardrail |
| Restrict interactions to a specific language | Language guardrail |
| Detect and filter incoherent or low-quality model responses | Gibberish detector |
| Prevent source code from appearing in requests or responses | Contains-code detector |
| Audit all requests that contain PII without modifying traffic | NLP PII detector with action: "flag" |
PII protection
Myra AI Workspace provides two approaches to protecting personally identifiable information (PII).
Reversible tokenisation — PII Protector
PII Protector is the recommended approach for interactive use cases such as the Chat view. It detects PII in the user's message, replaces each value with an opaque token before forwarding the request to the provider, and restores the original values in the response. When the request egresses to an external AI provider the model never sees real PII; the user receives a natural, contextually coherent response.
Local models are served unmasked. When a request routes wholly to a first-party local (Myra/EU) model — the resolved model and every failover fallback are Myra-hosted — masking is skipped on both the request and the response: the data never leaves Myra, so masking would only corrupt the user's own text. See PII Protector — Local model legs are not masked. If any leg in the request's egress set is an external provider, masking applies as normal (fail-closed).
⭐ Example: A user types
"My email is alice@example.com". The provider receives"My email is [MYRA-REDACT-EMAIL_ADDRESS:a3f1c2:9b1d4f7a2c3e]". The model's response referencing the token is returned to the user withalice@example.comrestored.
Use PII Protector when:
- The AI model needs to reference PII contextually in its response (for example: summarising a document that mentions people's names).
- The end user should receive the original values in the response. Restoration works on both streaming and non-streaming responses (given
target: "both").
Permanent scrubbing — NLP PII detector
The NLP PII detector (Presidio) detects PII and replaces it with static labels such as <EMAIL_ADDRESS>. The replacement is permanent — the original value is not restored in the response.
Use the NLP PII detector with action: "scrub" when:
- The response does not need to reference the original PII values.
- Compliance requires that PII never appear in the response, even in restored form.
- You need an audit log of all detected PII (use
action: "flag").
Structured data — Regex guardrail
The Regex guardrail is the fastest and most precise tool for structured sensitive data formats — credit card numbers, SSNs, IBANs, routing numbers, passport numbers, and similar values with predictable patterns. It runs in-process (Tier 1) with sub-millisecond latency and supports a built-in library of named patterns with checksum validation.
Use the regex guardrail for structured PCI, PII, or domain-specific data patterns. Combine it with Tier 2 guardrails for comprehensive coverage: the regex guardrail blocks structured data first, then the NLP PII detector handles unstructured forms such as names and addresses.
File attachments on a PII-protected gateway
PII guardrails scan request content as text. They cannot see personal data inside a file sent as raw binary — a PDF or image embedded directly in the request as base64 bytes egresses to the AI provider without being masked. To close that gap, a gateway that runs an active PII protector (PII Protector, Custom PII blacklist, or the NLP PII detector with action scrub/block) rejects an inference request that carries inline base64 file media, with the block reason pii_media_unmaskable.
- What is rejected: inline base64 image, PDF, or audio in the request body — an OpenAI-style
image_url/video_urldata:URL, aninput_audio/file(file_data) block, or an Anthropic-nativeimage/documentblock with a base64source(including base64 nested inside atool_resultor adocumentsourceof typecontent). - What is accepted: text that has already been extracted from a file, and file references (a Files-API
file_id). Attach files through the Chat view: on a PII-protected gateway the gateway extracts each PDF/image to text server-side first, so the extracted text passes through PII masking before it leaves the gateway. Remote image URLs (http(s)://…, notdata:) are not inline binary and are not affected by this check. - Limitation: the check treats a gateway as PII-protected when it runs one of the guardrails above. A gateway that masks personal data only through a Regex or Keyword guardrail is not covered — add a PII Protector, Custom PII, or NLP PII detector to enable the attachment check.
- Context compaction strips, it does not reject: when a long conversation is summarised (context compaction, either leg), any inline binary in the turns being compacted is removed from the summarisation prompt and replaced with a short
[binary omitted: N bytes]marker before that prompt leaves the gateway — the surrounding text, the block type and the media type are preserved, so the summary still records that an attachment was there. Rejecting instead would mean a conversation that ever carried an attachment could never be compacted again, and raw base64 carries nothing a summary could use. The removal is verified, not assumed: if any payload survives the strip, the gateway refuses to send the summarisation prompt at all rather than forwarding it, and the compaction is declined. - Local model legs are exempt: when a request routes wholly to a first-party local (Myra/EU) model, this base64-media block does not apply — there is no third-party egress to fail closed for, so the attachment reaches the local model as-is (like the rest of the local-leg unmask above).
This block is a guardrail block: in streaming mode it returns HTTP 200 with the refusal delivered over the stream (see Error codes — guardrail_blocked).
Code interpreter and attached data files. When the code interpreter analyses an attached spreadsheet, the raw file bytes are staged into the isolated sandbox — they never enter the request that egresses to an AI provider. On a PII-protected gateway this staging is permitted only when the gateway runs a PII detector that masks the sandbox result fail-closed before it leaves the gateway; a gateway whose PII configuration cannot guarantee that (for example a fail-open detector — the staging decision reads the configuration alone, so a detector that is fail-open outside a PII mandate does not qualify even though a mandate would make it withhold — or only a Custom PII blacklist) refuses staging, and the assistant answers from the extracted, masked text instead.
Model-generated files show clear names, like the chat. When the assistant writes a text file for you to download — a write_file document (.md, .txt, .csv, .html) or a text artifact produced by the code interpreter — the file is unmasked for you exactly as the on-screen chat is: the reversible PII Protector tokens (names, e-mail addresses, phone numbers, and the other detected types) are restored to their clear values before the file is stored or delivered, so the download matches what you see in the conversation rather than showing redaction placeholders. This is the same selective unmask the chat reply already applies, and the file's stored text holds the clear values at rest exactly as the conversation's own messages do. Two deliberate limits:
- Custom PII blacklist keywords stay masked in the file (for now). A confidential keyword you configured on the Custom PII blacklist is shown in the file as a neutral
[redacted]placeholder, not its clear value — even though the chat reply unmasks it. This is intentional and fail-closed. The read-back path — a stored file re-opened via a follow-up, or retrieval over project knowledge — now does re-mask custom keywords before the content egresses to the AI provider (the tool-result egress seam re-applies the keyword tokenizer), so the leak this limit originally guarded against is closed. Restoring custom keywords to their clear value in the generated file (to match the chat) is a separate follow-up not yet shipped; until it lands, files keep custom keywords masked. The reversible PII Protector types are already re-masked on read-back and so are restored in the file. - Binary documents produced directly by the code interpreter are not rewritten. A
.docx/.pdf/.xlsx(and similar) the code interpreter generates as raw binary bytes is delivered as-is; a redaction placeholder embedded inside such a container is not restored, because rewriting bytes inside a compressed/structured document would corrupt it. This does not affect awrite_filedocument the gateway renders to.docx/.pdfat download time — that is rendered from the already-unmasked text above, so it carries the clear values. Only an artifact the model emitted directly as binary is affected.
Prompt injection and jailbreak protection
Jailbreak detector
The Jailbreak detector is a zero-configuration Tier 1 guardrail pre-loaded with 18 detection phrases drawn from known prompt injection and jailbreak attacks. It requires no configuration beyond adding it to the gateway.
Prompt Guard
Prompt Guard is a Tier 2 semantic safety classifier based on Meta's Llama Guard 3, running locally within the Myra perimeter. It classifies requests against a configurable set of harm categories and is more robust against novel or paraphrased attack patterns than the keyword-based Jailbreak detector.
For maximum coverage, use both guardrails: the Jailbreak detector (Tier 1) blocks known patterns instantly, and Prompt Guard (Tier 2) catches semantic variants.
Output policy enforcement
The following guardrails inspect model responses rather than — or in addition to — user requests.
| Guardrail | What it enforces |
|---|---|
| JSON schema | The model response must match a declared JSON schema. Non-conforming responses are blocked. |
| Language restriction | Requests or responses must use a specific writing system (Latin, Cyrillic, Arabic, etc.). |
| Gibberish detection | Low-quality or incoherent model responses are blocked or flagged before reaching the user. |
| Contains-code detection | Source code in requests or responses is detected and the configured action applied. |
Building a guardrail pipeline
A production data protection configuration typically combines multiple guardrails. A recommended starting point for a general-purpose gateway handling user data:
[
{
"type": "jailbreak",
"name": "block-jailbreaks",
"action": "block",
"target": "request"
},
{
"type": "regex",
"name": "block-structured-pii",
"action": "block",
"target": "request",
"patterns": ["pci_pan", "iban", "ssn", "routing_number"]
},
{
"type": "pii_protector",
"name": "tokenize-pii",
"target": "both",
"entities": ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "LOCATION"],
"fail_open": false
}
]
This pipeline:
- Blocks known jailbreak attempts instantly (Tier 1, no latency).
- Blocks structured financial and identity data before it leaves the network (Tier 1, no latency).
- Tokenises remaining personal data contextually and restores original values in the response (Tier 2, milliseconds).
For the full reference on how guardrails are configured, ordered, and managed, see Building a guardrail pipeline.
AI transparency labeling (EU AI Act Art. 50)
Myra AI Workspace makes AI-generated content recognizable as such and keeps users aware that they are interacting with an AI system, in line with the transparency obligations of Article 50 of the EU AI Act.
Users know they are interacting with an AI. A generative-AI disclosure is shown once after sign-in, and the product is presented throughout as "MYRA AI"; combined with the per-message marking below, a user is informed that responses come from an AI system.
AI-generated content is marked in the interface. Every assistant reply in the chat surface carries a visible AI-generated marker (German: KI-generiert). The marker is shown on the user's own chat, on the public shared-conversation view, and on streaming replies as they arrive. It is not shown on the user's own messages.
AI-generated content is marked in exports. When an AI reply, an AI-written artifact, or a conversation transcript is exported as a document, the exported file carries a footer notice stating that it contains AI-generated content. This applies to the prose document formats — PDF, Word (DOCX), OpenDocument text (ODT), plain text (TXT), PowerPoint (PPTX), OpenDocument presentation (ODP) — and to the Markdown (.md) download. The notice is localized to the user's interface language (English or German) and is worded to say the document contains AI-generated content, so a transcript that also includes the user's own prompts is not falsely labeled as entirely AI-authored.
Scope of the export marker. The marker is applied only to AI-content exports; it is deliberately not added to:
- User-uploaded file exports (for example, downloading a knowledge file from a project) — these are the user's own documents, not AI output, and must not be mislabeled.
- Clipboard copies — a copy-to-clipboard action is not an exported document, and injecting a notice into every copied snippet would break normal paste use.
- Raw original downloads of an AI-written artifact (the unmodified source bytes of a
.py,.csv, etc.) — a notice cannot be inserted into arbitrary file bytes without corrupting them; the exported document forms (PDF/DOCX/ODT/TXT/MD) carry the marker instead. - Data-only tabular exports (CSV/XLSX/ODS) — these contain extracted table data with no prose surface for a notice, so the marker is not appended to them.
Generated images fall under the separate synthetic-media provisions of Article 50(2) and are outside the scope of this text-content marking. The administrator model-testing Playground is an internal developer tool rather than an end-user conversational surface, and is likewise not in scope for the per-message marker.
Viewing guardrail events
Every guardrail verdict is recorded in the request log. To view guardrail activity, open Settings → Request Logs and filter by Blocked status or search for a specific guardrail name.
The following fields are available in each log entry:
| Field | Description |
|---|---|
blocked |
Whether the request was blocked by any guardrail |
blocked_by |
The name of the guardrail that issued the block |
block_reason |
The pattern or category that triggered the verdict |
detectors_fired |
All guardrails that produced a non-pass verdict |
The Dashboard also shows recent guardrail events in the Recent guardrail events card. The Cost Analytics view (Overview → Analytics → Cost Analytics) breaks down blocked requests by gateway and time period.
See also
- Guardrails — the concept: tiers, verdict lifecycle, and targets
- Building a guardrail pipeline — full configuration reference, verdict behaviour, and pipeline management
- PII Protector — reversible PII tokenisation
- NLP PII detector — permanent scrubbing via NER
- Regex guardrail — structured data pattern matching
- Keyword guardrail — exact string matching
- Jailbreak detector — zero-config prompt injection protection
- Prompt Guard — semantic safety classification
- Request logs — viewing guardrail events