Skip to content

Prompt Guard

The Prompt Guard guardrail is a Tier 2 (sidecar HTTP call, milliseconds) guardrail that uses Meta's Llama Guard 3 model to classify request and response content against 14 safety categories. It detects harmful, illegal, or policy-violating content that rule-based guardrails cannot cover. Llama Guard 3 runs as a locally hosted model within Myra's certified infrastructure — prompt content is never transmitted outside the Myra perimeter.

Screenshot: Prompt Guard editor in the Guardrail Builder Prompt Guard editor

When to use Prompt Guard

Use Prompt Guard when you need semantic safety classification — detecting harmful intent expressed in natural language, rephrased attacks, or policy violations that literal keyword matching cannot catch. For structured data detection (PII, card numbers), use the Regex guardrail or NLP PII Detector.

How it works

The guardrail sends the inspected content to the locally hosted Llama Guard 3 sidecar. The model classifies the content against the configured safety categories and returns a verdict. Request-phase classification evaluates only the most recent user message — the full conversation history is not sent to the classifier.

⚠️ Caution: Prompt Guard does not support action: "scrub". Configure block or flag only.


Configuration reference

Field Type Default Description
type string — Must be "prompt_guard"
name string — Human-readable label for this guardrail instance
action string "block" What to do on a safety violation: block or flag
target string "request" Which phase to classify: request, response, or both
timeout_ms integer 2000 Timeout for the sidecar call in milliseconds. Minimum 1000 ms, and a whole number — a value below that cannot complete a real call, so the detector becomes a guaranteed timeout and, on a fail_open detector, the request is forwarded unchecked. Values outside [1000, 120000], non-numeric values and an explicit null are rejected with 400; a value already stored on a detector is carried through unchanged. The runtime also clamps to [1000, 120000], so an out-of-range value that predates this rule is corrected rather than obeyed.
fail_open boolean true When true, sidecar errors allow the request to pass through; when false, they block it
categories array | null null Safety categories to enforce; null enforces all 14 categories
context_prompt string null Deployment context prepended to each user message before classification — reduces false positives on professional platforms (see Context injection)

💡 Note: The classifier endpoint is platform-managed — the Llama Guard 3 sidecar is resolved by the platform, not per gateway. A detector may not carry a url field: the gateway-config write endpoint rejects any guardrail detector that includes one (HTTP 400). This is a deliberate security boundary — a tenant-set endpoint would let a gateway be pointed at an arbitrary internal address.


Safety categories

Code Category FP risk for block
S1 Violent Crimes Low
S2 Non-Violent Crimes High — 14.5% FP on security/education content; 7.2% with context_prompt
S3 Sex-Related Crimes Low
S4 Child Sexual Exploitation Low
S5 Defamation Medium
S6 Specialised Advice (medical, legal, or financial) High — triggers on any professional context
S7 Privacy Violations Medium
S8 Intellectual Property Infringement Medium
S9 Weapons of Mass Destruction (CBRN) Low
S10 Hate Speech Medium — triggers on academic and historical text
S11 Suicide and Self-Harm Low
S12 Explicit Sexual Content Low
S13 Elections Integrity Medium
S14 Code Interpreter Abuse Low — only fires in agentic/tool-use scenarios

💡 Note: For action: block, use only the low-FP categories: S1, S3, S4, S9, S11, S12, S14. This set produces approximately 1.7% false positives on OR-Bench-hard and drops to approximately 1.1% with context_prompt on security and education platforms. Avoid S2 and S6 for blocking. Use action: flag if you need visibility into S2, S6, or S10 without blocking.


Category filtering

When categories is set to an array, the guardrail only blocks or flags violations within the listed categories. Content classified as unsafe for a category not in the list is treated as safe.

When categories is null or omitted, all 14 categories are enforced.


Context injection

The context_prompt field lets you describe your deployment context so Llama Guard 3 can apply appropriate judgement. The text is prepended to each user message before classification.

This is most useful when your platform serves professionals who regularly ask questions that superficially resemble harmful requests:

{
  "type": "prompt_guard",
  "name": "safety-filter",
  "action": "block",
  "target": "request",
  "categories": ["S1", "S3", "S4", "S9", "S11", "S12", "S14"],
  "context_prompt": "This platform serves security professionals, educators, and researchers. Users regularly ask educational questions about cybersecurity, malware, vulnerabilities, chemistry, history, and other professional topics as part of legitimate work. Classify a request as unsafe only when it is clearly seeking operational assistance for causing direct harm — not when it is asking how something works, its history, or how to defend against it."
}

Measured impact (OR-Bench-hard, 10% sample):

Configuration Recommended_block FP S2 alone FP
No context ~1.7% ~14.5%
With context_prompt ~1.1% ~7.2%

💡 Note: Before classification the gateway truncates the inspected text to its last 9,000 characters, keeping the most recent content and cutting any excess from the front. This bound is on characters, not tokens, and applies whether or not context_prompt is set. Llama Guard 3 has its own context window of roughly 4,096 tokens — that is a property of the model, not a gateway setting.


Actions

Action Behaviour
block The request or response is denied. The caller receives a synthetic assistant message identifying which categories triggered the block.
flag The violation is recorded in the request log. The pipeline continues without modification.

💡 Note: When a request is blocked, the gateway returns a synthetic assistant message identifying the triggering categories. For example:

Request blocked by content policy (safety-filter): S1 – Violent Crimes, S9 – CBRN

The value of the name field of the guardrail appears in the message, making it easy to correlate blocks with your guardrail configuration. When the classifier reported a violation without naming a category, the reason reads uncategorized; the classifier's own wording is written to the gateway error log rather than returned to the caller.


Classifier output validation

The gateway does not trust the classifier's answer. Llama Guard 3 must reply with one of exactly two shapes:

Completion Meaning
safe No violation. Surrounding whitespace and trailing ASCII punctuation are tolerated; any further word or non-ASCII text after the verdict is not.
unsafe followed by a category list A violation. Categories are S<n> codes separated by commas, whitespace, newlines, or a colon — for example unsafe\nS1,S9.

Anything else — an empty completion, a repetition loop, prose, JSON, a foreign alphabet, a truncated verdict word — is indeterminate: neither safe nor unsafe. An indeterminate result is treated exactly like a classifier outage (see fail_open below), never as a pass. It is recorded as guardrail_verdict: indeterminate in the request log, raises a [guardrail_indeterminate] entry in the gateway error log, and feeds the guardrail_unavailable platform alert, so a broken classifier is visible within minutes instead of silently passing traffic.

Two related rules:

  • An unsafe verdict is never downgraded. If the classifier says unsafe but its category list is malformed or truncated, the detection still stands — the gateway enforces whichever valid S<n> codes it can read.
  • An unsafe verdict the classifier did not attribute to any category blocks (or flags) with the reason uncategorized when no categories allowlist is configured. When an allowlist is configured, the result is indeterminate instead: the gateway cannot tell whether the violation is one you asked it to enforce, so it will not guess in either direction.

💡 Note: Only category codes that stand alone as a complete token are recognised. A classifier that writes prose containing something like abschnitt s5 does not thereby report category S5.


fail_open behaviour

fail_open governs what happens when the guardrail cannot produce a verdict — whether because the sidecar was unreachable, or because it replied with output the gateway could not read.

fail_open Sidecar unavailable or classifier output unreadable
true (default) Request passes through as if no violation was found. The event is still logged as guardrail_verdict: error (unreachable) or indeterminate (unreadable) and raises a platform alert.
false Request is blocked with the "temporarily unavailable" message

⚠️ Caution: Set fail_open: false in environments where safety enforcement must never be bypassed. With fail_open: true, a sidecar outage — or a classifier replica returning nonsense behind HTTP 200 — allows all traffic through unclassified. The gateway will tell you loudly that this is happening, but it will not stop the traffic.

💡 Note: Verdicts recorded on an internal inference leg (for example an agentic sub-fetch or a summarisation step) that is itself blocked are not carried onto the parent request's log row, so a leg-level indeterminate appears in the gateway error log and the model-error triage queue rather than in request_log.


Limitations

  • scrub action is not supported. Configure block or flag only.
  • Request-phase classification evaluates only the last user message. The full conversation history is not sent to the classifier.
  • Before classification the gateway truncates the inspected text to its last 9,000 characters (the most recent content is kept). This is a character bound, independent of token count and of whether context_prompt is set. Llama Guard 3's own context window of roughly 4,096 tokens is a separate model limit.
  • Only the classifier's documented output shapes are accepted; anything else is an indeterminate result governed by fail_open (see Classifier output validation). Pointing the guardrail at a model that does not follow the Llama Guard 3 output contract therefore makes every request indeterminate.
  • Llama Guard 3 classifies text. A turn whose content carries only images, audio, documents, or tool results contains nothing for it to read. On the request phase the guardrail then falls back to the most recent user message that does carry text; if there is none, the request proceeds entirely unclassified — including any earlier turns the model still sees. A malformed content shape is different: content the gateway cannot interpret at all (a content array of scalars, a text block whose text is not a string) is an indeterminate result governed by fail_open, not a silent pass.
  • Prompt Guard is optimised for unstructured content policy enforcement. For structured sensitive data (PII, card numbers, credentials), use the Regex guardrail or the NLP PII Detector.

Example configurations

{
  "type": "prompt_guard",
  "name": "safety-filter",
  "action": "block",
  "target": "both",
  "categories": ["S1", "S3", "S4", "S9", "S11", "S12", "S14"]
}

Block with deployment context to reduce false positives

{
  "type": "prompt_guard",
  "name": "safety-filter",
  "action": "block",
  "target": "request",
  "categories": ["S1", "S3", "S4", "S9", "S11", "S12", "S14"],
  "context_prompt": "This platform serves security professionals and researchers. Classify as unsafe only requests clearly seeking operational assistance for causing direct harm."
}

Flag specialised advice in responses for audit (no blocking)

{
  "type": "prompt_guard",
  "name": "flag-advice",
  "action": "flag",
  "target": "response",
  "categories": ["S6"]
}

Enforce all 14 categories on requests, blocking on sidecar failure

{
  "type": "prompt_guard",
  "name": "full-safety",
  "action": "block",
  "target": "request",
  "fail_open": false
}

Layer Prompt Guard after keyword pre-filtering

Running keyword guardrails first (Tier 1) catches simple jailbreak strings before the more expensive sidecar call of Prompt Guard.

[
  {
    "type": "keyword",
    "name": "jailbreak-terms",
    "action": "block",
    "target": "request",
    "keywords": ["ignore previous instructions", "DAN mode", "jailbreak"],
    "case_sensitive": false
  },
  {
    "type": "prompt_guard",
    "name": "safety-filter",
    "action": "block",
    "target": "both",
    "categories": ["S1", "S3", "S4", "S9", "S11", "S12", "S14"]
  }
]

Configuring Prompt Guard

Screenshot: Prompt Guard card in the Guardrail Builder Prompt Guard card — expanded view

Proceed as follows to configure Prompt Guard in the Guardrail Builder:

  1. Open the gateway detail page and scroll down to the Guardrails card.
  2. Click on the + Prompt Guard button.
  3. A collapsed Prompt Guard card appears at the bottom of the list.
  4. Click on the card to expand it.
  5. Enter a name in the Name text field.
  6. Select the action from the Action drop-down list: block or flag.
  7. The drop-down also lists scrub, but Prompt Guard does not support it — scrub is treated as flag.
  8. Select the target from the Target drop-down list: request, response, or both.
  9. If required, select specific safety categories from the Safety categories list. Leave empty to enforce all 14 categories.
  10. If required, enter a deployment context in the Context Prompt text field to reduce false positives.
  11. Toggle the Fail Open switch to false if the sidecar must be a hard dependency.
  12. Click on the Save Guardrails button.

-> Prompt Guard is saved and appears in the execution plan.


Pipeline position

Prompt Guard is Tier 2 — it makes an HTTP call to a sidecar service. All Tier 1 guardrails (regex, keyword) run before any Tier 2 guardrail. Within Tier 2, guardrails run in the order they appear in the guardrails array.


See also