Skip to content

Gibberish detector

The gibberish detector is a Tier 1 (in-process, sub-millisecond) guardrail that identifies low-quality, incoherent, or machine-garbled text in model responses. It is suited for ensuring response quality on customer-facing deployments and for catching model failure modes such as token repetition loops, encoding artefacts, or near-empty responses.

Screenshot: Gibberish detector editor in the Guardrail Builder Gibberish detector editor

When to use the gibberish detector

Use the gibberish detector when you need to block or log degraded model output — responses dominated by repeated tokens, symbol floods, binary encoding artefacts, or near-empty text. It inspects model responses only.

How it works

Three independent heuristic signals are computed against the response text. Each signal that exceeds its threshold counts as one signal hit.

Signal Detects Threshold
Shannon entropy Character repetition — e.g. aaaaaaa… or repeated tokens entropy < entropy_threshold (default 2.5)
Word repetition ratio Vocabulary collapse — few unique words relative to total word count unique_words / total_words < word_repeat_ratio (default 0.15); evaluated only when the response has at least 5 words
Alpha character ratio Non-text content — encoding artefacts, symbol flooding, binary data alpha_chars / total_chars < alpha_ratio (default 0.6)

Verdict rule:

Signal hits Result
1 Always flagged — recorded in detectors_fired, pipeline continues
2 or more Configured action applied — block or flag

A single signal hit never blocks on its own, regardless of the configured action. Blocking requires at least two signals. This prevents false positives from short, terse, or numeric responses.

💡 Note: Responses shorter than 20 characters are always passed without inspection. Very short responses (e.g. "Yes.") produce unreliable heuristic scores.

💡 Note: The alpha-ratio check is skipped when the response is predominantly non-Latin script — when more than 30% of its bytes are high-bytes (for example Cyrillic, Arabic, or CJK text). Such scripts are legitimately low in Latin alphabetic characters, so the ratio check would otherwise misfire.


Configuration reference

Field Type Default Description
type string — Must be "gibberish"
name string — Human-readable label for this guardrail instance
action string "block" What to do when 2 or more signals fire: block or flag
target string "request" Set this to "response" — the detector inspects model output only. The orchestrator default is "request", in which phase the detector never runs, so omitting target makes it a silent no-op. "target": "response" is required.
entropy_threshold number 2.5 Shannon entropy threshold. Responses with character entropy below this value are flagged as repetitive
word_repeat_ratio number 0.15 Minimum unique-word ratio. Responses where fewer than this fraction of words are unique are flagged
alpha_ratio number 0.6 Minimum alphabetic character ratio. Responses that are mostly non-alphabetic are flagged

💡 Note: The gibberish detector inspects model responses only, so "target": "response" must be set explicitly. The orchestrator defaults an absent target to "request", in which phase the detector does nothing — omitting target silently disables it.


Threshold tuning

Signal Default Loosen (fewer blocks) Tighten (more blocks)
entropy_threshold 2.5 Lower (e.g. 2.0) Raise (e.g. 3.0)
word_repeat_ratio 0.15 Lower (e.g. 0.05) Raise (e.g. 0.25)
alpha_ratio 0.6 Lower (e.g. 0.4) Raise (e.g. 0.75)

Example configurations

Block gibberish responses (default configuration)

{
  "type": "gibberish",
  "name": "quality-check",
  "action": "block",
  "target": "response"
}

Blocks responses that score badly on at least two of the three signals.

Flag only — monitoring mode

Records suspected gibberish in the log without disrupting the client:

{
  "type": "gibberish",
  "name": "quality-monitor",
  "action": "flag",
  "target": "response"
}

Strict quality enforcement

Tighter thresholds to catch more marginal responses — useful for structured-output pipelines where responses should contain well-formed prose or data:

{
  "type": "gibberish",
  "name": "strict-quality",
  "action": "block",
  "target": "response",
  "entropy_threshold": 3.0,
  "word_repeat_ratio": 0.25,
  "alpha_ratio": 0.7
}

Configuring the gibberish detector

The gibberish detector has no card in the visual Guardrail Builder. The builder exposes only the PII protector, Presidio, Prompt Guard, regex, keyword, custom PII, and jailbreak detector types — gibberish is configured through the gateway configuration API only, by adding a detector object to the gateway's guardrails array.

Proceed as follows:

  1. Build the detector as a JSON object — see Example configurations above.
  2. Add the object to the guardrails array in the gateway config and send it with a PATCH /admin/v1/gateways/{id} request:
curl -X PATCH https://<your-gateway-host>/admin/v1/gateways/{id} \
  -H "Content-Type: application/json" \
  -d '{
    "config": {
      "guardrails": [
        {
          "type": "gibberish",
          "name": "quality-check",
          "action": "block",
          "target": "response"
        }
      ]
    }
  }'

💡 Note: The guardrails array is replaced in full on each PATCH — include every detector the gateway should keep, not only the new one.

-> The gibberish detector is saved and appears in the execution plan.


Pipeline position

The gibberish detector is Tier 1 — it runs in-process with no external calls. It executes in the response phase only. A block verdict prevents the malformed response from reaching the client.


See also