{"slug": "a-real-time-llm-stream-guard-that-catches-llm-hallucinations-mid-generation", "title": "A real-time LLM stream guard that catches LLM hallucinations mid generation", "summary": "HAL-X AI researchers F. Aghayev and E. Ahmadbayli released SIMURG, a streaming integrity monitor that detects LLM decoding corruption in real time, cutting the stream mid-flight to prevent users from seeing corrupted output. The tool processes 197,632 characters per second on a laptop CPU, detects corruption within ~590 characters of onset, and requires no model, GPU, or training, using only numpy. SIMURG is designed for production LLMs, especially quantized or small models, and offers a zero-leak guarantee by holding the stream's opening in a buffer until verified clean.", "body_md": "**Streaming Integrity Monitor & Universal Regeneration Guard**\n\nCatch LLM decoding corruption **while the answer is still being generated** and cut the\nstream **mid-flight**: corruption that starts in the hold window never reaches the\nuser, and mid-stream corruption is aborted within a few hundred characters of onset,\nso the host regenerates the answer.\n\n| throughput | detection latency | false-alarm budget | footprint | setup |\n|---|---|---|---|---|\n197,632 chars/sec on a laptop CPU |\n~590 chars past corruption onset |\nconfigurable, conformal-calibrated | numpy only, no model, no GPU | 3 lines, zero training |\n\nThe guard runs hundreds of times faster than a typical LLM produces text, so it is never the bottleneck: a model streaming at 50 tokens/sec writes ~250 chars/sec, and SIMURG reads 197,000.\n\n**Table of contents**\n\nWhen you run an LLM in production, especially a **quantized, small, or self-hosted**\nmodel, it sometimes **derails mid-generation**. The decoded stream stops doing the\ntask and collapses into one of a handful of pathologies:\n\n| failure mode | what it looks like |\n|---|---|\nrepetition collapse |\nthe same phrase, list, or token repeated until the token budget runs out |\ncross-lingual drift |\nan English answer that quietly slides into Chinese, Arabic, or Cyrillic |\nregurgitation |\nthe model dumps a README, boilerplate, or training text |\nstructural breakdown |\n`#REF! -0.00 -0.00 ... 0.00` : number and symbol garbage |\ntemplate leakage |\n`< |\n\nThis is **not** factual hallucination. A fluent-but-wrong sentence (see\n[What SIMURG is NOT](#what-simurg-is-not)) has no statistical scar. What is shown\nabove is **decoding corruption**, and it leaves a *statistical signature in the\ntoken stream*: repetition rate, lexical variety, script distribution,\ncompressibility, and predictive surprise all move in measurable ways.\n\nSIMURG watches that signature character by character, decides in real time whether\nthe stream has gone bad, tells you **where** it started, and lets you **abort and\nretry** before the user ever sees the corruption.\n\nThe full technical report, with the complete evaluation, per-class analysis, onset-localization study, and the zero-leak protocol specification:\n\nSIMURG: Zero-Leak Online Detection of LLM Decoding Corruption in Production Streams, F. Aghayev, E. Ahmadbayli, HAL-X AI, 2026.[Read the paper (PDF, 13 pages)]\n\nSIMURG |\npost-hoc linter | LLM-as-judge | perplexity threshold | |\n|---|---|---|---|---|\n| when it fires | mid-generation, ~590 chars past onset |\nafter the full answer | after the full answer | post-hoc, or needs logprob access |\n| what the user sees | zero bad tokens when onset is in the hold window; otherwise the clean prefix plus a bad tail of at most ~900 chars, replaced by the retry |\nthe whole corrupt answer | the whole corrupt answer | varies |\n| why it fired | a named, human-readable reason on every alarm | a pattern list | the judge's opinion, if any | one number |\n| model-agnostic | any OpenAI-compatible endpoint, or any stream you feed | any | any | needs a logprob-capable backend |\n| overhead | numpy-only, ~197k chars/sec on one CPU core | trivial | one extra LLM call per answer | per-token logprobs |\n\nThe zero-leak property is the point: post-hoc checks can only tell you that the\nanswer was bad *after the user read it*. SIMURG holds the opening of every stream\nin a buffer, releases it only once it is verified clean, keeps re-checking, and\ncuts the stream the moment it crosses the calibrated threshold.\n\nSIMURG makes **one O(1)-per-character pass** over the stream, maintaining a set of\nincremental features (digit fraction, foreign-script fraction, repetition rate,\ncompressibility, type-token ratio, script-switch rate, structural-artifact density,\n...), and feeds a pluggable detector ensemble on top of them:\n\n``` php\nflowchart TD\n    A[\"token stream\"] --> B[\"stream features<br/>one O(1) per character incremental pass\"]\n    B --> C1[\"char n-gram surprise<br/>self-calibrating, no reference corpus\"]\n    B --> C2[\"Count-Min repetition sketch<br/>constant memory, 8k counters\"]\n    B --> C3[\"rolling SimHash drift<br/>topic collapse detection\"]\n    B --> C4[\"robust-z self-calibration<br/>baselines frozen on the clean prefix\"]\n    B --> C5[\"rule tier<br/>interpretable thresholds, zero training\"]\n    C1 --> D[\"conformal fusion<br/>finite-sample false-alarm budget\"]\n    C2 --> D\n    C3 --> D\n    C4 --> D\n    C5 --> D\n    L[\"learned tier<br/>15-weight online logistic model\"] --> D\n    D --> E[\"CLEAN / SUSPECT / CORRUPT<br/>plus Page-Hinkley onset localization\"]\n    E --> F[\"zero-leak protocol<br/>HOLD first 350 chars, RELEASE if clean,<br/>re-check every 400, ABORT on corrupt\"]\n    F --> G[\"bad tokens never reach the UI\"]\n```\n\n| detector | what it measures | why it catches corruption |\n|---|---|---|\nchar n-gram surprise |\npredictive surprise of each char against an in-stream 3-gram model | loops and garbage drive surprise toward zero |\nCount-Min repetition |\nn-gram repetition rate in a constant-memory sketch | repetition collapse is the most common production failure |\nrolling SimHash drift |\ndistance of a 48-token fingerprint from the clean-prefix baseline | topic collapse and regurgitation move the fingerprint |\nrobust-z self-calibration |\nevery feature z-scored against its own frozen clean-prefix baseline | no hand-tuned magic numbers, adapts to any domain |\nrules |\ninterpretable thresholds (digit fraction, script switch, template markers, ...) | day-one coverage, every alarm is a sentence a human can read |\n\n**Rule tier.** Interpretable thresholds on the stream features. Works on day one with**zero training**, and every alarm is explainable:`\"repetition loop rate=0.71\"`\n\n,`\"digit fraction 0.57\"`\n\n,`\"script switch en to zh\"`\n\n.**Learned tier.** A small**online logistic regression**(15 weights, a few KB) that adds robustness and** keeps learning in production**via`partial_fit`\n\n.\n\nThe fusion layer sets its thresholds from the score distribution on *clean*\nstreams, which gives a **finite-sample guarantee on the false-alarm rate**. \"Flag\nat most 2% of clean outputs\" is a knob you set and the calibration enforces, not a\nthreshold you hope holds.\n\n**HOLD** the first 350 characters. A stream that is corrupt from the start is killed before a single character reaches the UI.**RELEASE** the prefix if it scores clean, and freeze the self-calibrated baselines on it.**Re-check** every 400 characters for the rest of the stream.**ABORT** on a calibrated threshold crossing (with a 2-hit or hard-rule hysteresis so a single noisy checkpoint does not kill a good answer).\n\nA synthetic stream that is clean prose and then collapses into a repetition loop\nat character 339. SIMURG holds the opening, verifies the clean prefix, scores\nthe stream at every 400-char checkpoint, and aborts 821 characters after the\nloop starts. Corrupt streams that are already bad at the 350-char checkpoint\nare blocked fully (12 of 21 in the benchmark, see below); for this mid-stream\nonset the user sees the clean prefix plus a short bad tail, and the guard's\ncontract with the host is a **retry**: `GuardedLLM`\n\nregenerates the answer and\nthe host replaces the shown text, so the bad tail never becomes the final\noutput:\n\nEvery alarm carries the reasons that fired it. For the stream above:\n\n```\nrepetition loop rate=0.66 zlib=0.10\nvocabulary collapse ttr=0.09\nsurprise collapse low_frac=1.00\n```\n\nReproducible end-to-end benchmark: builds the **CorruptBench** synthetic set\n(243 streams, 4 failure classes), trains the learned tier, calibrates the\nconformal thresholds, and reports the full table:\n\n```\npip install -e .\npython3 -m simurg.data.evaluate          # seed 7, deterministic dataset\n```\n\nTest split (81 streams), seed 7:\n\n| metric | value |\n|---|---|\n| stream-level TPR | 78/80 = 0.975 |\n| recall, repetition collapse | 16/18 = 0.89 |\n| recall, cross-lingual drift | 25/25 = 1.00 |\n| recall, regurgitation | 19/19 = 1.00 |\n| recall, structural breakdown | 18/18 = 1.00 |\n| detection latency past onset | median 590, p90 868 chars |\n| onset localization error | median 532 chars |\n| zero-leak (onset inside hold window) | 12/21 blocked fully |\n| throughput | 197,632 chars/sec |\n| stream-level AUROC (final score) | 0.55, dragged down by ties at p=1.0 and a 1-stream clean test split; TPR/FPR at the calibrated threshold is the operating metric |\n\nIn addition, the shipped detector **flagged 0 false alarms on 121 real production\ntexts** from a self-hosted reasoning-model deployment.\n\n**Those numbers describe the bundled domain.** The detector is only as good as the\nclean corpus it calibrates against, so retrain on your own traffic before you\ntrust it in production. It takes seconds, see\n[below](#teach-it-your-domain-and-your-failure-modes).\n\n```\npip install simurg        # numpy only\npip install simurg[figures]   # + matplotlib, for the paper plots\npip install simurg[test]      # + pytest\n```\n\nFrom source:\n\n```\ngit clone https://github.com/doofzoff/SIMURG.git\ncd SIMURG\npip install -e .\n```\n\nWorks with **vLLM, llama.cpp server, TGI, Ollama, OpenAI, OpenRouter**: anything\nthat speaks `/v1/chat/completions`\n\n. Batteries included: the zero-leak protocol\nplus an **abort, retry, fallback-model** ladder.\n\n``` python\nfrom simurg import GuardedLLM\n\nllm = GuardedLLM(\n    \"http://localhost:8000/v1\", model=\"my-model\",\n    retries=1,\n    fallback=GuardedLLM(\"https://openrouter.ai/api/v1\",\n                        model=\"qwen/qwen3\", api_key=\"sk-...\"),   # optional\n)\n\nresult = llm.chat(\n    [{\"role\": \"user\", \"content\": \"Explain how oil prices affect a small economy.\"}],\n    on_token=lambda t: print(t, end=\"\", flush=True),             # only CLEAN text is ever forwarded\n)\n\nprint(result.ok)        # True if a clean answer was produced\nprint(result.verdict)   # \"clean\" | \"suspect\" | \"corrupt\"\nprint(result.attempts)  # the full ladder: what each attempt did and why\n```\n\nIf an attempt corrupts, **nothing from it reaches on_token**. A corrupt attempt\nis retried; if all retries fail, the fallback model is tried.\n\nNot on an OpenAI-style API? Wrap your own token loop:\n\n``` python\nfrom simurg import Simurg\n\ns = Simurg()                          # rule tier works with zero setup\nfor token in my_llm_stream():\n    v = s.feed(token)\n    if v.state == \"corrupt\":\n        abort_and_retry(reason=v.reasons, onset=v.onset_char)\n        break\n    ui.write(v.released)              # text cleared for display (may lag while holding)\nfinal = s.finish()\nui.write(final.released)\npython\nfrom simurg import Simurg\n\ns = Simurg()\ns.feed(whole_text)\nprint(s.finish().state)               # \"clean\" / \"suspect\" / \"corrupt\"\n```\n\nFeed the calibration step **your** good outputs so the thresholds fit your domain:\n\n```\n# bring your own clean corpus (.jsonl with a \"text\" field per line)\nSIMURG_CORPUS_JSONL=/path/to/my_clean_outputs.jsonl python3 -m simurg.data.evaluate --save\n```\n\nFull guide, including the quick path, the live dashboard, and the production\nflywheel: ** docs/TRAINING.md**.\n\nGive SIMURG examples of *your* model's bad outputs. It tells you **whether that\nfailure is even catchable** in stream statistics, and hands you a fitted detector\nif it is:\n\n``` python\nfrom simurg import fit_custom_detector\n\nreport, detector = fit_custom_detector(\n    \"template_leak\",\n    clean_texts   = my_good_outputs,     # 50+\n    corrupt_texts = my_bad_outputs,      # 20+\n)\nprint(report)\n#  verdict: DETECTABLE   held-out AUROC: 0.98   -> auto-registered into every Simurg()\n```\n\nThe gate is the point: fluent factual lies come back ** NOT DETECTABLE** instead\nof a false promise. Details, plus the zero-training\n\n`LexiconDetector`\n\nfor known\nbad markers like `<|im_start|>`\n\n: **.**\n\n[docs/CUSTOM.md](/doofzoff/SIMURG/blob/main/docs/CUSTOM.md)\n\n```\npython3 -m simurg.training.train_live      # writes metrics for the bundled dashboard\n```\n\nA real-time web dashboard: log-loss, accuracy, AUROC, **all 15 weights animating\nper epoch**, memory, and the final held-out TPR/FPR verdict.\n\nA second web page for *runtime*: connect it to any OpenAI-compatible endpoint,\nsend a prompt, and watch the answer get guarded while it is generated. The\ndashboard renders in real time:\n\n- the\n**released stream text**(what the user would actually see), - the\n**fused corruption score** with the calibrated SUSPECT/ABORT thresholds and the 350-char hold zone, - the\n**corruption onset marker** and the human-readable**reasons**, **all 15 stream features** as sparklines, sampled at every checkpoint.\n\nEvery run is recorded as a **session** (timestamped frames with score, state,\nreleased text, features and reasons). The sessions panel lists them, deletes\nthem, and **replays any session at up to 128x** for postmortem analysis, so a\ncorrupt answer from Tuesday can be re-watched the way a crash log is read.\n\n```\npython3 -m simurg.guard_dashboard --port 8321\n# open http://127.0.0.1:8321, point it at your endpoint, guard a stream\n```\n\nPasted texts can also be analyzed at full speed in the same UI. Same self-contained dark style as the training dashboard, zero new dependencies: the server is stdlib-only and acts as a CORS-free proxy to your endpoint.\n\nSIMURG detects **corrupt or degenerate decoding**, not **factual wrongness**. A\nfluent, well-formed sentence that is simply *false* (\"the capital of Australia is\nSydney\") has no stream-statistical signature: it looks exactly like a true\nsentence. For that you need **grounding** (constrain the model to retrieved facts\nand make it quote them), retrieval verification, or a factuality checker.\n\nSIMURG guards the *delivery*; grounding guards the *content*. Use both.\n`fit_custom_detector`\n\nwill explicitly refuse to pretend it can catch this class.\n\n```\nsrc/simurg/\n├── core.py              taxonomy, detector protocol, registry\n├── features.py          the single O(1)/char stream-feature pass\n├── signals/             the raw estimators: n-gram surprise, Count-Min sketch,\n│                        rolling SimHash, robust-z calibration, Page-Hinkley\n├── detection/           rules, detectors, conformal fusion, sentinel (protocol)\n├── learning/            online logistic model, custom-failure-mode training (BYOC)\n├── integrations/        GuardedLLM, the OpenAI-compatible drop-in guard\n├── data/                CorruptBench synth, dataset builder, benchmark, generator\n├── training/            live-training run + real-time web dashboard\n├── guard_dashboard.py   live guard dashboard server (stdlib-only, SSE, sessions)\n├── guard_ui/            live guard dashboard front-end + recorded sessions\n└── weights/             shipped model + conformal thresholds (use as a pair)\ndocs/                    TRAINING.md, CUSTOM.md\nexamples/                runnable quickstart\ntests/                   sentinel regressions + end-to-end dashboard tests\nfigures/                 benchmark figures referenced by this README\npaper/                   the full technical report (PDF)\n.github/workflows/       CI: test matrix on 3.10 / 3.12 / 3.13 + build check\nCHANGELOG.md             release history\n```\n\nIdeas under active consideration, in rough priority order:\n\n**Engine-level abort.** Ship integrations that stop generation*inside*the inference engine (a vLLM streaming hook and a generic SSE middleware proxy), so an abort frees GPU time instead of just saving the UI. The guard already exposes everything a host needs; what is missing is the wiring.**Fleet telemetry.** Export`p(corrupt)`\n\n, verdict transitions, and onset positions as Prometheus metrics or OpenTelemetry spans, so a Grafana panel can show a*corruption rate per model and endpoint*and alert when a quantization or a prompt change starts producing bad streams.**Zero-dependency runtime.** Export the guard core (features, sketches, fusion) to ONNX or a small C library that runs inside the inference server with no Python, for hosts that cannot take a numpy dependency on the hot path.**CI regression suite.** A golden corpus of labeled clean and corrupt streams with fixed expected verdicts, plus latency and throughput budgets, run as a GitHub Action on every pull request: the build fails when a threshold tweak quietly degrades detection.**Multi-stream fleet mode.** Guard N parallel live streams in one process, with per-stream sessions and a single dashboard that compares corruption rates across endpoints, so a bad quantization shows up as one lane going red while the others stay green.\n\n**Will it catch factual hallucinations?**\nNo, and it will tell you so. Factual errors have no stream-statistical signature.\nUse grounding or a factuality checker for content, SIMURG for delivery.\n\n**What is the overhead?**\nOne O(1) pass per character, ~197k chars/sec on a laptop CPU. A 50 tok/s model\nwrites ~250 chars/sec, so the guard is hundreds of times faster than the model it\nguards. Memory is bounded per stream: 8,192 sketch counters, a 48-token SimHash\nwindow, and an n-gram table capped at 60k contexts.\n\n**Does it only work with English?**\nNo. Script features are language-agnostic (per-script fractions, switch rates),\nand you can declare your expected scripts at construction time\n(`Simurg(expected_scripts=(\"cyrillic\",))`\n\n). Retrain on your traffic for best\nresults.\n\n**What is the SUSPECT state for?**\nIt is a non-blocking warning tier between CLEAN and CORRUPT. Your host can use it\nto slow the UI down, show a subtle indicator, or pre-stage a retry, without\ndiscarding a stream that may still turn out clean.\n\n**How do I retrain on my own domain?**\n`SIMURG_CORPUS_JSONL=... python3 -m simurg.data.evaluate --save`\n\nover your clean\noutputs. It rebuilds the weights and the conformal thresholds in seconds. Full\nguide: [docs/TRAINING.md](/doofzoff/SIMURG/blob/main/docs/TRAINING.md).\n\n```\n@techreport{aghayev2026simurg,\n  title       = {SIMURG: Zero-Leak Online Detection of LLM Decoding Corruption in Production Streams},\n  author      = {Aghayev, Farid and Ahmadbayli, Elturan},\n  institution = {HAL-X AI},\n  year        = {2026},\n  url         = {https://github.com/doofzoff/SIMURG},\n  note        = {technical report, see paper/simurg_paper.pdf}\n}\n```\n\n**Apache-2.0**. See [LICENSE](/doofzoff/SIMURG/blob/main/LICENSE). Developed by **doofZ (Farid Aghayev)**,\nHAL-X AI.", "url": "https://wpnews.pro/news/a-real-time-llm-stream-guard-that-catches-llm-hallucinations-mid-generation", "canonical_source": "https://github.com/doofzoff/SIMURG", "published_at": "2026-08-24 21:00:25+00:00", "updated_at": "2026-08-24 21:13:06.231725+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["HAL-X AI", "F. Aghayev", "E. Ahmadbayli", "SIMURG"], "alternates": {"html": "https://wpnews.pro/news/a-real-time-llm-stream-guard-that-catches-llm-hallucinations-mid-generation", "markdown": "https://wpnews.pro/news/a-real-time-llm-stream-guard-that-catches-llm-hallucinations-mid-generation.md", "text": "https://wpnews.pro/news/a-real-time-llm-stream-guard-that-catches-llm-hallucinations-mid-generation.txt", "jsonld": "https://wpnews.pro/news/a-real-time-llm-stream-guard-that-catches-llm-hallucinations-mid-generation.jsonld"}}