{"slug": "llm-judges-verify-presence-not-absence-omission-blindness-in-ai-clinical-notes", "title": "LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes", "summary": "A new benchmark of 500 single-error clinical note pairs shows LLM judges detect added or altered content with paired discrimination scores of 0.79-0.94 but fail on omissions, scoring 0.50-0.63, according to a paper submitted to arXiv on 31 Aug 2026. Restructuring the task—listing facts from the transcript and checking each—recovers omission detection: a per-fact pipeline flags missing facts at 2.7% false alarms, while a GEPA-evolved single-call prompt detects 36.9% of omissions versus 24.6% for the pipeline (p=0.002) at 6.2% false alarms and a tenth of the cost per note. A physician author validated 70 items and sided with the pipeline on 10 of 10 disagreements (p=0.002), and a second clinician blind-graded the severity rubric to within one grade.", "body_md": "# Computer Science > Computation and Language\n\n[Submitted on 31 Aug 2026]\n\n# Title:LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It\n\n[View PDF](/pdf/2608.31016)\n\n[HTML (experimental)](https://arxiv.org/html/2608.31016v1)\n\nAbstract:Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the transcript and flags problems. We ask whether judges detect omissions. Public corpora cannot supply the answer key: their clinician reference notes and transcripts are materially discrepant. Our benchmark has 500 single-error note pairs from audited fact sheets, 298 with a named fact certainly absent and 202 added-or-altered controls. Across eight judge designs, paired discrimination (the flawed note below its clean twin, 0.5 a coin flip) reads 0.79-0.94 on added or altered content and 0.50-0.63 on omissions. On single notes, no design flags omissions reliably more often than perfect notes. Wording changes, voting and GEPA prompt optimisation move the operating point without creating usable detection. Restructuring the task recovers it: list the facts the transcript establishes, then check the note for each. Two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call. The pipeline's flags name the missing fact and its severity at 2.7% false alarms. The single call detects more (36.9% against 24.6%, p=0.002) at 6.2% false alarms and a tenth of the cost per note. A physician author validated 70 items and, where the two routes disagree, sided with the pipeline on 10 of 10 (p=0.002). A second clinician, not an author, graded the severity rubric blind and agrees to within a grade. On real vendor notes from a companion census no benchmark threshold transfers, but the re-calibrated single call detects more than the best of the eight at half its false-alarm rate. Omissions whose fact is restated elsewhere defeat both routes. We release the benchmark, prompts and judgements.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/llm-judges-verify-presence-not-absence-omission-blindness-in-ai-clinical-notes", "canonical_source": "https://arxiv.org/abs/2608.31016", "published_at": "2026-09-02 11:07:07+00:00", "updated_at": "2026-09-02 11:24:24.776203+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv", "GEPA"], "alternates": {"html": "https://wpnews.pro/news/llm-judges-verify-presence-not-absence-omission-blindness-in-ai-clinical-notes", "markdown": "https://wpnews.pro/news/llm-judges-verify-presence-not-absence-omission-blindness-in-ai-clinical-notes.md", "text": "https://wpnews.pro/news/llm-judges-verify-presence-not-absence-omission-blindness-in-ai-clinical-notes.txt", "jsonld": "https://wpnews.pro/news/llm-judges-verify-presence-not-absence-omission-blindness-in-ai-clinical-notes.jsonld"}}