{"slug": "one-in-three-ai-scribe-notes-carries-a-verified-clinical-error", "title": "One in three AI scribe notes carries a verified clinical error", "summary": "A preprint audit of three commercial AI scribes on 142 consultations found that 31.3% of 565 notes carried a verified clinical error, with failures concentrated in allergy and medication information, invented patient identity, and history written as examination on telephone consultations. The study, released on arXiv on 31 Aug 2026, released all 618 findings with transcript-side evidence, prompts, and model versions, and found that the review instruction alone could move the verified candidate share from 9.3% to 79.0%.", "body_md": "# Computer Science > Computation and Language\n\n[Submitted on 31 Aug 2026]\n\n# Title:One note in three: a verified census of three deployed AI scribes, and the instrument that counted it\n\n[View PDF](/pdf/2608.31017)\n\n[HTML (experimental)](https://arxiv.org/html/2608.31017v1)\n\nAbstract:Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.\n\n### Current browse context:\n\ncs.CL\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/one-in-three-ai-scribe-notes-carries-a-verified-clinical-error", "canonical_source": "https://arxiv.org/abs/2608.31017", "published_at": "2026-09-02 03:07:07+00:00", "updated_at": "2026-09-02 03:21:53.866240+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research", "ai-products"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/one-in-three-ai-scribe-notes-carries-a-verified-clinical-error", "markdown": "https://wpnews.pro/news/one-in-three-ai-scribe-notes-carries-a-verified-clinical-error.md", "text": "https://wpnews.pro/news/one-in-three-ai-scribe-notes-carries-a-verified-clinical-error.txt", "jsonld": "https://wpnews.pro/news/one-in-three-ai-scribe-notes-carries-a-verified-clinical-error.jsonld"}}