An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study A multi-site prospective study evaluating large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction found that LLM accuracy declined as question ambiguity increased, with mean question-level accuracy of 91.5% (SD 13.4%) across 157 questions, falling from 96% for Medication/Event Flag questions to 62% for Event Timing questions. The study, involving 9,430 abstractor answers reconciled to 4,715 consensus answers across two centers using American College of Cardiology National Cardiovascular Data Registry (ACC NCDR) registries, reported that 87% of LLM answers exactly matched consensus, 2% partially, and 9% did not, while human inter-rater agreement was approximately 98%. The authors concluded that LLMs achieved far lower accuracy than human abstractors, with accuracy steadily decreasing as ambiguity and required clinical reasoning increased. arXiv:2608.20373v1 Announce Type: new Abstract: Objective: To evaluate large language model LLM performance on unprocessed electronic medical record EMR data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry ACC NCDR . In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers 501 pilot; 4,214 validation . In the pilot, candidate data sources per question averaged between 14.6 SD 13.9 for demographics and 89.2 SD 56.1 for history and risk factors. In validation, human inter-rater agreement was approximately 98\% while 87\% of LLM answers exactly matched consensus, 2\% partially, and 9\% did not. Mean question-level accuracy was 91.5\% SD 13.4\% across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\% for Medication/Event Flag to 62\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.