LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
A new benchmark of 500 single-error clinical note pairs shows LLM judges detect added or altered content with paired discrimination scores of 0.79-0.94 but fail on omissions, scoring 0.50-0.63, according to a paper submi…