Via wisedocs.ai
Claude Fable 5 leads the pack at 64.4%, but most models can't even crack 15% on complex medical case file analysis
If you’ve ever tried to read a 150-page medical file and piece together a coherent timeline of what happened to a patient, you already know it’s brutal. Now imagine asking an AI to do it. Turns out, most of them are pretty bad at it too. The MLCR-AA leaderboard, launched on August 21 by Artificial Analysis, ranks AI models on their ability to reason through lengthy, complex medical and insurance case files. Anthropic’s Claude Fable 5 claimed the top spot with a score of 64.4%. That might not sound like a gold-star performance, but when the median model on the leaderboard scores below 15%, it starts to look a lot more impressive.
What the benchmark actually measures #
The MLCR-AA leaderboard draws from Wisedocs’ broader Medical Long Context Reasoning benchmark, which the company first introduced on June 18, 2026. The full benchmark includes 250 questions spread across six difficulty tiers, designed to simulate the kind of analytical grunt work that medical professionals and insurance claims adjusters do daily.
The leaderboard specifically tests models on the two hardest tiers, Expert and Compound, using 60 synthetic medical and insurance case questions. The underlying case files average 70 to 150 pages in length.
Wisedocs designed the benchmark from a position of deep domain knowledge. The company has trained its systems on over 100 million claim documents and uses an expert-in-the-loop validation process to ensure its outputs hold up in real-world settings like insurance and legal claims processing.
The accuracy-completeness paradox #
One of the more interesting findings from the leaderboard is a tension between two metrics that you’d expect to go hand in hand: accuracy and completeness. They don’t.
Nearly 40% of models surpass 80% accuracy on the questions they actually attempt to answer. But most models score below 50% on completeness. In plain terms, many models are good at getting individual answers right but terrible at answering all the questions a case file demands.
GPT-5.6 Terra (max) is the poster child for this disconnect. It posted the highest accuracy score on the leaderboard at 93.7%, meaning when it answered a question, it was almost always correct. But its completeness was so low that it ranked just 10th overall.
Claude Fable 5’s 64.4% composite score reflects a better balance between the two metrics, which is why it sits at the top. Other models in the Claude family scored between 53.9% and 59.4%. Among open-weights models, Kimi K3 (max) from Moonshot AI earned the highest ranking at 38.3%.
Why this matters beyond the leaderboard #
Wisedocs open-sourced the foundational elements of the MLCR benchmark on Hugging Face in June, including case files and lower-tier questions.
The top-performing model, Claude Fable 5, costs approximately $1 per task. Even so, the best model gets roughly a third of the job wrong. For an industry where errors can mean denied claims, missed injuries, or legal liability, 64.4% is a starting point, not a finish line.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our