{"slug": "myocardbench-a-real-world-data-benchmark-for-evaluating-large-language-models-in", "title": "MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios", "summary": "GPT-5.4 achieved the highest macro-average score of 62.55 on MyoCardBench, a new real-world benchmark for evaluating large language models in cardiovascular care, according to a study published on arXiv. The benchmark includes 2,263 items from 13 task-specific datasets and assessed seven LLMs across clinical dimensions, with GPT-5.4 also ranking first in all three dimensions. The study found that CardioAuxReport performed best at 86.38, while CardioECGRead and CardioEthics scored lowest at 17.25 and 17.34, respectively.", "body_md": "arXiv:2607.25186v1 Announce Type: new\nAbstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop MyoCardBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: MyoCardBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, MyoCardBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.", "url": "https://wpnews.pro/news/myocardbench-a-real-world-data-benchmark-for-evaluating-large-language-models-in", "canonical_source": "https://arxiv.org/abs/2607.25186", "published_at": "2026-07-29 04:00:00+00:00", "updated_at": "2026-07-29 04:26:51.265549+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-ethics", "ai-safety"], "entities": ["MyoCardBench", "GPT-5.4", "Gemini 3.1 Pro", "Qwen 3.6 27B", "CardioAuxReport", "CardioECGRead", "CardioEthics", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/myocardbench-a-real-world-data-benchmark-for-evaluating-large-language-models-in", "markdown": "https://wpnews.pro/news/myocardbench-a-real-world-data-benchmark-for-evaluating-large-language-models-in.md", "text": "https://wpnews.pro/news/myocardbench-a-real-world-data-benchmark-for-evaluating-large-language-models-in.txt", "jsonld": "https://wpnews.pro/news/myocardbench-a-real-world-data-benchmark-for-evaluating-large-language-models-in.jsonld"}}