04:00
2026-07-29
arxiv.org
large-language-models
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
GPT-5.4 achieved the highest macro-average score of 62.55 on MyoCardBench, a new real-world benchmark for evaluating large language models in cardiovascular care, according to a study published on arXβ¦