{"slug": "holtercare-bench-a-multimodal-benchmark-for-evaluating-long-term-dynamic-ecg", "title": "Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis", "summary": "Researchers at Zhejiang University introduced Holtercare-23K, a multimodal dynamic ECG dataset with 22,980 QA pairs from 788 clinical Holter records, and Holtercare-Bench, a benchmark evaluating MLLMs on temporal localization, clinical diagnosis, and global summarization. Zero-shot evaluations of leading MLLMs revealed significant performance gaps on ultra-long pathological sequences, but fine-tuning improved results. The project is available at https://github.com/ZJU4HealthCare/Holtercare-Bench.", "body_md": "arXiv:2608.19297v1 Announce Type: new\nAbstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high-quality datasets and benchmarks. To address this, we introduce (i) Holtercare-23K, a large-scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records and featuring a novel signal-video-text tri-modal alignment. Based on this dataset, we present (ii) Holtercare-Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization. Zero-shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra-long pathological sequences. However, fine-tuning representative models yields substantial improvements. This work illuminates the limitations of current MLLMs in electrophysiology and provides a foundational benchmark for long-term medical MLLMs. Our project is available at https://github.com/ZJU4HealthCare/Holtercare-Bench.", "url": "https://wpnews.pro/news/holtercare-bench-a-multimodal-benchmark-for-evaluating-long-term-dynamic-ecg", "canonical_source": "https://arxiv.org/abs/2608.19297", "published_at": "2026-08-21 04:00:00+00:00", "updated_at": "2026-08-21 04:16:03.176477+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-tools"], "entities": ["Zhejiang University", "Holtercare-23K", "Holtercare-Bench", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/holtercare-bench-a-multimodal-benchmark-for-evaluating-long-term-dynamic-ecg", "markdown": "https://wpnews.pro/news/holtercare-bench-a-multimodal-benchmark-for-evaluating-long-term-dynamic-ecg.md", "text": "https://wpnews.pro/news/holtercare-bench-a-multimodal-benchmark-for-evaluating-long-term-dynamic-ecg.txt", "jsonld": "https://wpnews.pro/news/holtercare-bench-a-multimodal-benchmark-for-evaluating-long-term-dynamic-ecg.jsonld"}}