MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams Researchers introduced MIRA-Ev, a clinical argument mining benchmark built on Spanish MIR licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in Spanish, English, and Basque. The benchmark organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification, addressing the limitation of multiple-choice question answering that cannot detect when a model grounds correct diagnoses in irrelevant or contradictory evidence. arXiv:2607.19201v1 Announce Type: new Abstract: Clinical NLP evaluation remains dominated by multiple-choice question answering MCQA , which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish M\'edico Interno Residente MIR licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in parallel Spanish native , English, and Basque versions, the first clinical argumentation resource in Basque. MIRA-Ev organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification.