Position: Medical AI Neglects Real Treatment Outcomes A position paper on arXiv (2608.14598v1) argues that medical AI is neglecting real treatment outcomes, relying instead on human opinions and text syntheses, which limits its potential and causes deficiencies in frontier models and major benchmarks. The authors call for incorporating actual treatment outcome data from observational databases and randomized experiments into training and evaluation, reemphasizing outcome improvement as the downstream goal. arXiv:2608.14598v1 Announce Type: new Abstract: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses especially texts such as biomedical publications and clinical practice guidelines rather than actual underlying data on treatment outcomes. This neglect seriously limits the potential of medical AI, and is already causing deficiencies in both frontier models and major benchmarks, as argued in this position paper. Real treatment outcomes, drawn from sources such as observational databases and randomized experiments, should be substantially incorporated into both training and evaluation. Improving these outcomes should be reemphasized as the downstream goal of all medical AI.