arXiv:2609.38181v1 Announce Type: new Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readily accessible through consumer applications. However, whether LLMs can provide accurate survival predictions has not been rigorously evaluated. We introduce Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes and prompts pre-trained LLMs to predict survival zero-shot. We benchmark Survprompt against conventional survival models, including random survival forests (RSF), across two multi-institutional pan-cancer cohorts: the publicly available MSK-CHORD cohort and a newly curated cohort from the Providence St. Joseph Health Network constructed using an LLM-based medical abstraction framework. We report censored mean absolute error (cMAE) and concordance index (c-index) and conduct feature ablations to identify variables influencing LLM predictions. Frontier LLMs achieved surprisingly competitive cMAE for individual survival times. For example, GPT-5.6-Sol achieved cMAE within 10% of state-of-the-art RSF models specifically trained for survival prediction for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD. Feature ablations revealed that LLMs prioritized clinical variables similarly to specialized survival models. However, LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated between high- and low-risk patients (lower c-index). Zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance across cancer types and institutions remains an important limitation for clinical use.
Large Language Models are Approximate Survival Estimators
A new arXiv paper (2609.38181v1) introduces Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes to prompt pre-trained LLMs for zero-shot survival prediction, benchmarked against random survival forests (RSF) on the MSK-CHORD cohort and a newly curated Providence St. Joseph Health Network cohort. GPT-5.6-Sol achieved censored mean absolute error (cMAE) within 10% of state-of-the-art RSF models for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD, but LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated high- from low-risk patients (lower c-index). The authors conclude that zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance remains an important limitation for clinical use.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.