{"slug": "large-language-models-are-approximate-survival-estimators", "title": "Large Language Models are Approximate Survival Estimators", "summary": "A new arXiv paper (2609.38181v1) introduces Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes to prompt pre-trained LLMs for zero-shot survival prediction, benchmarked against random survival forests (RSF) on the MSK-CHORD cohort and a newly curated Providence St. Joseph Health Network cohort. GPT-5.6-Sol achieved censored mean absolute error (cMAE) within 10% of state-of-the-art RSF models for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD, but LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated high- from low-risk patients (lower c-index). The authors conclude that zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance remains an important limitation for clinical use.", "body_md": "arXiv:2609.38181v1 Announce Type: new \nAbstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readily accessible through consumer applications. However, whether LLMs can provide accurate survival predictions has not been rigorously evaluated. We introduce Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes and prompts pre-trained LLMs to predict survival zero-shot. We benchmark Survprompt against conventional survival models, including random survival forests (RSF), across two multi-institutional pan-cancer cohorts: the publicly available MSK-CHORD cohort and a newly curated cohort from the Providence St. Joseph Health Network constructed using an LLM-based medical abstraction framework. We report censored mean absolute error (cMAE) and concordance index (c-index) and conduct feature ablations to identify variables influencing LLM predictions. Frontier LLMs achieved surprisingly competitive cMAE for individual survival times. For example, GPT-5.6-Sol achieved cMAE within 10% of state-of-the-art RSF models specifically trained for survival prediction for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD. Feature ablations revealed that LLMs prioritized clinical variables similarly to specialized survival models. However, LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated between high- and low-risk patients (lower c-index). Zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance across cancer types and institutions remains an important limitation for clinical use.", "url": "https://wpnews.pro/news/large-language-models-are-approximate-survival-estimators", "canonical_source": "https://arxiv.org/abs/2609.38181", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:18:55.354916+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "artificial-intelligence"], "entities": ["arXiv", "Survprompt", "GPT-5.6-Sol", "MSK-CHORD", "Providence St. Joseph Health Network", "random survival forests"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/large-language-models-are-approximate-survival-estimators", "markdown": "https://wpnews.pro/news/large-language-models-are-approximate-survival-estimators.md", "text": "https://wpnews.pro/news/large-language-models-are-approximate-survival-estimators.txt", "jsonld": "https://wpnews.pro/news/large-language-models-are-approximate-survival-estimators.jsonld"}}