cd /news/large-language-models/large-language-models-are-approximat… · home › topics › large-language-models › article
[ARTICLE · art-142981] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Large Language Models are Approximate Survival Estimators

A new arXiv paper (2609.38181v1) introduces Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes to prompt pre-trained LLMs for zero-shot survival prediction, benchmarked against random survival forests (RSF) on the MSK-CHORD cohort and a newly curated Providence St. Joseph Health Network cohort. GPT-5.6-Sol achieved censored mean absolute error (cMAE) within 10% of state-of-the-art RSF models for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD, but LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated high- from low-risk patients (lower c-index). The authors conclude that zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance remains an important limitation for clinical use.

by read1 min views1 publishedOct 1, 2026

arXiv:2609.38181v1 Announce Type: new Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readily accessible through consumer applications. However, whether LLMs can provide accurate survival predictions has not been rigorously evaluated. We introduce Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes and prompts pre-trained LLMs to predict survival zero-shot. We benchmark Survprompt against conventional survival models, including random survival forests (RSF), across two multi-institutional pan-cancer cohorts: the publicly available MSK-CHORD cohort and a newly curated cohort from the Providence St. Joseph Health Network constructed using an LLM-based medical abstraction framework. We report censored mean absolute error (cMAE) and concordance index (c-index) and conduct feature ablations to identify variables influencing LLM predictions. Frontier LLMs achieved surprisingly competitive cMAE for individual survival times. For example, GPT-5.6-Sol achieved cMAE within 10% of state-of-the-art RSF models specifically trained for survival prediction for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD. Feature ablations revealed that LLMs prioritized clinical variables similarly to specialized survival models. However, LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated between high- and low-risk patients (lower c-index). Zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance across cancer types and institutions remains an important limitation for clinical use.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/large-language-model…] indexed:0 read:1min 2026-10-01 · —