cd /news/artificial-intelligence/prompting-is-not-enough-supervised-b… · home › topics › artificial-intelligence › article
[ARTICLE · art-100832] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

A study of 21 audio-recorded pediatric surgical encounters found that zero-shot prompting of Qwen 2.5 32B achieved a macro Cohen's kappa of 0.139 (95% CI 0.111-0.164) for detecting shared decision-making behaviors, while a supervised classifier over frozen sentence embeddings reached 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119), and a logistic stack of the two reached 0.242 (0.198-0.284). The authors conclude that zero-shot prompting alone is insufficient and that patient-level grouping does not prevent leakage when labeled exemplars are precomputed outside the outer evaluation loop.

read1 min views20 publishedAug 18, 2026

arXiv:2608.14792v1 Announce Type: new Abstract: Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @qwen 2.5 32b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/prompting-is-not-eno…] indexed:0 read:1min 2026-08-18 · —