cd /news/artificial-intelligence/cite-what-you-explore-budget-aware-l… · home › topics › artificial-intelligence › article
[ARTICLE · art-146618] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence

Researchers introduced BAR, a Budget-Aware LLM Reasoning framework that refines medical knowledge graphs into disease-specific evidence graphs with support scores and provenance records, then has an LLM reason over them via a plan-navigate-verify loop. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improved AUPRC by 3.4 points over the strongest baseline, raised citation precision from 59.8% to 77.9%, and consumed only 62-65% of the budget cap, according to the arXiv paper 2610.07739v1.

by read1 min views1 publishedOct 7, 2026

arXiv:2610.07739v1 Announce Type: new Abstract: Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three contributions. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @bar 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cite-what-you-explor…] indexed:0 read:1min 2026-10-07 · —