cd /news/natural-language-processing/clinical-reasoning-under-a-partially… · home topics natural-language-processing article
[ARTICLE · art-129813] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation

A maxillofacial cone beam CT report generation system scored 0.4122 against a composite objective that weights 80% on large language model factual entailment and 20% on lexical overlap, versus 0.2909 for a report selected on the visible lexical ranking alone, according to an arXiv paper (arXiv:2609.13238v1). The authors report that optimizing for n-gram overlap drives entailment precision from 0.522 down to 0.266, and that acquisition centre alone predicts sentence choice at 0.718 versus 0.663 for an image-derived model, identifying dictation convention rather than anatomy as what lexical metrics reward. The delivered system emits eight unconditional statements and five gated on header geometry and reaches METEOR 0.3542 over 50 held-out cases from an unseen centre, with dataset and code at https://github.com/GIND123/CBCT-Clinical-Reasoner.

by read1 min views2 publishedSep 15, 2026

arXiv:2609.13238v1 Announce Type: new Abstract: Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which only the lexical fifth is visible during development. The grader's BLEU-4 and METEOR routines are reproduced in pure Python and match the reference to machine precision, and an offline entailment surrogate, which tells a report written for one patient from one written for another at an area under the curve of 0.987, makes the composite objective cheap enough to optimise directly. Over the 622-case public release, a report selected against the visible lexical ranking scores 0.2909, whereas one selected against the composite objective scores 0.4122, because pursuing n-gram overlap drives entailment precision from 0.522 down to 0.266. A 29 million parameter encoder fine-tuned on the release reaches a prevalence-weighted out-of-fold area under the curve of 0.486 over 985 statements, indistinguishable from the corpus prior, while nine numbers read from the image header reach 0.945 for mandible coverage and 0.872 for condyle coverage, and acquisition centre alone predicts sentence choice at 0.718 against 0.663 for the image-derived model, identifying dictation convention rather than anatomy as the quantity the lexical metrics reward. The delivered system emits eight unconditional statements and five gated on header geometry under polarity, laterality and tooth-level consistency constraints, and reaches METEOR 0.3542 over 50 held-out cases from an unseen centre. The dataset and code are available at https://github.com/GIND123/CBCT-Clinical-Reasoner

── more in #natural-language-processing 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/clinical-reasoning-u…] indexed:0 read:1min 2026-09-15 ·