cd /news/artificial-intelligence/confidence-estimation-for-financial-… · home topics artificial-intelligence article
[ARTICLE · art-89832] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

A new arXiv study (2608.06532v1) evaluating seven confidence estimators across five open-weight LVLMs on three financial visual question-answering benchmarks finds that only trained internal probes produce thresholdable calibration, while inference-only baselines are badly overconfident. The best estimator varies by model and task, with none leading more than eight of twenty (model, condition) cells, and a bilingual contrast shows apparent language robustness is a composition artifact. Under a strict 5% error budget, deferral clears a real share of the easiest condition but almost none of the hardest, with only the grounding-aware probe lowering confidence on answers given without using the figure.

read1 min views1 publishedAug 10, 2026

arXiv:2608.06532v1 Announce Type: new Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the exhibit. The operational question is therefore trust, not accuracy: which answers can be acted on, and which escalated to a reviewer. We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer. Three findings hold. First, the scarce property is calibration, not ranking: the inference baselines rank correct above incorrect answers competitively but are badly overconfident, calibration error far above what a threshold can tolerate, and only the trained probes produce a thresholdable score. Second, reliability is structured rather than global, along two axes a practitioner can read directly: the best estimator shifts with both model and task, none leading more than eight of twenty (model, condition) cells, and a controlled bilingual contrast exposes an apparent language robustness as a composition artifact that dissolves once models are read one at a time. Third, cast as deferral under an error budget, how much can be safely automated is set first by the model's competence and only narrowed by its confidence, so deferral clears a real share of the easiest condition and almost none of the hardest, near zero at a strict 5% budget. Two trained probes carry the calibration a deferral policy needs, and among them only the grounding-aware one lowers its confidence on answers a model gives without using the figure, separating detected non-grounding from a fluent guess.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/confidence-estimatio…] indexed:0 read:1min 2026-08-10 ·