{"slug": "toward-auditable-and-calibrated-ai-for-dementia-related-crash-severity-a-to", "title": "Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review", "summary": "A study on arXiv (2609.22694v1) reframes dementia-related crash severity prediction as a decision-aware triage problem, using 4,781 Texas crash records with structured fields and police narratives under a stratified 70/15/15 split. The leakage-controlled Gemma model achieved the highest observed macro-F1 of 0.545 (95% bootstrap CI [0.507, 0.583]), while the best calibrated fusion model reached macro-F1 of 0.522 with expected calibration error of 0.033. Selective deferral at 70% coverage raised macro-F1 to 0.573 and lowered severity cost to 0.577, with deferred cases routed to a proposed human-review process rather than further evaluated.", "body_md": "arXiv:2609.22694v1 Announce Type: new \nAbstract: Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.", "url": "https://wpnews.pro/news/toward-auditable-and-calibrated-ai-for-dementia-related-crash-severity-a-to", "canonical_source": "https://arxiv.org/abs/2609.22694", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:26:58.637247+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-safety", "ai-research", "large-language-models"], "entities": ["arXiv", "Gemma", "BERT", "Texas"], "alternates": {"html": "https://wpnews.pro/news/toward-auditable-and-calibrated-ai-for-dementia-related-crash-severity-a-to", "markdown": "https://wpnews.pro/news/toward-auditable-and-calibrated-ai-for-dementia-related-crash-severity-a-to.md", "text": "https://wpnews.pro/news/toward-auditable-and-calibrated-ai-for-dementia-related-crash-severity-a-to.txt", "jsonld": "https://wpnews.pro/news/toward-auditable-and-calibrated-ai-for-dementia-related-crash-severity-a-to.jsonld"}}