arXiv:2609.22694v1 Announce Type: new Abstract: Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.
Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review
A study on arXiv (2609.22694v1) reframes dementia-related crash severity prediction as a decision-aware triage problem, using 4,781 Texas crash records with structured fields and police narratives under a stratified 70/15/15 split. The leakage-controlled Gemma model achieved the highest observed macro-F1 of 0.545 (95% bootstrap CI [0.507, 0.583]), while the best calibrated fusion model reached macro-F1 of 0.522 with expected calibration error of 0.033. Selective deferral at 70% coverage raised macro-F1 to 0.573 and lowered severity cost to 0.577, with deferred cases routed to a proposed human-review process rather than further evaluated.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.