arXiv:2610.03482v1 Announce Type: new Abstract: Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, using Gaokerena-V and Gaokerena-R as two previously developed backbones. We first examine response variability on a 168-question Iranian medical entrance examination and observe substantially lower five-run consistency for Gaokerena-V than for Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. We then construct two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and train lightweight claim-level heads on frozen backbone attention maps and token probabilities. On held-out test splits, the heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The heads require neither retrieval nor repeated sampling at inference time. These results provide an initial study of single-pass claim-level uncertainty estimation for Persian medical language models; the test splits are small and the labels are automatically generated.
Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models
Researchers adapted the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, training lightweight claim-level hallucination detection heads on frozen backbone attention maps and token probabilities that achieved PR-AUCs of 0.4820 and 0.4652 — 2.30 and 2.66 times their respective random baselines — and ROC-AUCs of 0.7852 and 0.7810 on held-out test splits. The study used Gaokerena-V and Gaokerena-R as backbones and built two paired Persian claim-level hallucination datasets of 1,600 responses each, after observing substantially lower five-run consistency for Gaokerena-V than Aya-Expanse-8B on a 168-question Iranian medical entrance examination. The heads require neither retrieval nor repeated sampling at inference time, though the authors note the test splits are small and labels automatically generated.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.