cd /news/natural-language-processing/language-specificity-vs-domain-diver… · home › topics › natural-language-processing › article
[ARTICLE · art-139477] src=machinebrief.com ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

Fine-tuned XLM-RoBERTa achieved an F1-score of 0.5959 on Bangla medical named entity recognition, setting a new state of the art above the previously reported best of 0.5848, according to an arXiv benchmark evaluating BanglaBERT, mBERT, XLM-RoBERTa and GPT-4o mini across the full 3,179-sample test set. The language-specific BanglaBERT model underperformed its multilingual counterparts with an F1-score of 0.4937, indicating pretraining domain diversity can outweigh language specificity in specialized clinical settings, while fine-tuned transformers beat the best prompting configuration by a factor of 3.76. Per-entity analysis showed Medicine and Specialist categories above 0.83 F1, with Symptom the hardest at 0.4367 despite being the most frequent training class.

by read1 min views1 publishedSep 25, 2026

arXiv:2609.29101v1 Announce Type: new Abstract: Medical Named Entity Recognition (NER) for low-resource languages remains a challenging task due to high linguistic variability and a scarcity of domain-specific annotated corpora. This work presents a comprehensive empirical benchmark evaluating three fine-tuned transformer encoders-BanglaBERT, multilingual BERT (mBERT), and XLM-RoBERTa-against GPT-4o mini under zero-shot and few-shot prompting configurations for Bangla medical NER. In contrast to prior studies that evaluated large language models on limited subsets of only 50 samples, we conduct a large-scale evaluation across the full test set of 3,179 samples, providing statistically robust and reproducible baselines. Our fine-tuned XLM-RoBERTa model achieves an F1- score of 0.5959, establishing a new state-of-the-art and surpassing the previously reported best result of 0.5848. Crucially, we demonstrate that the language-specific BanglaBERT model consistently underperforms its multilingual counterparts with an F1-score of 0.4937, indicating that pretraining domain diversity can outweigh language specificity in highly specialized clinical settings. Furthermore, we present a detailed per-entity-type analysis for this task, revealing that Medicine and Specialist categories are recognized with high reliability, achieving F1- scores above 0.83, while the Symptom category remains the most challenging with an F1-score of 0.4367 despite being the most frequent training class. Finally, fine-tuned transformer models outperform the optimal prompting configuration by a factor of 3.76, confirming that prompt-only pipelines remain inadequate for structured clinical entity extraction in low-resource language environments.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @xlm-roberta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/language-specificity…] indexed:0 read:1min 2026-09-25 · —