{"slug": "language-specificity-vs-domain-diversity-benchmarking-transformers-for-bangla", "title": "Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER", "summary": "Fine-tuned XLM-RoBERTa achieved an F1-score of 0.5959 on Bangla medical named entity recognition, setting a new state of the art above the previously reported best of 0.5848, according to an arXiv benchmark evaluating BanglaBERT, mBERT, XLM-RoBERTa and GPT-4o mini across the full 3,179-sample test set. The language-specific BanglaBERT model underperformed its multilingual counterparts with an F1-score of 0.4937, indicating pretraining domain diversity can outweigh language specificity in specialized clinical settings, while fine-tuned transformers beat the best prompting configuration by a factor of 3.76. Per-entity analysis showed Medicine and Specialist categories above 0.83 F1, with Symptom the hardest at 0.4367 despite being the most frequent training class.", "body_md": "arXiv:2609.29101v1 Announce Type: new \nAbstract: Medical Named Entity Recognition (NER) for low-resource languages remains a challenging task due to high linguistic variability and a scarcity of domain-specific annotated corpora. This work presents a comprehensive empirical benchmark evaluating three fine-tuned transformer encoders-BanglaBERT, multilingual BERT (mBERT), and XLM-RoBERTa-against GPT-4o mini under zero-shot and few-shot prompting configurations for Bangla medical NER. In contrast to prior studies that evaluated large language models on limited subsets of only 50 samples, we conduct a large-scale evaluation across the full test set of 3,179 samples, providing statistically robust and reproducible baselines. Our fine-tuned XLM-RoBERTa model achieves an F1- score of 0.5959, establishing a new state-of-the-art and surpassing the previously reported best result of 0.5848. Crucially, we demonstrate that the language-specific BanglaBERT model consistently underperforms its multilingual counterparts with an F1-score of 0.4937, indicating that pretraining domain diversity can outweigh language specificity in highly specialized clinical settings. Furthermore, we present a detailed per-entity-type analysis for this task, revealing that Medicine and Specialist categories are recognized with high reliability, achieving F1- scores above 0.83, while the Symptom category remains the most challenging with an F1-score of 0.4367 despite being the most frequent training class. Finally, fine-tuned transformer models outperform the optimal prompting configuration by a factor of 3.76, confirming that prompt-only pipelines remain inadequate for structured clinical entity extraction in low-resource language environments.", "url": "https://wpnews.pro/news/language-specificity-vs-domain-diversity-benchmarking-transformers-for-bangla", "canonical_source": "https://www.machinebrief.com/news/language-specificity-vs-domain-diversity-benchmarking-transf-ns7l", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 05:00:44.253578+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-research", "large-language-models"], "entities": ["XLM-RoBERTa", "BanglaBERT", "mBERT", "GPT-4o mini", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/language-specificity-vs-domain-diversity-benchmarking-transformers-for-bangla", "markdown": "https://wpnews.pro/news/language-specificity-vs-domain-diversity-benchmarking-transformers-for-bangla.md", "text": "https://wpnews.pro/news/language-specificity-vs-domain-diversity-benchmarking-transformers-for-bangla.txt", "jsonld": "https://wpnews.pro/news/language-specificity-vs-domain-diversity-benchmarking-transformers-for-bangla.jsonld"}}