cd /news/large-language-models/a-diversity-diet-for-a-healthier-mod… · home › topics › large-language-models › article
[ARTICLE · art-40517] src=aclanthology.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT

Researchers at French institutions found that diversity-driven sampling can reduce pre-training dataset size by up to 94% and training time by 73% while maintaining performance in ModernBERT models. In some tasks, diversity sampling improved model quality by up to 10 points over random sampling.

read2 min views31 publishedJun 22, 2026
A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT
Image: Aclanthology (auto-discovered)
[A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT](https://aclanthology.org/2026.findings-acl.1707.pdf)

[Louis Estève](https://aclanthology.org/people/louis-esteve/),
[Christophe Servan](https://aclanthology.org/people/christophe-servan-7075/),
[Thomas Lavergne](https://aclanthology.org/people/thomas-lavergne/unverified/),
[Agata Savary](https://aclanthology.org/people/agata-savary/)
Abstract

Diversity has been gaining interest in the NLP community in recent years. At the same time, state-of-the-art transformer models such as ModernBERT use very large pre-training datasets, which are driven by size rather than by diversity. This summons to investigate theimpact of diversity on pre-training. We do so in this study, with the express intent of reducing pre-training dataset size, while retaining atleast comparable performance. We compare diversity-driven sampling algorithms, and we use the best one to pre-train several ModernBERT models on French with a fixed compute budget. We fine-tune and evaluate them on a variety of French benchmarks. We compare them with models pre-trained on randomly sampled data of commensurate size, with the same compute budget. We find that both random and diversity-driven sampling may reduce the pre-training dataset by up to 94% and the pre-training time by up to 73% while maintaining performance. Moreover, in some tasks, the inherent quality of models, estimated via head-only fine-tuning, is up to 10 points higher with diversity sampling than with random sampling.- Anthology ID:

- 2026.findings-acl.1707
- Volume:
[Findings of the Association for Computational Linguistics: ACL 2026](https://aclanthology.org/volumes/2026.findings-acl/)- Month:
[Findings](https://aclanthology.org/venues/findings/)- SIG:
- Publisher:
  • Association for Computational Linguistics
- Note:
- Pages:
  • 34168–34181
- Language:
- URL:
[https://aclanthology.org/2026.findings-acl.1707/](https://aclanthology.org/2026.findings-acl.1707/)- DOI:
- Cite (ACL):
[A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT](https://aclanthology.org/2026.findings-acl.1707/)(Estève et al., Findings 2026)- PDF:
[https://aclanthology.org/2026.findings-acl.1707.pdf](https://aclanthology.org/2026.findings-acl.1707.pdf)
── more in #large-language-models 4 stories · sorted by recency
── more on @louis estève 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-diversity-diet-for…] indexed:0 read:2min 2026-06-22 · —