cd /news/large-language-models/freqblimp-frequency-controlled-minim… · home topics large-language-models article
[ARTICLE · art-125490] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

A new arXiv paper, arXiv:2609.07153v1, introduces FreqBLiMP, a frequency-controlled extension of the BLiMP benchmark that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, the authors find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood but only a modest reduction in overall contrastive acceptability accuracy. That aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena requiring lemma-specific information.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.07153v1 Announce Type: cross Abstract: Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones. However, these benchmarks largely ignore lexical frequency variation, despite lexical frequency being a pervasive and highly skewed property of natural language use. Consequently, existing evaluations do not test whether grammatical preferences remain stable when contrasts involve rare lexical items. We introduce FreqBLiMP, a frequency-controlled extension of BLiMP that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal-pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, we find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood, but only a modest reduction in overall contrastive acceptability accuracy. However, this aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena that require lemma-specific information.

── more in #large-language-models 4 stories · sorted by recency
── more on @freqblimp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/freqblimp-frequency-…] indexed:0 read:1min 2026-09-10 ·