FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity A new arXiv paper, arXiv:2609.07153v1, introduces FreqBLiMP, a frequency-controlled extension of the BLiMP benchmark that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, the authors find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood but only a modest reduction in overall contrastive acceptability accuracy. That aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena requiring lemma-specific information. arXiv:2609.07153v1 Announce Type: cross Abstract: Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models LMs prefer acceptable sentences over minimally different unacceptable ones. However, these benchmarks largely ignore lexical frequency variation, despite lexical frequency being a pervasive and highly skewed property of natural language use. Consequently, existing evaluations do not test whether grammatical preferences remain stable when contrasts involve rare lexical items. We introduce FreqBLiMP, a frequency-controlled extension of BLiMP that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal-pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, we find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood, but only a modest reduction in overall contrastive acceptability accuracy. However, this aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena that require lemma-specific information.