cd /news/artificial-intelligence/cracks-in-the-foundation-seemingly-m… · home topics artificial-intelligence article
[ARTICLE · art-93005] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

A new arXiv preprint (arXiv:2608.10296v1) reveals that four minor architectural choices in dense transformer models—normalization, GQA, pretraining context length, and sliding window attention—can compound to reduce long-context performance by up to 47% when combined, despite having little effect on short-context metrics. The authors trained over 170,000 GPU hours and released OlmPool, a set of 26 comparable 7B models, showing that some architectures outperform Llama 3 on long-context extensibility.

read1 min views1 publishedAug 12, 2026

arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set of four minor architectural decisions --- all made by at least one of the Olmo, Llama, and Qwen dense model families --- have a compoundingly negative effect on long context extensibility. Any one of these choices alone has a minor impact on long context performance, but combining three or more can drop the performance downstream by up to 47%. Furthermore, these differences are not detectable from short-context loss or validation datasets. We show that much of the variation in long context ability across model families is driven by these architectural features and detectable from applying context extension early in pretraining. We demonstrate this with controlled ablations that hold data, tokenizer, and extension recipe fixed while varying normalization, GQA, pretraining context length, and sliding window attention. After over 170,000 GPU hours of training, we release the resulting set of models as OlmPool, a set of 26 comparable 7B models with checkpoints before and after long-context extension. This pool includes several architectures that outperform the Llama 3 architecture on long context extensibility. In an analysis of our ablation models, we identify patterns in attention sink behavior and attention distributions across context that are attributable to specific architectural differences.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cracks-in-the-founda…] indexed:0 read:1min 2026-08-12 ·