cd /news/large-language-models/curveshift-is-agent-progress-scalar-… · home topics large-language-models article
[ARTICLE · art-85571] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

A new arXiv preprint (2608.00355v1) finds that apparent progress toward harder tasks in large language models is mostly a ceiling effect, but a smaller hard-task effect persists after controlling for overall ability. Using LiveCodeBench, models released after September 2024 gain about +0.40 logits on the hardest problems beyond predictions from easy and medium performance, raising the hard-problem solve rate from roughly 18% to 25%. The authors release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and analysis code.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/curveshift-is-agent-…] indexed:0 read:1min 2026-08-04 ·