cd /news/artificial-intelligence/general-capability-and-capabilities-… · home topics artificial-intelligence article
[ARTICLE · art-86280] src=lesswrong.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

General capability - and capabilities generally - have no good y-axis

A new essay argues that no existing metric provides a valid interval scale for AI capability, meaning claims about exponential progress or stagnation lack a reliable y-axis. The author, writing on the Effective Altruism Forum, contends that benchmark scores, Elo ratings, the Epoch Capabilities Index, log-loss, and METR's Time Horizon all measure proxies that can distort true capability, and that the apparent acceleration in 2024 may stem from measurement artifacts. The piece concludes that current AI measurements are 'numerical gestures toward, not readings of' actual capability, and calls for a theory akin to thermodynamics for temperature.

read3 min views1 publishedAug 4, 2026

BLUF:

  • To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big leap/small increment’). We need a good y-axis: an interval scale of AI capability which means +1 unit always represents the same degree of ‘how much better’, in the same way +1 degree Celsius is always the same amount of ‘how much hotter’.

  • Yet there is no good y-axis for AI capability. All our measures are of something related-to but clearly not identical-with it, thus ‘true’ AI capability can be a funhouse-mirror reflection of whatever was measured. Specifically:

  • Benchmark score: One small step in benchmark score can be a giant leap in capability, or the opposite, or whatever else. (My 6/10 vs. your 4/10 ≠ I’m 50% better at maths than you).

  • Elo et al: Can give a real y-axis in terms of winning chances, but doesn’t translate outside of beating others. (Going from 50% to 73% to 88% chance to get a higher score than you on a maths test ≠ gaining 0 → 1 → 2 units of maths ability over you)

  • Epoch Capabilities Index: Analogous to IQ, so the y-axis intervals are dictated by modelling assumptions (IQ 115 → 130 not really the same increment of smarter as IQ 85 → 100). In any case, latent trait(propensity to get high scores across benchmarks) still a funhouse mirror of Capability(generally). Benchmarks are substantially endogenous to the models, so (e.g.) the 2024 acceleration observed in both ECI and TH may be explained by common measurement artefact rather than mutual corroboration.

  • Log-loss/prediction: Analogous to reaction time, so in the same way reaction time/digit span/vocab size is non-linear in human intelligence, prediction accuracy non-linear in AI.

  • METR Time Horizon: Measured time horizon ~ 10^(k * total score on METR task suite), likely explained by human task-completion psychometrics. If TH is linear in AI capability, then Opus 4.5 → 4.6 is a bigger advance than dawn-of-time → Opus 4.5.

  • Perhaps talk of ‘AI capability’ is better deflated, or maybe we await the theory which could do to intelligence what thermodynamics managed for temperature. Either way, our current measurements of AI are numerical gestures toward, not readings of, whatever is really going on.

Consider these two graphs:

These graphs paint very different pictures of AI progress: the shallow straight line of the ECI plot suggests steady incremental improvement; the (supra?)exponential sweep upwards for time horizons suggests screaming towards the singularity. Both are used (perhaps more than the researchers behind them would like) as summaries of AI in general. Yet which picture is, for want of a better term, right? Is the true (functional) form of AI capabilities best captured by linear-ish ECI or exponentially-increasing time horizons - or maybe something else?

I provide a counsel of despair. These two graphs are essentially the same picture with a transformed y-axis, and neither really measures AI capability. For what the right y-axis is, and the right transform of our measurements to get it, I only have varieties of scare-quotes and question-marks to offer.

[[Continued](https://forum.effectivealtruism.org/posts/CQvdadxjCpd7i7kjA/general-capability-and-capabilities-generally-have-no-good-y)...]

[Discuss](https://www.lesswrong.com/posts/RxfTG5jcHH3azKQTA/general-capability-and-capabilities-generally-have-no-good-y#comments)
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @effective altruism forum 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/general-capability-a…] indexed:0 read:3min 2026-08-04 ·