# General capability - and capabilities generally - have no good y-axis

> Source: <https://www.lesswrong.com/posts/RxfTG5jcHH3azKQTA/general-capability-and-capabilities-generally-have-no-good-y>
> Published: 2026-08-04 14:25:44+00:00

**BLUF:**

- To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big leap/small increment’). We need a good y-axis: an interval scale of AI capability which means +1 unit always represents the same degree of ‘how much better’, in the same way +1 degree Celsius is always the same amount of ‘how much hotter’.
- Yet there is no good y-axis for AI capability. All our measures are of something related-to but clearly not identical-with it, thus ‘true’ AI capability can be a funhouse-mirror reflection of whatever was measured. Specifically:
- Benchmark score: One small step in benchmark score can be a giant leap in capability, or the opposite, or whatever else. (My 6/10 vs. your 4/10 ≠ I’m 50% better at maths than you).
- Elo et al: Can give a real y-axis in terms of winning chances, but doesn’t translate outside of beating others. (Going from 50% to 73% to 88% chance to get a higher score than you on a maths test ≠ gaining 0 → 1 → 2 units of maths ability over you)
- Epoch Capabilities Index: Analogous to IQ, so the y-axis intervals are dictated by modelling assumptions (IQ 115 → 130 not really the same increment of smarter as IQ 85 → 100). In any case, latent trait(propensity to get high scores across benchmarks) still a funhouse mirror of Capability(generally). Benchmarks are substantially endogenous to the models, so (e.g.) the 2024 acceleration observed in both ECI and TH may be explained by common measurement artefact rather than mutual corroboration.
- Log-loss/prediction: Analogous to reaction time, so in the same way reaction time/digit span/vocab size is non-linear in human intelligence, prediction accuracy non-linear in AI.
- METR Time Horizon: Measured time horizon ~ 10^(k * total score on METR task suite), likely explained by human task-completion psychometrics. If TH is linear in AI capability, then Opus 4.5 → 4.6 is a bigger advance than dawn-of-time → Opus 4.5.

- Perhaps talk of ‘AI capability’ is better deflated, or maybe we await the theory which could do to intelligence what thermodynamics managed for temperature. Either way, our current measurements of AI are numerical gestures toward, not readings of, whatever is really going on.

# Introduction

Consider these two graphs:

These graphs paint very different pictures of AI progress: the shallow straight line of the ECI plot suggests steady incremental improvement; the (supra?)exponential sweep upwards for time horizons suggests screaming towards the singularity. Both are used (perhaps more than the researchers behind them would like) as summaries of AI in general. Yet which picture is, for want of a better term, right? Is the true (functional) form of AI capabilities best captured by linear-ish ECI or exponentially-increasing time horizons - or maybe something else?

I provide a counsel of despair. These two graphs are essentially the same picture with a transformed y-axis, and neither really measures AI capability. For what the right y-axis is, and the right transform of our measurements to get it, I only have varieties of scare-quotes and question-marks to offer.

[[Continued](https://forum.effectivealtruism.org/posts/CQvdadxjCpd7i7kjA/general-capability-and-capabilities-generally-have-no-good-y)...]

[Discuss](https://www.lesswrong.com/posts/RxfTG5jcHH3azKQTA/general-capability-and-capabilities-generally-have-no-good-y#comments)
