How big are the frontier models? I tried to answer this question using a statistical model of intelligence indices, a No-CoT reasoning benchmark, API prices, compute trends, and a poll of 20 AI researchers and engineers. Here are the results:
- The motivation is simple: people speculate on this topic endlessly, but I've yet to see a proper attempt that is grounded in evidence. Intelligence indices offer a good starting point, however they measure RL as much as the base model. I decided to take the Gould et al. No-CoT
- glad I could contribute to the survey!