Member-only story
Scaling parameters hit a wall. The replacement (thinking longer at inference) is already plateauing too. The actual leaps are hiding somewhere else. #
Read the article for free
[.]here In 2019, a researcher named François Chollet built a benchmark called ARC-AGI. The puzzles were simple enough that a child could solve them. They were mostly colored grids where you had to find the pattern and fill in the missing square. He designed them specifically to defeat the one strategy AI labs were betting everything on, which was making the model bigger. For four years, that is what they did, and for four years, ARC-AGI did not move. GPT-3 scored close to zero. GPT-4o scored 5%.
Then in December 2024, OpenAI showed a model called o3 that scored 87.5% on the same puzzles. It used no new training data and also belonged to the same model class as the one that had scored 5% the month before. The only thing it did differently was think longer. Where the older model spent a few hundred tokens deciding on an answer, o3 spent five and a half billion. The high-compute run was reportedly trained on a large portion of the public training set, so the leap is not pure inference, but the shape of the breakthrough still held.