The Next AI Breakthrough May Not Be a Bigger Model OpenAI's o3 model scored 87.5% on the ARC-AGI benchmark in December 2024, up from GPT-4o's 5%, by using 5.5 billion tokens of inference-time compute instead of scaling model parameters. The benchmark, created by researcher François Chollet in 2019, was designed to resist the strategy of making models bigger, which had failed for four years. Member-only story The Next AI Breakthrough May Not Be a Bigger Model Scaling parameters hit a wall. The replacement thinking longer at inference is already plateauing too. The actual leaps are hiding somewhere else. Read the article for free . here In 2019, a researcher named François Chollet built a benchmark called ARC-AGI. The puzzles were simple enough that a child could solve them. They were mostly colored grids where you had to find the pattern and fill in the missing square. He designed them specifically to defeat the one strategy AI labs were betting everything on, which was making the model bigger. For four years, that is what they did, and for four years, ARC-AGI did not move. GPT-3 scored close to zero. GPT-4o scored 5%. Then in December 2024, OpenAI showed a model called o3 that scored 87.5% on the same puzzles. It used no new training data and also belonged to the same model class as the one that had scored 5% the month before. The only thing it did differently was think longer. Where the older model spent a few hundred tokens deciding on an answer, o3 spent five and a half billion. The high-compute run was reportedly trained on a large portion of the public training set, so the leap is not pure inference, but the shape of the breakthrough still held.