Improving 25 square packing upper bounds through evolutionary LLM program search A developer used LLM-driven evolutionary program search, similar to AlphaEvolve, to break 25 records for packing n unit squares into the smallest enclosing square, including one record that had stood for 47 years. The search ran for 128 generations over two days, using 7.5k CPU-hours and $125 of tokens, with Claude Haiku agents mutating programs scored by distance from the best known packing. What is the smallest square that encloses n unit squares? For some values of n, the answer is clear. 16 squares pack into a 4x4. In contrast, here is the beautiful, recently proven to be optimal https://github.com/Queuingtheorydotcom/11SquaresFormalized , answer for n=11. Records for different values of n were historically maintained https://web.archive.org/web/20230530194618/https://erich-friedman.github.io/packing/squinsqu/ by Erich Friedman, and more recently https://kingbird.myphotos.cc/packing/squares in squares.html by David Ellsworth. Driven by a flurry of recent excitement around square packing, records are now also nicely maintained https://jlevy.github.io/squares/ by Joshua Levy. How did we break these records? TL;DR: We achieve these results https://github.com/ry-xu/square packing through LLM-driven evolutionary program search, similar to AlphaEvolve https://arxiv.org/abs/2506.13131 . The programs were seeded with primitives vibe-coded by yours truly, and the search was lightly guided between meetings and insufficient sleep. Motivating program search Following recent state-of-the-art approaches to solving open problems, I began by asking claude to break a record and then let it crunch away for an hour to no success. Seared into my brain from years of ML is to always look at the data . Looking at some of the resulting packings, it became apparent that the style of search was inefficient, often landing in uninteresting minima. To build some intuition, I vibe-coded a web app and decided to try for the records myself. Over the course of many hours and many packed squares, I requested various features that I thought may help the algorithm to break out of degenerate solutions, such as: - being able to round the corners of the squares - a control for shaking intensity - “wind” - restricting the shaking direction - “fade” - shaking less as the box shrinks This packing lab is memorialized here https://ryanxu.net/packing lab/ . While packing, I often found myself repeating certain actions that consistently led to interesting packings e.g. rounding the squares, then letting them sharpen over and over again . Eventually it occurred to me that I was simply executing a program on top of these vibe-coded toggles, and that an LLM could search through the space of programs to rediscover mine, and hopefully many better. The initial evolutionary search was implemented and within an hour, the record for n=51 had been broken on my little MacBook Air without a fan. Soon after, n=103 and n=105 had also fallen. The next morning, as I ran around sharing my packings with the office, I was kindly gifted a bit more compute by the boss https://x.com/Lifrordi and encouraged to take the day to scale up the approach. The run The run began with a seeded baseline program. This program places n squares at random in a box, then slowly shrinks the box. Whenever squares overlap, an optimizer L-BFGS finds a local arrangement where they no longer overlap. When this converges, we have a packing Alongside this code was a set of unused functions developed for the packing lab for future programs to use. Every generation, we tasked a fleet of claude haiku agents with generating mutations of the previous generation’s best programs. Best is decided by a fitness score—the distance from the best known packing for a fixed sample of n values. We ran this evolutionary loop for 128 generations. In the end, a total of 25 records were broken over the course of 2 days, using 7.5k CPU-hours and $125 of tokens. The oldest record broken had stood for 47 years. Although the records were broken in one continuous evolutionary search, a few modifications were made throughout: 1. Scaling up from 16 candidates per generation to 64 2. Sampling instructions from a set of roles for candidate generation, since the diversity of evolved programs appeared low 1. default: try to improve the program 2. crossover: merge two candidates 3. invent: create and use a new primitive 4. moonshot: make a high risk high reward change 3. Starting to include n values 200 4. Increasing max runtime 10s → 60s, as I noticed that large n were timing out 5. Upweighting unbeaten n in the fitness score In many cases, we broke our own records multiple times The evolved programs In the first half of the search, evolved programs kept the initial random search and squeeze engine, opting to add increasingly strong and diverse packing-refining algorithms on top of standard basin hopping. Once the fitness score was modified to include larger values of n, programs began to initialize squares in more structured patterns, nicely matching what real world approaches might look like. Much of the fun in working with RL or evolutionary algorithms is seeing what novel behaviors are discovered. Here we list a few: - Initially rounding the squares, then sharpening them as the box shrinks - Writing the trivial grid immediately, so that an incomplete packing is not penalized—some light reward hacking - Tricks during polishing - Moving a square into the most empty spot—some programs choose random squares, others choose edge squares - Swapping the angles or positions of two squares. - Better initialization - Initializing squares at nice angles such as atan 1/2 , atan 1/3 , atan 2/3 - Initializing with structures such as staircases or diagonal columns - Fixing squares during squeeze - For large n, fixing staircases and optimizing only the rest - Squares that cause overlaps during perturbations are jittered less, leading to edges stabilizing fast In general, we see that the evolution agents had a tendency to add code and complexity. We see an explosion of program diversity around generation 20 after expanding the set of evolution prompts. Around ten years ago I spent an embarrassing number of hours trying to pack n=17 on paper. While I’m unsure what the future of human involvement in math looks like, I’m content knowing that this push began with an insight from a human in a coffee shop while packing squares in squares. I’d like to thank Mamacoffee Vodičkova, all the square packers out there, and Equilibre Technologies https://equilibre.ai/ for the time and a slice of the company cluster.