cd /news/artificial-intelligence/even-ai-cant-perfectly-build-ikea-fu… · home › topics › artificial-intelligence › article
[ARTICLE · art-140951] src=fastcompany.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Even AI can’t perfectly build Ikea furniture—yet

Epoch AI researchers Aiden Ament and Greg Burnham released the Furniture Assembly Benchmark, which tests AI models on spotting errors in partially assembled Ikea furniture, with Anthropic's Claude Opus 4.5 initially leading at 28% accuracy before OpenAI's GPT-6 Astra reached 80% within 10 months. The benchmark used 60 photos across three Ikea products — the Ställ shoe rack, Tonstad bed frame, and Gullaberg dresser — fed to models from OpenAI, Anthropic, Google, Moonshot, and Alibaba, with GPT-6 Astra averaging about 3 minutes per photo, up to 10 times faster than other models. The researchers say the task is a proxy for economically important visual and spatial reasoning work such as fixing a car or repairing household appliances.

by read4 min views3 publishedSep 28, 2026

Building Ikea furniture is not for the faint of heart. So much so that folklore suggests assembling a dresser, a bed, or anything from the Swedish brand can create enough tension it may even destroy a relationship.

But beyond the stuff’s being a pain to build, it turns out that the problem-solving and spatial reasoning involved in building Malm dressers or Billy bookshelves are remarkably good benchmarks for assessing just how intelligent AI models are.

Aiden Ament and Greg Burnham, researchers at Epoch AI, an artificial intelligence research nonprofit, recently released a report detailing the Furniture Assembly Benchmark, a metric developed to test the actual intelligence of AI models. It centers on how well the model can identify errors in partially assembled Ikea furniture.

At a time where tech companies love bragging about how intelligent their models are, it can be complicated to assess how “smart” each model actually is—particularly compared with others. Asking Chat GPT or Claude to spot errors in Ikea furniture assembly puts the models to the test in real-world applications. And while no model got it right 100% of the time, it does seem they are quickly improving.

The researchers first set out to develop a benchmark while on a company retreat, figuring out how to set up an experiment to test AI’s capabilities.

“First, we tried acting like ‘the AI’s robot’ where we asked it what to do and tried to follow its directions somewhat literally. The results were often messy,” Burnham tells Fast Company.

“I like the final idea better, which was due to my colleague Aiden Ament: asking it to spot mistakes,” Burnham says. “This reflects the reality that it’s OK to make mistakes midassembly so long as you can catch and correct them.”

Researchers used three Ikea products for the experiment, picking them based on varying degrees of difficulty according to Ikea’s own Complexity Index: the Ställ shoe rack for an easy level, Tonstad bed frame for a medium level, and the Gullaberg dresser as the most complex.

They went about building each piece of furniture, intentionally making mistakes along the way, and taking 60 pictures of various steps to feed them to the various AI models from OpenAI, Anthropic, Google, Moonshot, and Alibaba. In addition to the images, the researchers provided each AI model with a PDF of the assembly instructions, a tool to zoom in on each image, and a Python interpreter.

Each AI model was tasked with identifying what mistakes were made by the user and at what point during assembly, and was scored accordingly against the benchmark. Additionally, the researchers asked the model to provide users with an explanation of each mistake identified.

The benchmark isn’t just about building Ikea furniture. It’s intended to go beyond that, to measure how well these models could guide and troubleshoot more complex builds.

“If models can reason quickly and accurately enough in this domain, it is plausible that AI could soon provide reliable, real-time guidance for complex physical assembly and repair tasks,” Ament and Burnham wrote in their report.

“We think this challenge serves as a good proxy for a number of economically important tasks that require similar visual and spatial reasoning, like fixing a car or repairing household appliances,” they wrote.

Initially, the best scoring model on the Ikea test was Anthropic’s Claude Opus 4.5, scoring at just 28% accuracy. But within 10 months, OpenAI quickly caught up, with its GPT-6 Astra model reaching 80%. Among the most notable findings was just how quickly models could provide a response, with the winning GPT-6 Astra model taking around 3 minutes per photo, up to 10 times as fast as other models.

Chinese models seemed to particularly lag in progress, trailing behind U.S. models.

Still, while progress on spatial reasoning and troubleshooting within such a short time seems promising, the models still have limitations, particularly on speed and consistent accuracy.

“AI development is a very capital-intensive process, and AI companies growing their revenue is required for continued investment. So we’re always looking to benchmark areas where we don’t currently see much impact from AI,” Burnham tells Fast Company. “If AI can upskill factory workers to repair complex machinery, for instance, that could lead to less downtime.”

Still, rest assured, your job putting furniture kits together is safe—for now.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @epoch ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/even-ai-cant-perfect…] indexed:0 read:4min 2026-09-28 · —