Epoch AI's Furniture Assembly Benchmark: GPT-6 Astra spots 80 percent of IKEA assembly mistakes

On September 23, 2026, Epoch AI published “Can AI spot mistakes in IKEA assembly?” by Aiden Ament and Greg Burnham, introducing the Furniture Assembly Benchmark (FAB). Each of its 60 items is a photo of a partially assembled IKEA product, drawn from three builds of differing complexity. The model gets the assembly manual, a zoom tool and a Python interpreter, and must find every misstep in the photo and describe it. Grading is staged: did it identify the right errors, did it name the right step numbers, and does an LLM validator accept the description.

Progress over ten months was steep. The best score in November 2025 was 28 percent, from Claude Opus 4.5. In September 2026 GPT-6 Astra scored 80 percent, ahead of Claude Fable 5.1 at 70 percent and Claude Opus 5 at 61 percent, with a median of about three minutes per image. Epoch estimates Chinese open-weight models trail the frontier on this task by roughly seven months, and notes that several lacked image input altogether.

Why it matters: most headline benchmarks are text, code or math. Checking a physical build against instructions is a grounded visual reasoning task close to real inspection work, and the jump from 28 to 80 percent in under a year is a datapoint that perception of the physical world is improving as fast as other skills.

What it does not show: 60 photos from three pieces of furniture is a small test set, and Epoch says it is unclear how far the result generalizes to other physical tasks. Spotting an error in a still image is also far from doing the assembly.

Sources

Last verified October 5, 2026