Assembly mistakes spotted in a photograph

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]


83.3 per cent

Held by Claude Opus 5.5. Set 2026-09-22, and has stood for 7 days. Written down here 2026-09-27.

Everything else on this page asks a machine to do something machines were built for - arithmetic, search, symbols, a board with sixty-four squares on it. This one hands it a photograph of a half-built shoe cabinet and asks what went wrong. The benchmark was assembled by buying three pieces of IKEA furniture of increasing difficulty and photographing the builds as they went: some of the photographs show a correct build, and in the others a realistic mistake was made deliberately, often with assembly carried on past it so that the error is visible but is not the step just finished. The model gets the official manual as a PDF, a tool for zooming into the image and a Python interpreter, and it has to connect a technical diagram to a real object in a room - which is what fixing an appliance or a car actually requires, and which almost nothing else tests. The grading is deliberately lenient: name the right step, describe the problem roughly, and it counts. So this is the generous reading of the ability rather than the strict one, and the best score has still only just passed four fifths, while plenty of models released in the very same week sit nearer a fifth. That spread is the reason this is worth a page. It is not a skill the field has broadly; it is a skill two or three models have, this year, for the first time. In 1993, the year this page is pretending to be from, machine vision meant reading typed digits off a cheque under a controlled lamp.

That description was written when this frontier was opened, on 2026-09-27, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.


What counts
The best score any model has recorded on Epoch AI's Furniture Assembly benchmark, as published in Epoch AI's own benchmark file. A sample is a photograph of a part-built piece of IKEA furniture, handed to the model together with the official assembly manual; the model must say whether the build is correct so far and, where it is not, name the step the mistake was made in and describe it. The score is the share of samples it gets right. One rule belongs on the number rather than under it: the file carries a task version, scores from different versions are not comparable, and every row read here is version 1.0.0 - if a later version ever appears in that column this frontier needs a decision before it needs a reading.
What moves it
A model records a higher score on the same benchmark and Epoch AI publishes it.

Where to check it: https://epoch.ai/benchmarks/furniture-assembly

The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.


Last checked: 2026-09-27


Previous: The most storage operations ever done in a second   Next: The longest task an AI can finish on its own

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]