[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]
1,045 minutes
Held by claude-mythos-preview-early. Set 2026-04-07, and has stood for 175 days. Written down here 2026-09-07.
Give a model a job, and time how long the same job takes a competent human. Models finish short jobs and fail long ones, so the length at which a model succeeds half the time is a single number describing how much work it can be left alone with. In 2019 that number was six seconds. It has been doubling roughly every four months ever since, and it has now passed seventeen hours, which is longer than a working day. The measure comes from METR, which evaluates models for exactly this purpose. It says nothing about whether the work was any good, only that it was finished. That is a real limitation and it is still the closest thing anyone publishes to a speedometer for the field.
That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.
Where to check it: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
The program reads it from https://metr.org/assets/benchmark_results_1_1.yaml, which is the same figure in a form a script can parse.
Last checked: 2026-09-28
Previous: Assembly mistakes spotted in a photograph Next: Open problems in mathematics solved
[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]