[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]
72.0 per cent
Held by GPT-6 Astra. Set 2026-09-03, and has stood for 26 days. Written down here 2026-09-07.
A chess puzzle has one winning line and it can be checked absolutely, which makes it a clean test with nothing to argue about in the marking. These are positions a strong club player would be expected to find. Language models are not chess engines and do not search the way one does; a purpose-built engine has handled positions like these to a standard no human has matched since the 1990s. The interesting part is that a machine built to predict text manages seven in ten anyway, with no board, no search tree, and no design intent pointing at chess at all. This page keeps the number partly because it is an unusually fair test, and partly because a homepage from 1995 would have had a chess page on it.
That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.
Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard
The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.
Last checked: 2026-09-29
Previous: Questions answered without inventing an answer Next: The largest training run ever disclosed
[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]