[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]
THIS ONE IS FINISHED. It stopped being a frontier on 2026-09-07: saturated. The page stays, because what it recorded still happened.
Above ninety per cent, while tier four of the same set stands at 47.9. The tier, not the test, was the frontier.
93.7 per cent
Held by GPT-6 Astra. Set 2026-09-03. Written down here 2026-09-07.
The lower three tiers of the same private research-mathematics set whose top tier is still open, and still under half. Problems at these tiers take a graduate student somewhere between minutes and hours. Models now clear ninety per cent of them. The same examination, one tier higher, remains the hardest thing on this site - which is the clearest demonstration available that a benchmark is not one thing, and that solved is a question about where somebody drew the line.
That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.
Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard
The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.
Previous: Olympiad qualifying papers Next: The most arithmetic per watt
[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]