Research mathematics, lower tiers

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]


THIS ONE IS FINISHED. It stopped being a frontier on 2026-09-07: saturated. The page stays, because what it recorded still happened.

Above ninety per cent, while tier four of the same set stands at 47.9. The tier, not the test, was the frontier.


93.7 per cent

Held by GPT-6 Astra. Set 2026-09-03. Written down here 2026-09-07.

The lower three tiers of the same private research-mathematics set whose top tier is still open, and still under half. Problems at these tiers take a graduate student somewhere between minutes and hours. Models now clear ninety per cent of them. The same examination, one tier higher, remains the hardest thing on this site - which is the clearest demonstration available that a benchmark is not one thing, and that solved is a question about where somebody drew the line.

That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.


What counts
The best score any model has recorded on the private tiers one to three of FrontierMath version two, as published by Epoch AI.
What moves it
Nothing. It is finished.

Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard

The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.


Previous: Olympiad qualifying papers   Next: The most arithmetic per watt

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]