Hard competition mathematics

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]


THIS ONE IS FINISHED. It stopped being a frontier on 2026-09-07: saturated. The page stays, because what it recorded still happened.

Above ninety-eight per cent. The remaining errors are closer to marking noise than to mathematics.


98.1 per cent

Held by GPT-5. Set 2025-08-07. Written down here 2026-09-07.

The hardest fifth of a competition mathematics set which was, for several years, the standard measure of whether a language model could do mathematics at all. It is now answered very nearly perfectly. What survives of it is a useful date: the moment competition mathematics stopped being a frontier and became a formality.

That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.


What counts
The best score any model has recorded on the level-five subset of the MATH set, as published by Epoch AI.
What moves it
Nothing. It is finished.

Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard

The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.


Previous: Graduate-level science questions   Next: Olympiad qualifying papers

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]