[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]
75.6 per cent
Held by GPT-6 Astra. Set 2026-09-03, and has stood for 26 days. Written down here 2026-09-07.
SimpleQA asks short factual questions that have one correct answer, and scores whether the model gets them right. The interesting half is what happens when it does not know. A model can be wrong, or it can say it does not know, and this test rewards the second. That makes it the closest public measure of the failure people actually complain about: a machine producing a fluent, confident sentence with nothing behind it. Three quarters is a long way from where this needs to be for anything consequential, and it has climbed more slowly than the reasoning scores elsewhere on this page. Solving research mathematics and still inventing a date is not a contradiction. They are separate skills and they are improving at different rates.
That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.
Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard
The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.
Last checked: 2026-09-29
Previous: Real software issues resolved Next: Chess puzzles solved
[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]