[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]
83.5 per cent
Held by Claude Opus 4.7. Set 2026-04-16, and has stood for 166 days. Written down here 2026-09-07.
SWE-bench takes real bug reports from real open-source projects, hands a model the repository exactly as it stood before anyone fixed the problem, and checks whether the patch it writes makes the project's own tests pass. Nothing about it is simulated. The Verified subset is the portion human engineers confirmed is genuinely solvable from the information supplied. Two organisations publish a number for this and they do not agree: on the day this was written the official leaderboard said 79.2 per cent and Epoch AI's own evaluation said 83.5. Neither is dishonest; they run different harnesses. This page follows Epoch and says so, because the alternative is a number that moves whenever somebody picks the more flattering source.
That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.
Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard
The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.
Last checked: 2026-09-28
Previous: Research mathematics solved Next: Questions answered without inventing an answer
[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]