Real software issues resolved

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]


83.5 per cent

Held by Claude Opus 4.7. Set 2026-04-16, and has stood for 166 days. Written down here 2026-09-07.

SWE-bench takes real bug reports from real open-source projects, hands a model the repository exactly as it stood before anyone fixed the problem, and checks whether the patch it writes makes the project's own tests pass. Nothing about it is simulated. The Verified subset is the portion human engineers confirmed is genuinely solvable from the information supplied. Two organisations publish a number for this and they do not agree: on the day this was written the official leaderboard said 79.2 per cent and Epoch AI's own evaluation said 83.5. Neither is dishonest; they run different harnesses. This page follows Epoch and says so, because the alternative is a number that moves whenever somebody picks the more flattering source.

That description was written when this frontier was opened, on 2026-09-07, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.


What counts
The best score on SWE-bench Verified as evaluated and published by Epoch AI. The publisher is named deliberately: the official SWE-bench leaderboard runs a different harness and reports a different figure, and a frontier with two publishers and no named one cannot be checked.
What moves it
Epoch AI records a higher score on the same benchmark.

Where to check it: https://epoch.ai/data/ai-benchmarking-dashboard

The program reads it from https://epoch.ai/data/benchmarks.csv, which is the same figure in a form a script can parse.


Last checked: 2026-09-28


Previous: Research mathematics solved   Next: Questions answered without inventing an answer

[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]