[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]
96,096,261 bytes
Held by zmix 1.0. Set 2026-09-14, and has stood for 15 days. Written down here 2026-09-19.
Compression is prediction. To squeeze text you have to guess what comes next, and the better the guess, the fewer bits it costs to record whether you were right - so a contest to make a file small is a contest to model the language in it, which is why this benchmark says in its own first paragraphs that its goal is to encourage research in artificial intelligence rather than to find the best compressor. The file is the first billion bytes of an English Wikipedia dump from 2006, and it has not changed since, so results twenty years apart are directly comparable. What counts is the archive plus the program that unpacks it, which is the rule that makes the whole thing honest: a program cannot win by already knowing Wikipedia, because whatever it knows it has to carry. The leaders were hand-tuned mixtures of statistical models for the first decade, then recurrent networks, and are now transformers trained on the very file they are about to compress. In 1993 gzip would have got this file down to about a third of its size. It is under a tenth now, and every further percent has cost years.
That description was written when this frontier was opened, on 2026-09-12, and is never rewritten. Where it and the number above it disagree, the number is the reading and the description is history.
Where to check it: https://www.mattmahoney.net/dc/text.html
Last checked: 2026-09-27
Previous: The largest quantum volume Next: The most data ever read and written in a second
[ Front page ] [ The edge ] [ Finished ] [ About ] [ Links ]