All of Human Cooking in 2MB: What Radical Compression Says About Knowledge
Imagine the entire culinary inheritance of humanity — every braise, every fermentation, every grandmother's pasta sauce — squeezed into a file smaller than a single phone photograph. Two megabytes. Less data than the email thread you ignored this morning.
That provocation has been making the rounds online, and while it sounds absurd at first, it points to something genuinely interesting happening in computer science: our definitions of what counts as "storing knowledge" are being rewritten. Compression, AI models, and even biology are colliding in ways that make the old gigabyte-counting instinct feel quaint.
The difference between data and knowledge
A high-resolution scan of a single cookbook page might run to several megabytes. A video of someone kneading dough can easily eat a gigabyte. So how could the entirety of cooking fit in 2MB?
The answer is that "cooking" — as a body of human practice — isn't really about pixels or frames. It's about a relatively small number of techniques (sear, simmer, emulsify, ferment, bake) applied to a finite pantry of ingredients, governed by rules of chemistry, heat, and time. Strip away the photography, the prose, the personality, and what's left is closer to a compact rulebook than a library.
Two megabytes is roughly two million characters of text — about 20 average novels' worth of plain words. That's plenty of room for the structural grammar of cooking. What it can't store is the lived experience: the smell of caramelising onions, the muscle memory of folding pastry, the cultural weight of a dish.
This is the key distinction the 2MB thought experiment surfaces. Compression doesn't preserve knowledge; it preserves structure. The richness has to be reconstructed by whoever — or whatever — reads it.
Why AI changes the compression equation
For most of computing history, compression was a mathematical exercise: find redundancy, replace it with shorter codes, reverse the process to get the original back. ZIP files, JPEGs and MP3s all work this way. The ceiling was set by information theory — you couldn't squeeze below the actual entropy of the data.
Large language models have quietly broken that assumption. A modern AI model is, in a sense, a lossy compression of a huge chunk of the written internet. Ask it for a chocolate chip cookie recipe and it will produce one — not by retrieving a stored file, but by reconstructing something plausible from learned patterns. The "recipe" was never explicitly saved; it emerged from billions of tiny statistical relationships.
This is why the "2MB of cooking" idea isn't quite the joke it sounds. If you had a sufficiently smart decoder — a model that already understood chemistry, language, and culture — you wouldn't need to send it recipes at all. You'd send it the differences that make cuisines distinct, and it would generate the rest. The decoder carries the weight; the file just nudges it in the right direction.
That's a profound shift. Storage is no longer about preserving every byte. It's about preserving the smallest set of instructions that a capable reader can expand into something useful.
DNA: the other extreme of density
The compression conversation usually stops at silicon, but it doesn't have to. As Science magazine reported, researchers have shown that DNA could store all of the world's data in a single room. Every photo, every film, every database on Earth, encoded into the same four-letter alphabet — A, T, C, G — that already runs every cell in your body.
DNA is staggeringly dense. A gram of it can theoretically hold hundreds of petabytes. It is also, by accident of evolution, the most durable storage medium we know of: scientists have read genetic material from animals that died tens of thousands of years ago. Compare that to a hard drive, which is doing well to last a decade.
Put the two ideas together — AI-driven semantic compression and DNA as a physical substrate — and the 2MB cooking file looks less whimsical. You could store the structural rulebook of human cuisine in a smear of synthesised DNA the size of a punctuation mark. The bottleneck wouldn't be space. It would be writing and reading speed, which remain slow and expensive.
What gets lost when we compress culture
For an Australian audience, the cultural stakes here are worth dwelling on. Australian cooking is, more than most national cuisines, a layered import: Indigenous knowledge of native ingredients going back tens of thousands of years, British colonial habits, post-war Mediterranean migration, then waves of South-East Asian, Middle Eastern and African influence. A pho here is not a pho in Hanoi. A lamington recipe varies between Queensland and Victoria.
A 2MB "everything cookbook" would, almost by definition, be a flattening. Compression algorithms — whether mathematical or neural — reward the average and penalise the unusual. The more idiosyncratic a tradition is, the more bits it costs to encode, and the more likely it is to be smoothed away into something more generic.
This is already a known problem with large AI models. They reproduce the mainstream confidently and the marginal poorly. A model trained mostly on American food blogs will give you a perfect mac-and-cheese and a vague, faintly wrong damper. Compression is never neutral; it always reflects the priorities of whoever designed the codec.
The real lesson of the 2MB cookbook
So can all of human cooking actually fit in 2MB? Honestly, no — not in any form that does the subject justice. But the question is more useful than the answer.
It forces us to think about three things at once:
- How much of what we call "information" is really redundancy waiting to be stripped out by a smart enough encoder.
- How much of meaning lives in the reader, not the file — a point AI has made unavoidable.
- How storage media themselves are about to change, with DNA and other molecular approaches promising densities that make today's data centres look like warehouses full of filing cabinets.
The 20th century treated knowledge as something you accumulated: more books, more archives, more servers. The 21st is starting to treat it as something you distil. The interesting unit isn't the gigabyte anymore; it's the kilobyte that, fed into the right model, expands into a world.
Whether that's a triumph or a quiet cultural loss depends on what gets put into the 2MB — and, crucially, what doesn't. The recipe for your grandmother's Christmas pudding, with its specific brandy and its specific argument about whether to add carrot, is exactly the kind of detail that compression hates. Keeping it alive will probably require something stubbornly old-fashioned: writing it down, in full, and cooking it with someone who remembers.