Interactive tool

One protein could take years to map. There are 200 million of them.

A protein's job is decided by the exact shape it folds into — and for fifty years, working out that shape was one of biology's hardest problems. Then DeepMind's AlphaFold predicted the structure of nearly every protein known to science in about a year, and gave the whole library away for free. This is the story of the problem, the breakthrough, and the hockey-stick moment when a fifty-year climb went vertical. Every figure and quote is real; sources are at the bottom.

01 — Shape is everything

A protein is a string that snaps into a machine.

Your body runs on proteins — they digest food, fight infection, carry oxygen, fire neurons. Each one starts life as a flat chain of chemical beads called amino acids (there are 20 kinds). In a fraction of a second, that chain folds itself into a precise three-dimensional shape, and that shape is what the protein does. Get the shape wrong and the machine jams — misfolded proteins are implicated in Alzheimer's, Parkinson's, and cystic fibrosis. The hard problem: given only the flat chain, what shape does it fold into?

A flat chain of amino acids — press Fold

The sequence is easy to read. The shape is not. Modern machines can read the amino-acid sequence of a protein quickly and cheaply. But the sequence is just a list of beads — it doesn't tell you how the string folds up in space, and the folded shape is the only thing that matters for how the protein behaves.

The beads pull on each other in thousands of tiny ways at once — some repel water, some attract each other, some carry charge — and the chain settles into the one shape that balances every force. Predicting that final shape from the sequence alone is the protein folding problem.

Water-avoiding Water-loving Charged
02 — The impossible search

Too many shapes to ever try them all.

In 1969 the molecular biologist Cyrus Levinthal pointed out something strange. A protein could fold into an astronomical number of possible shapes — yet real proteins find the right one in milliseconds. If a protein tried every possibility one at a time, it would take longer than the universe has existed. This is Levinthal's paradox. Drag the length below and watch the search space explode.

100 amino acids
Possible folded shapes
5.15 × 1047
assuming just 3 options per bead — Levinthal's own estimate
Time to try them all, one per picosecond
1.1 × 1018×
the current age of the universe
A 100-amino-acid protein has about 5 × 1047 possible shapes. Even flipping through a trillion of them every second, you'd need a billion billion times the age of the universe to check them all — yet the real protein folds in under a thousandth of a second.

Nature solves this instantly; computers could not. For decades, the only reliable way to learn a protein's shape was to stop calculating and start measuring it in a lab — atom by painstaking atom.

03 — Fifty years by hand

The old way: one protein, and often one whole PhD.

Before AI, biologists determined a protein's shape experimentally — growing crystals of it and firing X-rays through them, or freezing it for an electron microscope. The work is exacting and slow. A single structure could take months to years and cost hundreds of thousands of dollars; determining one protein was often enough to fill a researcher's entire doctorate.

Months–years
Per single structure
The experimental effort to pin down one protein's shape, using X-ray crystallography, NMR, or cryo-electron microscopy.
~200,000
Structures in ~50 years
The Protein Data Bank, founded in 1971, accumulated roughly 200,000 experimentally solved structures by 2023 — the sum of the entire field's work.
~0.1%
Of known proteins covered
Against 200 million known protein sequences, fifty years of experiments had mapped only a tiny sliver. The gap was widening, not closing.

Put those two numbers together and the scale of the problem is clear: at the pace of hands-on experiments, mapping every known protein wasn't a matter of years or decades. It was effectively never.

04 — The breakthrough

AlphaFold learned to predict the shape in minutes.

DeepMind trained a neural network on the ~170,000 structures biologists had painstakingly solved, teaching it to predict a folded shape directly from an amino-acid sequence. At the field's blind test in 2020, it didn't just compete — it effectively solved the 50-year problem, then scaled from a handful of proteins to hundreds of millions.

November 2020 · CASP14
A 50-year problem, effectively solved
At CASP — a biennial, blind competition where teams predict structures that have been solved experimentally but not yet published — AlphaFold2 scored a median 92.4 GDT (out of 100), with an average error of about 1.6 ångströms: roughly the width of a single atom. Its predictions were, in many cases, indistinguishable from real experiments. Organisers called the problem solved.
July 2021 · London & Hinxton
Open-sourced, and 350,000 structures released free
DeepMind published AlphaFold in Nature (a paper now cited over 40,000 times) and, with EMBL-EBI, launched the free AlphaFold Protein Structure Database — starting with the entire human proteome and ~350,000 structures in total.
July 2022 · The full catalogue
200 million proteins — nearly every one known to science
In a single expansion, the database jumped from ~350,000 to over 200 million predicted structures, spanning more than a million species — almost the entire catalogued protein universe, all freely downloadable.
October 2024 · Stockholm
The Nobel Prize in Chemistry
Demis Hassabis and John Jumper of DeepMind were awarded the 2024 Nobel Prize in Chemistry for AlphaFold, sharing it with David Baker for computational protein design.
05 — The hockey stick

Fifty years of climbing, then straight up.

Here is the whole story in one line. Every protein structure humanity had ever determined by experiment, accumulated over five decades, versus what AlphaFold added in two years. On a normal (linear) axis, half a century of Nobel-grade laboratory work barely lifts off the floor. Switch to a logarithmic axis to see that history — and the near-vertical jump that dwarfs it.

Experiments (Protein Data Bank) AlphaFold (predicted)
On a linear axis, fifty years of experiments — about 200,000 structures — is a line you can barely see above zero. AlphaFold added 200 million in a single step. That's the hockey stick: not a faster climb, but a different curve entirely.
06 — A billion years, free

The part that still sounds made up.

Doing all of that experimentally would have taken the field an almost unimaginable amount of time. DeepMind's own framing, and its founder's, put a number on it — and then they gave the entire result away.

"It's kind of like a billion years of PhD time done in one year."

Sir Demis Hassabis — DeepMind CEO and 2024 Nobel laureate, on AlphaFold folding every known protein
200M+
Protein structures predicted and released — nearly every one known to science
Free
Open access to the entire database, via a partnership with EMBL's EBI
3M+
Researchers in 190+ countries have used it — over 1M in low- and middle-income nations
40,000+
Scientific papers cite AlphaFold — over 30% of them tied to disease research

DeepMind estimates the database has saved researchers hundreds of millions of years of collective effort. But the number isn't the point — the point is what all those saved years unlock. Here is what people are already using it for.

Drug discovery

Knowing a target protein's shape is the starting gun for designing a drug that fits it. Researchers are compressing early-stage discovery from years toward months.

Fighting drug resistance

Structures of proteins in malaria, tuberculosis, and antibiotic-resistant bacteria are helping scientists target diseases that kill millions each year.

Enzymes that eat plastic

Teams have used AlphaFold structures to engineer enzymes that break down single-use plastic, a route toward faster recycling.

Neglected diseases

Because it's free, labs in lower-income countries — often working on diseases the market ignores — get the same starting point as the best-funded institutions.

07 — Why it matters

Some "impossible" problems are just waiting for the right tool.

AlphaFold is the clearest case yet of a pattern worth understanding: a problem declared intractable for half a century, then reframed and dissolved by machine learning — and the value multiplied by giving the result away rather than hoarding it.

"Intractable" has an expiry date

The protein folding problem was a fifty-year grand challenge until, suddenly, it wasn't. It's worth asking which of your own "unsolvable" bottlenecks are next.

The moat can be the giveaway

DeepMind's influence grew by releasing AlphaFold free. Open assets can compound into more value — and goodwill — than locking them away.

Data plus judgment beats brute force

AlphaFold didn't simulate physics faster; it learned patterns from 170,000 hard-won examples. The lesson generalises to most real AI wins.

The hard part isn't believing AI can do remarkable things — AlphaFold settles that. It's identifying which problem in your organization is a protein-folding problem: enormous, expensive, and quietly ready to fall. That's the conversation we have with leadership teams before they commit budget.

Talk to us about your AI roadmap

Sources

Facts and quotes drawn from: Google DeepMind — AlphaFold (usage, database size, citations, impact examples), Jumper et al., Nature (2021), 200-million-structure release (2022), AlphaFold DB in 2024 (214M+ sequences), Levinthal's paradox (Wikipedia), Protein Data Bank growth (Wikipedia), Hassabis "billion years of PhD time" (Univ. of Cambridge), and the 2024 Nobel Prize in Chemistry. Age of the universe (~13.8 billion years) is a standard cosmological estimate.