AI for Science and Engineering

Protein structure prediction

Protein structure prediction guesses the 3D shape a chain of amino acids folds into, a problem AlphaFold2 transformed but did not fully close.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where this shows up
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Protein structure prediction guesses the 3D shape a long chain of amino acids folds itself into.

Think about dropping a long necklace into a box and shaking it. It does not land in a random tangle. Beads that attract each other pull together, and beads that repel push apart. It settles into a fairly consistent folded shape, every time you drop that same necklace.

A protein is a long chain of building blocks called amino acids that folds the same way. That shape decides almost everything about what the protein can do in a living cell.

Why it exists

A protein's job — carrying oxygen in blood, fighting an infection, speeding up a chemical reaction — depends on its 3D shape. It is not only about its chain of amino acids, written out as a sequence of letters. Two proteins with almost the same sequence can fold into very different shapes and do very different jobs.

For decades, the only reliable way to find a protein's real shape was to measure it directly in a lab. That meant slow, expensive techniques like X-ray crystallography. Meanwhile, sequencing a protein's chain of amino acids became fast and cheap. A huge gap opened up: millions of known sequences, and shapes known for only a small fraction of them.

Predicting the shape directly from the sequence, without the lab step, had been a famous unsolved problem in biology for fifty years. In 2020, DeepMind's AlphaFold2 system reached accuracy close to lab measurement on a wide range of proteins, a result the field had not expected so soon. It remains one of the most significant applications of deep learning to a hard scientific problem to date.

How it works

A protein's sequence: a chain of amino acid letters
   M  K  T  A  Y  I  A  K  Q  R  Q  I  S  F  V  K  ...
        |
        v
Search for similar sequences already known across many species
        |
        v
A large network reasons about which amino acids
end up physically close to each other in the folded shape
        |
        v
Predicted 3D structure, plus a confidence score per position

Real systems like AlphaFold2 are enormous — trained on the whole of known protein science, running on serious computing hardware. Nothing at that scale fits in a lesson meant to run in seconds on a laptop, and this lesson does not pretend otherwise. The developer section below builds a small, honest toy version of a much simpler, related task instead. It predicts a short local pattern from a sequence — the same basic idea, at a scale that actually runs.

Where this shows up

Structural biologists use predicted structures to understand how a mutation might disrupt a protein's function, guiding which experiments to run next. Drug discovery teams use predicted protein shapes to understand where a candidate drug molecule might physically bind. This feeds into the same design loop as molecular property prediction.

An honest warning

A predicted structure comes with a confidence score for a reason. It is not always right. It is least reliable exactly where a protein is most flexible, or where it interacts with other molecules — often the most biologically interesting parts. A predicted structure is a strong hypothesis for where to look next in the lab, not a replacement for the lab confirming it.

Remember this

  • A protein folds into a 3D shape because parts of its amino acid chain attract or repel each other. That shape decides what the protein does.
  • AlphaFold2 closed most of the gap between the huge number of known sequences and the much smaller number of experimentally measured shapes, a genuine scientific breakthrough.
  • A predicted structure is a hypothesis with a confidence score attached, not a guaranteed answer — real biology still confirms it in the lab.

What to learn next

  • Molecular property prediction — predicting a molecule's traits, the sibling problem to predicting a protein's shape.
  • Attention — the mechanism AlphaFold2's Evoformer component builds on, deciding which amino acids should influence each other.
  • Transformers — the family of architecture AlphaFold2's core network belongs to.

Developer — Code and libraries.

Real structure prediction is not reproducible on a laptop in seconds. What is reproducible, and genuinely instructive, is the smaller idea underneath it: predicting a per-position label from a sequence and its neighbours. This is a real, simplified miniature of per-residue prediction tasks used throughout structural biology — invented data, not real biology.

Setup

bash
pip install scikit-learn numpy

Minimal runnable code

Three-letter toy alphabet standing in for amino acid types, and three invented local patterns standing in for the real secondary-structure shapes proteins fold into — helices, sheets, and looser "coil" regions.

protein.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)

# A miniature, invented stand-in for a real protein task -- NOT real biology.
# Alphabet: H = hydrophobic residue, P = polar residue, C = charged residue.
# Toy rule: a repeating H-P-P pattern tends to fold into a HELIX ('h'),
# alternating H-C tends to form a SHEET ('s'), everything else is COIL ('c').
# Real secondary-structure prediction is this same idea -- label each amino
# acid using its neighbours -- run on real sequences with far richer signals.

AMINO = ["H", "P", "C"]

def make_sequence(length):
    pattern = rng.choice(["helix", "sheet", "coil"])
    seq, labels = [], []
    for i in range(length):
        if pattern == "helix":
            seq.append(["H", "P", "P"][i % 3]); labels.append("h")
        elif pattern == "sheet":
            seq.append(["H", "C"][i % 2]); labels.append("s")
        else:
            seq.append(rng.choice(AMINO)); labels.append("c")
    # sprinkle a little noise so the pattern is not perfectly clean
    seq = [rng.choice(AMINO) if rng.random() < 0.08 else r for r in seq]
    return seq, labels

def window_features(seq, i, window=2):
    # describe position i using its neighbours, padding at the sequence ends
    feats = []
    for offset in range(-window, window + 1):
        j = i + offset
        residue = seq[j] if 0 <= j < len(seq) else "-"
        feats.extend([1 if residue == a else 0 for a in AMINO])
    return feats

X, y = [], []
for _ in range(120):
    seq, labels = make_sequence(length=15)
    for i in range(len(seq)):
        X.append(window_features(seq, i))
        y.append(labels[i])

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0)
model = RandomForestClassifier(n_estimators=200, random_state=0)
model.fit(X_train, y_train)
accuracy = model.score(X_test, y_test)
print(f"per-residue accuracy on held-out positions: {accuracy:.3f}")

# predict one full new sequence, position by position
test_seq, true_labels = make_sequence(length=15)
predicted = [model.predict([window_features(test_seq, i)])[0] for i in range(len(test_seq))]
print("sequence: ", " ".join(test_seq))
print("true:     ", " ".join(true_labels))
print("predicted:", " ".join(predicted))
Output
per-residue accuracy on held-out positions: 0.867
sequence:  H C P H P P H P P H P P H P P
true:      h h h h h h h h h h h h h h h
predicted: c c h h h h h h h h h h h h h

Walkthrough

window_features describes each position using a small window of neighbours on either side — the toy version of the real idea that a residue's local chemical neighbourhood strongly influences what shape it folds into.

The printed example shows something genuinely instructive: the first two positions are misclassified as coil, even though the sequence turns out to be a clean helix. The window at those positions did not yet contain enough of the repeating H-P-P pattern to recognise it confidently — a small, honest reminder that position-by-position prediction depends on how much context is available at each point, exactly as it does for real sequence labelling problems.

random_state=0 makes both the data generation and the model reproducible, so this output is not a one-off lucky run.

Common mistakes

Mistaking this toy for real protein science. It is a deliberately invented alphabet and deliberately invented rule, chosen because it is easy to verify. Real amino acids, real secondary structure, and real folding physics are far more complex, and this example makes no claim otherwise.

Using too small a context window. window=2 looks two positions in each direction. A pattern with a longer period, like a real alpha helix's roughly 3.6-residue turn, needs a window wide enough to actually contain it.

Evaluating per-position accuracy without checking whole-sequence sensibility. A model can be individually correct at most positions while still producing a shape that makes no physical sense stitched together — real structure prediction checks both.

Assuming a small, symbolic toy scales to real biology by adding more data alone. Real structure prediction needed a fundamentally different approach — reasoning over related sequences across many species (multiple sequence alignments) and geometric reasoning over 3D coordinates — not more of this kind of local window classifier.

Try it yourself

Change window=2 to window=4 and rerun. Compare whether the first few positions of a fresh helix sequence get classified correctly now that the window can see further into the repeating pattern.

What to learn next

  • Attention — how real structure-prediction systems decide which distant residues in a sequence should influence each other, beyond a fixed local window.
  • Named entity recognition — another per-position labelling task, in text instead of amino acid sequences.
  • Molecular property prediction — predicting a molecule's traits once its structure, real or predicted, is known.

Researcher — Mathematics and papers.

Why the problem is hard

The number of possible 3D foldings of even a short amino acid chain is astronomically large — a fact known as Levinthal's paradox (Levinthal, 1969): if a protein searched every possible conformation randomly, folding would take far longer than the milliseconds to seconds it actually takes in a cell. Real proteins fold along efficient energy-guided pathways, not exhaustive search, and reproducing that efficiency computationally was the core difficulty for decades.

The AlphaFold2 approach, at a high level

AlphaFold2 (Jumper et al., 2021) does not predict structure from a single sequence in isolation. Given a query sequence, it first searches large sequence databases for evolutionarily related sequences across many species, building a multiple sequence alignment (MSA). Positions that co-vary together across related species — mutating together across evolution — are strong evidence that those positions are physically close in the folded structure, since a mutation at one is often only tolerated if compensated by a mutation at the other.

This MSA, together with structural templates when available, is processed by the Evoformer, a stack of attention-based blocks (see attention and transformers) that iteratively refines two representations: one over pairs of residues, one over the MSA itself, letting information flow between "which residues co-evolve" and "which residues are predicted to be spatially close." A structure module then converts this refined representation directly into 3D atomic coordinates.

Each predicted residue carries a pLDDT confidence score (predicted Local Distance Difference Test), trained to estimate how accurate that specific part of the structure prediction is likely to be — the mechanism behind the "confidence score per position" mentioned in the beginner section, and the reason flexible or poorly-conserved regions reliably show lower scores.

Complexity and cost

The developer example's window classifier is O(n · w) for a sequence of length n and window size w — trivial. Real structure prediction is dominated by the attention operations inside the Evoformer, which scale with the square of the number of residues, and additionally with the depth of the MSA (how many related sequences were found):

ComponentApproximate scaling
Window classifier (toy)O(n · w)
Evoformer pairwise attentionO(n^2) per layer, over n residues
MSA processingscales with MSA depth × n

This is why AlphaFold2 inference, while far faster than a lab structure determination, is not a laptop-seconds operation — a single prediction for a moderately sized protein can take minutes on serious accelerator hardware, and considerably longer for very large proteins or complexes.

Papers

  • Levinthal, C. (1969). How to Fold Graciously. Mössbauer Spectroscopy in Biological Systems, Proceedings. Origin of Levinthal's paradox.
  • Jumper, J. et al. (2021). Highly Accurate Protein Structure Prediction with AlphaFold. Nature 596. The AlphaFold2 paper.
  • Baek, M. et al. (2021). Accurate Prediction of Protein Structures and Interactions Using a Three-Track Neural Network. Science 373. RoseTTAFold, an independent, related approach published the same year.
  • Mirdita, M. et al. (2022). ColabFold: Making Protein Folding Accessible to All. Nature Methods 19. A widely used, more accessible implementation pipeline built on AlphaFold2's ideas.

Current state

AlphaFold2 and its successors have made structure prediction for a single protein chain a largely solved problem for many cases, verified against large benchmark sets of experimentally known structures. Open problems that remain active research include predicting how multiple proteins bind together as complexes, predicting the effect of a mutation on structure and function jointly, and predicting the many different shapes a flexible protein can take rather than one fixed structure. A predicted structure, however accurate on average, still requires experimental or functional validation before it informs a real drug-design or clinical decision — the pLDDT confidence score is a statistical estimate, not a guarantee for any one specific prediction.

What to learn next