AI for Science and Engineering

Generative design of molecules

Generative molecule design proposes brand-new candidate molecules likely to have a wanted property, rather than only ranking molecules that already exist.

On this page 6
  1. Why it exists
  2. How it works
  3. Where this shows up
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Generative molecule design proposes brand-new candidate molecules, rather than only ranking molecules someone already made.

Think about a cook who has tasted thousands of dishes. Instead of only judging recipes handed to them, they start proposing entirely new recipes they believe will taste good. Molecular property prediction is the judge — it scores a molecule that already exists. Generative molecule design is the cook proposing something new — every proposal still needs someone to actually cook and taste it before anyone can trust it.

Why it exists

The universe of chemically possible small molecules is estimated to be far larger than the number of atoms in the observable universe. Chemists cannot search that space by hand. They cannot even search it by testing every option. There is not enough time or money in the world to synthesise and test more than a tiny sliver of it.

Property prediction alone only ranks molecules that someone already thought to propose. Generative design flips the direction. It starts from what "good" looks like — a property score from a model like the ones in the previous lesson. Then it searches for new molecules likely to score well, before any of them exist in a lab.

How it works

A scoring function: molecule -> predicted score
        |
        v
Propose a batch of candidate molecules
        |
        v
Score every candidate
        |
        v
Keep the best few, make small changes to them, propose again
        |
        v  repeat many rounds
A short list of promising, never-before-made molecules
        |
        v
Sent to a real chemist to actually synthesise and test

This loop — generate, score, keep the best, repeat — is the same basic idea either way. The "generator" might be a genetic algorithm making small random edits, or a modern generative model like a diffusion model trained on huge molecule collections.

Where this shows up

Drug discovery teams use generative design to propose new candidate molecules aimed at binding a specific disease-related protein. This narrows an impossibly large search down to a shortlist worth attempting to make. Materials science teams use the same idea to propose new candidate compounds for batteries or solar cells.

An honest warning

A generated molecule is a suggestion, not a discovery. Many proposals are not even chemically valid — atoms with the wrong number of bonds, unstable arrangements that could never exist. Many valid ones would be extremely difficult or impossible to actually synthesise in a lab with known chemistry. That turns out to be as big a bottleneck as coming up with the idea in the first place.

Every single generated molecule that looks promising still needs real chemical synthesis and real laboratory testing. For anything intended for human use, it also needs full toxicology studies and regulatory approval, before it can be trusted anywhere near a patient.

Remember this

  • Generative molecule design proposes new candidates aimed at a wanted property, instead of only ranking molecules that already exist.
  • The core loop is generate, score, keep the best, repeat — whether the generator is simple or a large modern model.
  • A generated molecule is an untested idea. Validity, synthesisability, real lab testing, and — for medicine — regulatory approval all still have to happen after.

What to learn next

Developer — Code and libraries.

This example is not real chemistry — there is no chemistry library involved. It is the generate-score-select loop stripped down to its purest form, so the loop itself is what stays visible.

Setup

bash
pip install numpy

Minimal runnable code

A "molecule" here is a short string of fragment letters. The scoring function is invented, standing in for a real trained property predictor. The valence rule — no fragment repeating three times in a row — stands in for real chemical validity constraints, like an atom only being able to form so many bonds.

genmol.py
import numpy as np

rng = np.random.default_rng(0)

# A toy stand-in for real molecule generation. A "molecule" here is a
# short string of fragment letters -- not real chemistry -- but the loop
# (generate many candidates, score them, keep the best, mutate, repeat) is
# the same loop real generative chemistry tools run, using a real
# property predictor and a real chemical validity checker in place of ours.

FRAGMENTS = list("ABCDE")
LENGTH = 8

def is_valid(molecule):
    # toy "valence rule": the same fragment cannot repeat 3 times in a row
    return not any(molecule[i] == molecule[i+1] == molecule[i+2] for i in range(len(molecule) - 2))

def score(molecule):
    # a synthetic stand-in for a trained property predictor (e.g. predicted binding affinity)
    counts = {f: molecule.count(f) for f in FRAGMENTS}
    return counts["A"] * 1.5 + counts["C"] * 1.0 - counts["E"] * 2.0 - abs(counts["B"] - counts["D"])

def random_molecule():
    while True:
        m = "".join(rng.choice(FRAGMENTS, LENGTH))
        if is_valid(m):
            return m

def mutate(molecule):
    for _ in range(20):
        pos = rng.integers(0, LENGTH)
        candidate = molecule[:pos] + rng.choice(FRAGMENTS) + molecule[pos+1:]
        if is_valid(candidate):
            return candidate
    return molecule

population = [random_molecule() for _ in range(30)]

for generation in range(25):
    scored = sorted(population, key=score, reverse=True)
    survivors = scored[:10]
    children = []
    for parent in survivors:
        for _ in range(2):
            child = mutate(parent)
            children.append(child)
    population = survivors + children

best = max(population, key=score)
print(f"best molecule after 25 generations: {best}   score={score(best):.2f}")
print(f"best of one fresh random batch was:  {max([random_molecule() for _ in range(30)], key=score)}")

# how often does blind mutation propose something invalid, before we filter it?
attempts = 2000
invalid = sum(1 for _ in range(attempts)
              if not is_valid("".join(rng.choice(FRAGMENTS, LENGTH))))
print(f"\nof {attempts} randomly assembled candidates, {invalid} ({invalid/attempts:.1%}) broke the toy valence rule")
Output
best molecule after 25 generations: ACACAACA   score=10.50
best of one fresh random batch was:  DADBCCAB

of 2000 randomly assembled candidates, 369 (18.4%) broke the toy valence rule

Walkthrough

The genetic loop starts from 30 random valid candidates and, over 25 rounds, repeatedly keeps the top 10 by score, mutates each survivor twice, and moves on. Compare the winning score after optimisation, 10.50, to whatever a single fresh random batch of 30 candidates achieves — the optimisation loop consistently finds a better answer than random search of the same size, because it concentrates its search near what has already scored well.

is_valid is checked at every single mutation attempt, not only at the end — this matters, because roughly a fifth of purely random candidates break even this simple toy rule. A real system that ignored validity until the very end would waste most of its search budget on molecules that were never real options.

mutate tries up to 20 random edits before giving up and returning the parent unchanged, since not every position edit keeps a candidate valid — real molecular generators face a much harder version of the same problem, where most possible edits break chemical validity.

Common mistakes

Scoring invalid candidates instead of filtering them first. A model happily assigns a number to a nonsense candidate. Constraining generation to valid candidates, as this example does inside both random_molecule and mutate, is not optional polish — it is the difference between a real search and a search over mostly meaningless results.

Judging the loop by its best score alone. A generative search can find one excellent outlier while producing mostly poor candidates around it. Real systems typically report a whole shortlist, not only the single top result.

Believing a high predicted score is the finish line. The score here comes from an invented formula. In a real system it comes from a trained property predictor, which is itself an approximation — see the honest warning in molecular property prediction about scaffold generalisation.

Ignoring synthesisability. This toy has no notion of whether a "molecule" could actually be made by a chemist. Real generative-design systems increasingly score candidates on estimated ease of synthesis alongside predicted properties, because a molecule that scores perfectly but cannot be made is not useful.

Try it yourself

Add a fourth term to score that penalises candidates far from a target length ratio of fragment A, for example - abs(counts["A"] - 3) * 0.5. Rerun and see how the winning molecule's composition shifts to satisfy the new constraint alongside the old ones — a small illustration of how real design problems balance several wanted properties at once.

What to learn next

  • Molecular property prediction — replacing the toy scoring function with something trained on real data.
  • Diffusion models — the modern generative approach many current molecule-design systems use instead of a genetic algorithm.
  • AI agents — orchestrating a generate-score-select loop like this one as part of a larger automated discovery pipeline.

Researcher — Mathematics and papers.

The formal setting

Generative molecule design is commonly framed as a constrained optimisation problem over a molecular representation m:

m* = argmax_{m ∈ M_valid}  score(m)
  • M_valid — the set of chemically valid molecules under the chosen representation (valid valences, sensible ring structures)
  • score(m) — typically a learned property predictor, or a weighted combination of several (potency, solubility, synthesisability, novelty)

The developer example's is_valid and score are simplified, invented stand-ins for M_valid and score(m) respectively, over a toy string representation instead of a real molecular graph or SMILES string.

Representations and generators

Three representational choices dominate the field, each with different validity guarantees:

  • String-based (SMILES) — molecules encoded as a linear text string. Character-level generative models (RNNs, transformers) can propose syntactically invalid strings that do not correspond to any real molecule, requiring a validity filter after generation (Gómez-Bombarelli et al., 2018).
  • Graph-based — molecules generated directly as atom-and-bond graphs, allowing validity constraints (correct valence) to be enforced incrementally during generation rather than checked only afterward (Jin, Barzilay and Jaakkola, 2018, the Junction Tree VAE).
  • 3D diffusion-based — recent systems generate atoms directly in 3D space using denoising diffusion (see diffusion models), conditioned on a target binding pocket, producing molecules already shaped to fit a specific protein structure (Hoogeboom et al., 2022, EDM; Guan et al., 2023, on pocket-conditioned generation).

The synthesisability bottleneck

A molecule can be valid, score well on every predicted property, and still be practically impossible to make with known chemical reactions. Synthetic accessibility scores (Ertl and Schuffenhauer, 2009) estimate this from structural complexity, and are now commonly included as an additional term in the optimisation objective, or used as a post-hoc filter, precisely because early generative chemistry work found that unconstrained optimisation reliably produces molecules no chemist can actually synthesise (Gao and Coley, 2020).

Complexity and cost

For a population of size k evolved over g generations, each candidate of length n:

ComponentCost
Genetic algorithm (developer example)O(g · k · n) for mutation and scoring
Graph-based generative model, per candidateone forward pass through the generator, plus validity checking
3D diffusion generation, per candidateO(T) denoising steps, T typically in the tens to low hundreds
Property scoring, per candidateone forward pass of the property predictor, cheap relative to real lab synthesis

In every case, generation and scoring are computationally trivial compared to the real bottleneck: physically synthesising and testing even a small shortlist, which is measured in days to months, not milliseconds.

Papers

  • Gómez-Bombarelli, R. et al. (2018). Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules. ACS Central Science 4. An early, influential generative approach over SMILES strings.
  • Jin, W., Barzilay, R. and Jaakkola, T. (2018). Junction Tree Variational Autoencoder for Molecular Graph Generation. ICML.
  • Ertl, P. and Schuffenhauer, A. (2009). Estimation of Synthetic Accessibility Score of Drug-like Molecules. Journal of Cheminformatics 1.
  • Hoogeboom, E. et al. (2022). Equivariant Diffusion for Molecule Generation in 3D. ICML.
  • Gao, W. and Coley, C. (2020). The Synthesizability of Molecules Proposed by Generative Models. Journal of Chemical Information and Modeling 60. A widely cited critique showing many generative-model outputs are difficult or impossible to synthesise.

Current state

Pocket-conditioned 3D generation is a fast-moving current direction, aiming to generate molecules already shaped to bind a specific, known or predicted protein structure. A persistent, honest finding across the field is that optimisation pressure alone tends to produce molecules that exploit weaknesses in the scoring model rather than genuinely good candidates — a form of Goodhart's law specific to this domain — which is why experimental validation of any generative-design shortlist remains non-negotiable, and why no serious lab treats a generated molecule as more than a starting hypothesis.

What to learn next