AI for Science and Engineering
Predicting molecular properties
Molecular property prediction estimates traits like solubility or toxicity from a molecule's structure, and gets its evaluation wrong if the train-test split ignores which molecules share a chemical core.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Molecular property prediction guesses a trait of a molecule — like whether it dissolves in water — from its structure alone.
Think about guessing whether a fruit will taste sweet by looking at its shape and skin alone, without tasting it. An experienced fruit seller can often guess correctly from experience with thousands of fruits before. Molecular property prediction is the same guessing game, played on molecules. Given the structure — which atoms, connected how — it predicts a property without running the real chemistry experiment first.
Why it exists
Testing a real molecule's properties in a lab takes real time and real money, sometimes weeks per molecule. That includes how well it dissolves, whether it is toxic, and how strongly it binds to a target. Chemists exploring new drugs or materials might want to consider millions of candidate molecules. Testing all of them in a lab is not possible.
A model trained on molecules whose properties are already known can screen millions of untested candidates in minutes, ranking them by predicted promise. The lab then only needs to test the small number of candidates the model ranks highest. This does not remove the lab work — it decides which lab work is worth doing first.
How it works
Known molecules, each with a measured property
|
v train a model
Trained model: structure -> predicted property
|
v
Millions of untested candidate molecules
|
v screen them all in minutes
A short ranked list -> sent to the real lab for testingWhere this shows up
Drug discovery teams use property prediction to narrow down millions of candidate molecules. It picks out which ones are worth the weeks of lab time needed to actually synthesise and test them. Materials scientists use the same idea to screen candidate battery or solar-panel materials before building a single physical sample.
An honest warning
The way training data gets split into "practice" and "test" molecules can quietly make a model look far better than it is. Many molecules in a chemical library are small variations on the same core structure. If variations of the same core end up on both sides of the split, the model can partly recognise the core it already saw. The test score then becomes an overly optimistic measure of how well it handles a genuinely new core. This is one of the most common — and most easily hidden — mistakes in this field, and it is demonstrated directly below.
None of this replaces a real lab. A wrong prediction about toxicity is a real safety risk, not an inconvenience. Every promising candidate still needs real chemical synthesis and real laboratory testing. Anything intended for human use also needs the full regulatory approval and safety trial process, before it can be trusted with anyone's health.
Remember this
- Property prediction screens huge numbers of candidate molecules cheaply, so the real lab only tests the most promising few.
- A random split of training and test molecules can leak information between near-duplicate molecules, making accuracy look better than it really is.
- A model's prediction is a lead worth investigating, never a substitute for real laboratory testing and, for anything medical, real regulatory approval.
What to learn next
- Machine-learned interatomic potentials — predicting a different kind of molecular quantity, atomic energy, with its own accuracy demands.
- Generative design of molecules — proposing brand-new candidates instead of only ranking existing ones.
- Overfitting and underfitting — the general version of the leaking-split problem shown here.
Developer — Code and libraries.
The single most important engineering decision in this field is how you split your data. This example makes the consequence of getting it wrong visible in one run.
Setup
pip install scikit-learn numpy pandasMinimal runnable code
Real cheminformatics tools compute rich structural descriptors from a molecule's actual graph of atoms and bonds. This example uses a small stand-in — three hand-built numeric descriptors per molecule — so the code needs no extra chemistry library, while the core lesson about splitting stays exactly the same.
Fourteen scaffolds (core structures) each produce several close variants — this mirrors how real chemical libraries are built around a shared core.
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error
rng = np.random.default_rng(3)
n_scaffolds = 14
rows = []
for scaffold_id in range(n_scaffolds):
base_weight = rng.uniform(150, 400)
base_rings = rng.integers(1, 4)
base_polar = rng.integers(0, 6)
base_solubility = rng.uniform(-4, 0) # near-unrelated to weight/rings/polar on purpose
n_variants = rng.integers(7, 12)
for _ in range(n_variants):
weight = base_weight + rng.normal(0, 1.5) # tight noise: siblings look almost identical
rings = base_rings
polar_atoms = base_polar
solubility = base_solubility + rng.normal(0, 0.05)
rows.append([scaffold_id, weight, rings, polar_atoms, solubility])
df = pd.DataFrame(rows, columns=["scaffold", "weight", "rings", "polar_atoms", "solubility"])
X = df[["weight", "rings", "polar_atoms"]].values
y = df["solubility"].values
def evaluate(train_idx, test_idx, label):
model = RandomForestRegressor(n_estimators=200, random_state=0)
model.fit(X[train_idx], y[train_idx])
pred = model.predict(X[test_idx])
mae = mean_absolute_error(y[test_idx], pred)
print(f"{label:22s} test MAE = {mae:.3f} (n_test={len(test_idx)})")
rng2 = np.random.default_rng(1)
shuffled = rng2.permutation(len(df))
cut = int(0.8 * len(df))
evaluate(shuffled[:cut], shuffled[cut:], "random row split")
scaffolds = df["scaffold"].unique().copy()
rng2.shuffle(scaffolds)
test_scaffolds = set(scaffolds[-4:])
test_idx = df.index[df["scaffold"].isin(test_scaffolds)].to_numpy()
train_idx = df.index[~df["scaffold"].isin(test_scaffolds)].to_numpy()
evaluate(train_idx, test_idx, "scaffold split")
print()
print(f"molecules: {len(df)}, scaffolds: {n_scaffolds}")random row split test MAE = 0.269 (n_test=26) scaffold split test MAE = 1.022 (n_test=37) molecules: 128, scaffolds: 14
Walkthrough
Both evaluations use the exact same model, on the exact same underlying data — only the split differs, and the reported error nearly quadruples.
The random row split shuffles all 128 molecules together before cutting off the last 20% as a test set. Because siblings from the same scaffold have nearly identical descriptors and nearly identical solubility, some test molecules almost certainly have a near-duplicate twin sitting in the training set. The model does not need to generalise — it can recognise a familiar pattern.
The scaffold split holds out four entire scaffolds — every variant of those cores goes to the test set, none to training. Now the model faces descriptor patterns it has genuinely never encountered. The error jumps because this is a fair test of what the model can actually do with a brand-new chemical core, which is the real question anyone screening new candidate molecules cares about.
base_solubility was deliberately generated with almost no real relationship to the descriptors, and within-scaffold noise was kept small — this exaggerates the effect to make it visible in a small toy example. Real chemical data has genuine structure-property relationships mixed in with this same scaffold-leakage risk.
Common mistakes
Reporting only a random-split score. It is the single most common way an honest-looking result turns out to be optimistic once a model reaches new chemistry.
Assuming scaffold splitting is only relevant for graph neural networks. The leak has nothing to do with model architecture. It comes from how similar the training and test molecules are, regardless of what kind of model reads them.
Believing scaffold-split accuracy is the true ceiling. It is a more honest lower bound for genuinely novel chemistry, not a guarantee — a candidate's real-world behaviour can still surprise a model trained on a limited scaffold diversity.
Skipping the split question because the dataset "looks large." 128 molecules across 14 scaffolds averages under 10 molecules per scaffold — small enough that a handful of scaffolds ending up in the test set changes the result substantially, as this run shows.
Try it yourself
Change test_scaffolds = set(scaffolds[-4:]) to hold out only one scaffold instead of four, rerun, and compare the MAE. With less scaffold diversity held out, watch whether the gap between the two splits grows or shrinks, and think about why.
What to learn next
- Feature engineering — the general practice of building descriptors like the ones used here, done properly with real chemistry.
- Random forest — the model used for both splits in this example.
- Overfitting and underfitting — the general principle behind why a leaking split inflates apparent performance.
Researcher — Mathematics and papers.
The formal setting
Let a molecule be represented as a graph G = (V, E), where nodes V are atoms and edges E are bonds. A property predictor learns f: G -> y, where y is a scalar or vector property (solubility, binding affinity, toxicity class). Two representational choices dominate practice:
- Fixed fingerprints — a molecule is hashed into a fixed-length binary or count vector describing which substructures it contains (e.g. Extended-Connectivity Fingerprints, Rogers and Hahn, 2010), then fed to a conventional regressor such as the random forest used in the developer example.
- Learned graph representations — a graph neural network passes messages between bonded atoms over several layers, learning a representation end-to-end rather than relying on a hand-designed hashing scheme (Gilmer et al., 2017, the "Message Passing Neural Network" framework most modern architectures descend from).
Scaffold splitting, formally
Let each molecule i be assigned a Bemis-Murcko scaffold s(i) — the ring systems and connecting linkers of the molecule with side chains removed (Bemis and Murcko, 1996). A scaffold split partitions molecules by s(i) rather than by i directly:
Train = { i : s(i) ∈ S_train }, Test = { i : s(i) ∈ S_test }, S_train ∩ S_test = ∅This guarantees no scaffold appears on both sides, which is the property the developer example's "scaffold split" enforces directly on the toy scaffold column. Yang et al. (2019) and Wu et al. (2018, the MoleculeNet benchmark paper) both report substantially lower apparent performance under scaffold splitting than random splitting across a range of standard datasets and model families, confirming this is a property of the evaluation protocol, not of any one model architecture.
Complexity and cost
For n molecules with average k atoms:
| Representation | Featurisation cost | Model training cost (typical) |
|---|---|---|
| Fixed fingerprint + random forest | O(n · k) | O(n log n) |
| Message-passing GNN | O(n · k) per epoch | scales with epochs × k per molecule |
Graph neural networks typically outperform fixed fingerprints on large datasets with enough examples per scaffold family to learn transferable substructure representations, and offer little advantage on small datasets, where fixed fingerprints combined with strong regularisation are a competitive, cheaper baseline (Yang et al., 2019).
Papers
- Bemis, G. and Murcko, M. (1996). The Properties of Known Drugs. 1. Molecular Frameworks. Journal of Medicinal Chemistry 39. Defines the scaffold concept used throughout the field.
- Rogers, D. and Hahn, M. (2010). Extended-Connectivity Fingerprints. Journal of Chemical Information and Modeling 50. The standard fixed-fingerprint representation (ECFP).
- Gilmer, J. et al. (2017). Neural Message Passing for Quantum Chemistry. ICML. The MPNN framework underlying most graph neural network property predictors.
- Wu, Z. et al. (2018). MoleculeNet: A Benchmark for Molecular Machine Learning. Chemical Science 9. Establishes scaffold splitting as the field's standard, more honest evaluation protocol.
- Yang, K. et al. (2019). Analyzing Learned Molecular Representations for Property Prediction. Journal of Chemical Information and Modeling 59.
Current state
Scaffold-split, or the closely related time-split (train on molecules synthesised before a cutoff date, test on those synthesised after), is now the expected standard for reporting property-prediction results credibly; a random-split-only result is treated with suspicion by reviewers in the field. Foundation-model-style molecular representations pretrained on large unlabelled molecule collections, then fine-tuned on small labelled property datasets, are an active direction for improving low-data performance. None of this changes the field's central limitation: computational screening prioritises candidates for testing, and a computed property is not a substitute for wet-lab measurement, in vivo toxicology, or the regulatory approval process required before any candidate reaches real-world use.
What to learn next
- Generative design of molecules — using a property predictor as the scoring function inside a design loop.
- Group leakage — the general version of the scaffold-split problem: the same underlying group appearing in both train and test.
- Protein structure prediction — a related structural biology problem with its own, much larger-scale evaluation challenges.