AI for Science and Engineering
Machine-learned interatomic potentials
A machine-learned interatomic potential predicts the energy between atoms fast enough to simulate millions of tiny time steps, standing in for far slower quantum chemistry.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A machine-learned interatomic potential predicts the energy between atoms fast enough to simulate millions of tiny moments in a row.
Think about a chef who has cooked thousands of dishes. They can now judge whether a curry needs more salt by smell alone, without running a lab test every time. Working out the exact energy between atoms from the true laws of quantum mechanics is like that lab test — accurate, but slow. A machine-learned interatomic potential is the chef's trained instinct — fast, learned from many examples of the real calculation, close enough to use constantly.
Why it exists
Scientists simulate how atoms move over time — in a material, a drug molecule, a battery — using molecular dynamics. It steps forward in tiny fractions of a second, over and over, millions or billions of times, to see how the whole system evolves. At every single one of those tiny steps, the simulation needs to know the energy and forces acting on every atom.
The most accurate way to compute that energy is quantum chemistry, specifically a method called density functional theory. It is also extremely slow — far too slow to repeat billions of times. Before machine-learned potentials, scientists used simpler hand-written formulas as a fast approximation, accepting that these formulas were often crude and missed real chemistry.
A machine-learned interatomic potential is trained on a batch of expensive, accurate quantum-chemistry calculations, and learns to predict the same energies thousands of times faster. It brings quantum-level accuracy to simulations that need billions of fast steps. That works only as long as the model stays trustworthy at every one of those steps — which turns out to be the hard part.
How it works
Slow, accurate quantum chemistry calculation
|
v run it on a few thousand atomic arrangements
Training data: (arrangement of atoms, true energy) pairs
|
v fit a model
A trained interatomic potential (predicts in microseconds)
|
v
New atomic arrangement -> Potential -> predicted energy AND predicted forceNotice this needs both an energy and a force at every step — the force tells the simulation which direction to push each atom next. That extra requirement, not present in most machine learning problems, is what makes this domain different from an ordinary regression task.
Where this shows up
Materials scientists use these models to simulate how new battery materials behave, over millions of atomic movements. It costs a fraction of running full quantum chemistry the whole way.
Drug discovery teams use them to simulate how a candidate molecule wiggles and settles inside a protein pocket. That is where molecular property prediction and generative molecule design get some of their numbers from.
An honest warning
A model that predicts good energies is not automatically safe to use in a real simulation. If the predicted force is not smooth and consistent as atoms move slightly, the simulation can behave unphysically. Atoms can drift with no real force pushing them, or energy can appear from nowhere. Getting the energy roughly right is not the same as being safe to simulate with. Results from any such model should be checked against known experimental behaviour, before being trusted for a real materials or drug design decision.
Remember this
- A machine-learned interatomic potential replaces slow, accurate quantum chemistry with a fast, learned approximation, trained on a batch of the real calculations.
- Molecular dynamics needs both energy and force at every tiny time step, not only an energy guess.
- A model can look accurate on energy and still be unsafe to simulate with, if its predicted forces are not smooth and physically consistent.
What to learn next
- Surrogate models for simulation — the general idea this lesson specialises to atoms and energies.
- Molecular property prediction — using a trained model to predict a molecule's behaviour, not only one energy value.
- Derivatives and gradients — the mathematics behind why force is the derivative of energy.
Developer — Code and libraries.
The distinctive engineering problem here is not accuracy alone — it is whether a model's predicted force is well-behaved enough to simulate with. This example makes that failure visible.
Setup
pip install torch scikit-learn numpyMinimal runnable code
lennard_jones is real, simple physics: the energy between two atoms as a function of the distance between them. We train two very different models to predict this energy from distance, then compare the force each one implies — force being -dE/dr, the negative rate of change of energy with distance.
import numpy as np
import torch
import torch.nn as nn
from sklearn.ensemble import RandomForestRegressor
torch.manual_seed(0)
rng = np.random.default_rng(0)
def lennard_jones(r, epsilon=1.0, sigma=1.0):
# the real physics: energy between two atoms as a function of their distance r
return 4 * epsilon * ((sigma / r) ** 12 - (sigma / r) ** 6)
# training data: energies at a handful of sampled distances (this stands in
# for a set of expensive quantum-chemistry calculations, one per distance)
r_train = np.sort(rng.uniform(0.9, 2.5, 25))
e_train = lennard_jones(r_train)
# --- model 1: a random forest, a strong general-purpose regressor ---
forest = RandomForestRegressor(n_estimators=200, random_state=0)
forest.fit(r_train.reshape(-1, 1), e_train)
# --- model 2: a small smooth neural network ---
net = nn.Sequential(nn.Linear(1, 32), nn.Tanh(), nn.Linear(32, 32), nn.Tanh(), nn.Linear(32, 1))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
r_t = torch.tensor(r_train, dtype=torch.float32).reshape(-1, 1)
e_t = torch.tensor(e_train, dtype=torch.float32).reshape(-1, 1)
for _ in range(3000):
opt.zero_grad()
loss = ((net(r_t) - e_t) ** 2).mean()
loss.backward()
opt.step()
# force = -dE/dr. Compare a finite-difference force (works on any model,
# including the forest) against the neural network's own autograd force.
def finite_diff_force(predict_fn, r, h=1e-4):
return -(predict_fn(r + h) - predict_fn(r - h)) / (2 * h)
test_r = np.linspace(1.0, 2.2, 7)
print("r LJ energy forest E forest F(fd) net F(autograd)")
for r in test_r:
true_e = lennard_jones(r)
forest_e = forest.predict([[r]])[0]
forest_f = finite_diff_force(lambda x: forest.predict(np.array(x).reshape(-1, 1))[0], r)
r_g = torch.tensor([[r]], dtype=torch.float32, requires_grad=True)
net_e = net(r_g)
net_f = -torch.autograd.grad(net_e, r_g)[0].item()
print(f"{r:.2f} {true_e:8.4f} {forest_e:8.4f} {forest_f:10.4f} {net_f:10.4f}")r LJ energy forest E forest F(fd) net F(autograd) 1.00 0.0000 1.4840 -0.0000 23.0768 1.20 -0.8910 -0.8429 -0.0000 -2.1992 1.40 -0.4607 -0.4987 -0.0000 -1.7446 1.60 -0.2242 -0.2173 -0.0000 -0.7527 1.80 -0.1141 -0.1225 -0.0000 -0.3416 2.00 -0.0615 -0.0669 -0.0000 -0.1781 2.20 -0.0350 -0.0341 -0.0000 -0.1031
Walkthrough
Look at the forest's energy column first: it is reasonably close to the true Lennard-Jones energy at every distance — not perfect, but plausible.
Now look at the forest's force column: every single value is essentially zero. A random forest's predictions are piecewise constant — flat within each small region a tree splits on — so its local slope, and therefore its force, is zero almost everywhere. In a real molecular dynamics simulation, this model would tell every atom that no force is acting on it at all. The atoms would drift, unpushed, regardless of how close together they got. The energy predictions looked fine. The model was still unsafe to simulate with.
The neural network's force column, by contrast, is smooth and varies sensibly with distance, because net(r) is a smooth mathematical function and torch.autograd.grad computes its exact derivative — not an approximation, the real slope of the function the network represents.
Common mistakes
Judging a potential only by its energy accuracy. The forest's energy predictions were not bad. Its forces were unusable. Always check both.
Using a non-differentiable model family for a potential meant to drive dynamics. Random forests, k-nearest neighbours and other non-smooth models can be excellent regressors and terrible interatomic potentials, for exactly this reason.
Computing forces by finite differences on a model that was never meant to be differentiated. finite_diff_force here is deliberately model-agnostic to expose the forest's flatness. Production systems train models specifically so that the force can come from automatic differentiation of the energy, guaranteeing energy conservation is at least structurally respected.
Trusting a model far outside its training range. Both models here were trained on distances between 0.9 and 2.5. Neither has any idea what happens at 0.5, where real atomic repulsion becomes extremely steep.
Try it yourself
Add from sklearn.neural_network import MLPRegressor and swap in a scikit-learn multi-layer perceptron for the forest — a smooth model without needing PyTorch. Recompute finite_diff_force for it and compare to the random forest's flat zeros. The improvement in force smoothness should be visible immediately, even without exact autograd.
What to learn next
- Surrogate models for simulation — the general pattern this lesson specialises for atoms.
- Physics-informed neural networks — building a governing equation directly into training, the same instinct applied here to keep forces physical.
- Random forest — understanding why tree-based models are piecewise constant helps explain the flat-force failure above.
Researcher — Mathematics and papers.
The formal setting
Let a system have N atoms with positions {r_1, ..., r_N}. An interatomic potential is a function E({r_i}) giving the total potential energy of the configuration. The force on atom i is defined as the negative gradient of energy with respect to that atom's position:
F_i = - ∇_{r_i} E({r_1, ..., r_N})∇_{r_i}— the gradient operator with respect to atomi's three position coordinatesF_i— a 3-dimensional force vector on atomi
This is not an optional add-on. It is the physical definition of force as conservative, and it is what guarantees energy conservation in a simulation: if F_i is exactly the gradient of some smooth E, total energy is provably conserved along the simulated trajectory (up to numerical integration error). If F_i is fit independently of E — as the finite-difference force was in the developer example, decoupled from what the model actually represents — that guarantee is lost, and simulated energy can drift arbitrarily over long trajectories.
Modern architectures
Early machine-learned potentials (Behler and Parrinello, 2007) used hand-crafted, physically motivated descriptors of each atom's local environment — "symmetry functions" — fed into a small neural network per atom, with atomic contributions summed to a total energy, and forces obtained by automatic differentiation of that sum. This structurally enforces the energy-force relationship from the start, unlike the developer example's finite-difference force.
Modern architectures largely use graph neural networks, treating atoms as graph nodes and bonds as edges, with message-passing layers that respect physical symmetries. Equivariant architectures such as NequIP (Batzner et al., 2022) go further, guaranteeing that predicted forces rotate correctly when the input molecule is rotated — a symmetry real physics has and an ordinary neural network does not automatically respect.
Complexity and cost
For N atoms and a local cutoff radius capturing k neighbours per atom on average:
| Component | Typical cost |
|---|---|
| Descriptor / message-passing per atom | O(k) per atom, O(N·k) total |
| Energy prediction | O(N) given descriptors |
| Force via automatic differentiation | roughly 2-3x the cost of the energy forward pass |
| Quantum chemistry (DFT) comparison | O(N^3) or worse per configuration |
The near-linear scaling of a machine-learned potential in N, against DFT's cubic-or-worse scaling, is the entire practical case for using one — it is what makes simulating tens of thousands of atoms for nanoseconds of simulated time feasible at all.
Papers
- Behler, J. and Parrinello, M. (2007). Generalized Neural-Network Representation of High-Dimensional Potential-Energy Surfaces. Physical Review Letters 98. The foundational paper establishing per-atom neural network potentials.
- Batzner, S. et al. (2022). E(3)-Equivariant Graph Neural Networks for Data-Efficient and Accurate Interatomic Potentials. Nature Communications 13. NequIP, a widely used equivariant architecture.
- Batatia, I. et al. (2022). MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields. NeurIPS. A widely adopted, more expressive successor architecture.
Current state
The field has largely converged on equivariant graph neural networks trained end-to-end so that force is guaranteed to be the exact gradient of a predicted energy, precisely to avoid the failure mode demonstrated in the developer example. Foundation-model-style interatomic potentials trained across large swaths of the periodic table (rather than one material at a time) are an active area, aiming for the same "train once, use broadly" property that neural operators bring to differential equations. Deploying any such model on a real materials or drug-design decision still requires validating against experiment or high-accuracy quantum chemistry on the specific system of interest — a trained potential is a tool for exploring a huge space quickly, not a replacement for that final check.
What to learn next
- Neural operators — the closely related idea of training once to generalise across many cases, applied here to atomic energies.
- Molecular property prediction — using a trained potential's outputs as features for a downstream prediction task.
- Backpropagation — the general mechanism automatic differentiation relies on to compute exact forces.