AI for Science and Engineering

Physics-informed neural networks

A physics-informed neural network is a network that is graded on obeying a physics equation, not only on matching examples.

On this page 6
  1. Why it exists
  2. How it works
  3. Where this shows up
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A physics-informed neural network is a network graded on obeying a physics rule, not only on matching examples.

Think about a school exam where the rule is "your answers must add up to 100 marks total." A student could write any set of numbers that happens to add up to 100. They could still pass that particular check, even without solving a single question correctly.

A physics-informed neural network is graded the same way. It is penalised whenever its output breaks a known physics equation, in addition to being penalised for disagreeing with any data it is given.

Why it exists

An ordinary neural network learns purely from examples. Show it a thousand correct answers, and it learns to imitate the pattern between them. That works well when a thousand correct answers are cheap to get.

In science and engineering, correct answers are often not cheap. A single measurement from a real experiment, or a single run of an expensive simulator (see surrogate models), can cost real time and money. Sometimes almost no measurements exist at all for the exact situation you care about.

But physics does not go silent because data is scarce. Many physical processes — how heat spreads, how a fluid flows, how a solid bends — are described by a known equation. That equation holds even when nobody has measured this particular case. A physics-informed neural network, usually shortened to PINN, uses that equation itself as an extra source of training signal, on top of whatever real data exists. It learns from measurements when they are available, and from the equation everywhere else.

How it works

Known physics equation, e.g. "the rate of change equals minus itself"
        |
        v
Pick many points where the equation should hold
        |
        v
Network's own output, checked against the equation at those points
        |                                   |
        v                                   v
"How much does the network      "How much does the network's
 disagree with any real data?"   output disagree with the equation?"
        |                                   |
        +-------------------+---------------+
                             v
                    Combined training signal

The network is never shown the correct answer directly at most points. It is shown the rule the correct answer must obey, and works backward from that rule. It uses the same learning process ordinary networks use — see how neural networks learn.

Where this shows up

Engineers use physics-informed networks to estimate temperature or stress inside a material at points where no sensor was placed. They fill in the gaps between a handful of real measurements, using the heat or stress equation as a guide. Researchers modelling blood flow use them to estimate flow patterns inside vessels that are hard to measure directly, constrained by the known equations of fluid motion.

An honest warning

A physics-informed network is only as correct as the equation it is given. If the equation is a simplified version of reality — and most equations used in practice are simplifications — the network faithfully learns the simplification, confidently. It has no way of knowing it is missing something the real world does not skip.

Remember this

  • A physics-informed neural network is penalised for disagreeing with a known equation, not only for disagreeing with data.
  • It lets a model learn from very little — or even zero — real measurements, because the equation supplies extra training signal everywhere.
  • It is only as good as the equation itself. A simplified equation produces a confidently simplified model.

What to learn next

Developer — Code and libraries.

The clearest way to see what a PINN does differently: solve an equation with a network that is never once shown the correct numerical answer.

Setup

bash
pip install torch

Minimal runnable code

The equation is dy/dx = -y with y(0) = 1 — a rate of change proportional to the current value, with a known starting point. Its exact solution is y(x) = e^-x, used here only to check our work, never given to the network as training data.

pinn.py
import torch
import torch.nn as nn

torch.manual_seed(0)

# the ODE we want to solve, without ever giving the network a single (x, y) answer:
#   dy/dx = -y ,  y(0) = 1
# analytic solution, for checking our work only: y(x) = exp(-x)

net = nn.Sequential(nn.Linear(1, 16), nn.Tanh(), nn.Linear(16, 16), nn.Tanh(), nn.Linear(16, 1))
optimizer = torch.optim.Adam(net.parameters(), lr=0.01)

x_physics = torch.linspace(0, 3, 30).reshape(-1, 1).requires_grad_(True)
x_zero = torch.zeros(1, 1)

for step in range(2000):
    optimizer.zero_grad()

    y_pred = net(x_physics)
    # dy/dx via autograd -- the network's own output, differentiated with respect to its input
    dy_dx = torch.autograd.grad(y_pred, x_physics, grad_outputs=torch.ones_like(y_pred),
                                 create_graph=True)[0]
    physics_residual = dy_dx + y_pred          # should be zero if the ODE is satisfied
    physics_loss = (physics_residual ** 2).mean()

    initial_condition_loss = (net(x_zero) - 1.0) ** 2

    loss = physics_loss + initial_condition_loss.squeeze()
    loss.backward()
    optimizer.step()

    if step % 500 == 0:
        print(f"step {step:4d}  physics_loss={physics_loss.item():.6f}  "
              f"ic_loss={initial_condition_loss.item():.6f}")

print()
test_x = torch.tensor([[0.0], [1.0], [2.0]])
predicted = net(test_x).detach().squeeze().tolist()
analytic = [1.0, 0.3679, 0.1353]
for x, p, a in zip([0, 1, 2], predicted, analytic):
    print(f"x={x}  network={p:.4f}  analytic e^-x={a:.4f}")
Output
step    0  physics_loss=0.191191  ic_loss=0.406332
step  500  physics_loss=0.000005  ic_loss=0.000000
step 1000  physics_loss=0.000001  ic_loss=0.000000
step 1500  physics_loss=0.000039  ic_loss=0.000034

x=0  network=1.0000  analytic e^-x=1.0000
x=1  network=0.3678  analytic e^-x=0.3679
x=2  network=0.1354  analytic e^-x=0.1353

Walkthrough

x_physics.requires_grad_(True) is the line that makes this a PINN instead of an ordinary network. It tells PyTorch to track how the network's output changes as this input changes, so dy/dx can be computed exactly — not estimated, not looked up from a table, computed as a true derivative of the network's own function.

torch.autograd.grad(y_pred, x_physics, ..., create_graph=True) asks for that derivative and, critically, keeps it attached to the computation graph. create_graph=True is what allows physics_residual to be backpropagated through — without it, the physics loss could be computed but never used to update the network.

physics_residual = dy_dx + y_pred is the ODE itself, rearranged so the correct answer makes it zero. Squaring and averaging it turns "how far from zero" into an ordinary loss the optimiser can push down.

The result: by the final steps, physics_loss has fallen to almost nothing, and the network's predictions at x=0,1,2 land within 0.01% of the true e^-x curve — having been given zero direct (x, y) training pairs beyond the single starting point.

Common mistakes

Forgetting requires_grad_(True) on the input. Without it, torch.autograd.grad has nothing to differentiate with respect to, and raises an error about the tensor not requiring gradients.

Forgetting create_graph=True. The derivative computes fine, but it becomes a dead end — it cannot flow back into updating the network's weights, so the physics loss stays flat forever.

Weighting the physics loss and data loss badly. Real PINNs often need to scale one loss term up or down relative to the other, since an equation with steep gradients can dominate training and starve the network of attention on boundary or initial conditions. This example has only one equation and one boundary condition, so it is not visible here — it becomes a real tuning problem on harder equations.

Treating a converged loss as a converged solution. A physics loss near zero means the network satisfies the equation at the points it was checked at, not at every point in between. Always spot-check on a finer grid than the one used in training.

Try it yourself

Change the equation to dy/dx = -2*y, whose true solution is y(x) = e^(-2x). Update physics_residual to dy_dx + 2 * y_pred and rerun. The network should converge to the new, steeper curve — using exactly the same training loop, because nothing about the loop is specific to one equation.

What to learn next

  • Neural operators — training one network to solve an entire family of equations, not only this one instance.
  • Loss functions — how residual terms like physics_residual become trainable objectives.
  • Gradient descent — the optimisation process driving both loss terms down together.

Researcher — Mathematics and papers.

The formal setting

Consider a differential equation N[u](x) = 0 for x in domain Ω, with boundary or initial conditions B[u](x) = 0 for x on ∂Ω. u is the unknown solution function, N is a differential operator (possibly nonlinear), and B encodes known values on the boundary.

A PINN approximates u with a neural network u_θ(x), parameterised by weights θ. The training loss is:

L(θ) = w_r · (1/N_r) Σ_i  ( N[u_θ](x_r_i) )^2
     + w_b · (1/N_b) Σ_j  ( B[u_θ](x_b_j) )^2
     + w_d · (1/N_d) Σ_k  ( u_θ(x_d_k) - u_k )^2
  • x_r_i — "collocation points," sampled inside Ω, where the equation residual is checked
  • x_b_j — points on the boundary ∂Ω, where boundary/initial conditions are checked
  • x_d_k, u_k — any real data pairs available, which may be empty (N_d = 0)
  • w_r, w_b, w_d — weights balancing the three terms, tuned per problem
  • N[u_θ](x) — the differential operator applied to the network's own output, evaluated via automatic differentiation, exactly as the developer example computed dy/dx

The developer example is the special case N[u] = u' + u, B[u] = u(0) - 1, and N_d = 0 — no data term at all.

Why automatic differentiation matters here

Computing N[u_θ] for equations involving second or higher derivatives (common in physics — wave equations, elasticity, diffusion) requires differentiating the network's output with respect to its input multiple times. Automatic differentiation gives this exactly, at machine precision, as opposed to finite-difference approximation, which introduces discretisation error and requires choosing a step size. This exactness is the core mechanical reason PINNs became practical once general-purpose automatic differentiation frameworks matured.

Complexity and known failure modes

Training cost scales with the number of collocation points N_r and the cost of computing N[u_θ], which itself scales with the order of the derivatives required — a second-order PDE roughly doubles the backward passes needed per collocation point compared to a first-order one.

PINNs are known to suffer from spectral bias: gradient-descent-trained networks tend to fit low-frequency components of a solution before high-frequency ones (Rahaman et al., 2019), which makes PINNs slow to converge on solutions with sharp features or high-frequency oscillations. Wang, Teng and Perdikaris (2021) documented gradient pathologies where the different loss terms have wildly different gradient magnitudes, causing the optimiser to effectively ignore the boundary or data terms unless the weights w_r, w_b, w_d are carefully tuned or adaptively rebalanced.

Papers

  • Raissi, M., Perdikaris, P. and Karniadakis, G. (2019). Physics-Informed Neural Networks. Journal of Computational Physics 378. The paper that established the name and the modern formulation.
  • Rahaman, N. et al. (2019). On the Spectral Bias of Neural Networks. ICML. Explains the low-frequency-first learning dynamic that limits PINNs on sharp solutions.
  • Wang, S., Teng, Y. and Perdikaris, P. (2021). Understanding and Mitigating Gradient Pathologies in Physics-Informed Neural Networks. SIAM Journal on Scientific Computing.

Current state

Plain PINNs remain difficult to scale to complex, high-dimensional, or multi-physics problems, and are usually slower to train per-instance than a classical numerical solver run once. Their advantage is not raw speed on one fixed problem — it is flexibility with sparse or noisy data, and smooth handling of inverse problems (inferring an unknown physical parameter from partial data). Active research directions include adaptive loss weighting, domain decomposition for large domains, and — the subject of the next lesson — architectures that learn the solution operator for an entire family of equations at once, amortising the cost across many future queries instead of solving one instance at a time.

What to learn next