Neural radiance fields
A NeRF stores a scene as a function that answers "what colour and how solid is this point in space", and renders new views by marching rays through it.
- 18 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A NeRF stores a whole 3D scene inside a small neural network, then shows it from angles nobody photographed.
NeRF stands for neural radiance field: a network holding a light value for every point in space.
The analogy you have already lived
Think of fog lit by a street lamp. You cannot point at a surface. What reaches your eye is the sum of light from every wisp of fog along the way.
Now imagine the fog is thick in some places and absent in others. Where it is thick, that patch is what you see, and everything behind it is hidden. Where it is absent, your gaze passes through.
A NeRF describes the whole world as that kind of fog. Solid objects are places where the fog is extremely thick.
Why it exists
The 3D models from structure from motion are clouds of points, or triangles stitched over them. Those struggle with real scenes.
Thin things like leaves and wires fall between the triangles. Glass and shiny metal do not have a single colour. Fur and smoke have no surface at all.
NeRF sidesteps the whole question. It never asks "where is the surface". It only asks two things about every point in space: how solid is it, and what colour is it.
How a picture gets made
pick a pixel of the picture you want
|
shoot a ray from the camera through that pixel
|
take about a hundred sample points along the ray
|
ask the network at each one: how solid? what colour?
|
walk along the ray adding up colour, but stop
contributing once the ray is blocked
|
that total is the pixel's colourThe last part is the important bit. If the ray hits something solid at step 20, whatever is at step 30 is behind it and contributes nothing. That is handled automatically by tracking how much light is still getting through.
How it learns
You have a hundred real photos and you know where each camera was, from structure from motion.
The network starts knowing nothing. Render what it thinks a view should look like, compare with the real photo, and nudge it.
Every step is differentiable, so this ordinary training loop teaches the network a whole scene. Nothing else is needed — no 3D scanner, no depth sensor.
The trick that made it work
The first version of this idea produced blurry mush. The scene came out smooth and lifeless.
The cause is subtle. A plain neural network fed raw positions has a strong bias toward smooth outputs, so fine detail never appears.
The fix was to stop feeding raw positions. Instead, feed a set of waves at rising frequencies computed from the position. Suddenly the network has sharp detail to build from.
In the run below, that one change takes the fit from 22 to 49 on a quality score. Higher is better, and this is the difference between a blur and a photograph.
The honest part
NeRFs are slow, and the numbers say why.
One 800 by 800 picture is 640,000 rays. About 192 samples each. That is over 122 million network calls, for one frame. Training runs through a hundred images, tens of thousands of times.
The original work needed a day or two on a strong GPU per scene. Nearly every follow-up paper since has been about making that number smaller.
Where you have already seen it
- Property listings letting you fly through a room.
- Museums publishing walkable captures of artefacts.
- Film production replacing a set with a captured location.
- Map apps with smooth 3D flythroughs of landmarks.
Remember this
- A NeRF answers "how solid and what colour" for every point in space.
- Rendering marches a ray and adds up colour until the light is blocked.
- Feeding the network waves instead of raw positions is what made it sharp.
What to learn next
- Gaussian splatting — the same rendering maths, made real-time.
- Structure from motion — where the camera poses come from.
- Diffusion models — the other family now generating 3D scenes.
Developer — Code and libraries.
Setup
pip install torch==2.5.1 numpy==1.26.4Training a real NeRF needs a GPU and a dataset. Its two central ideas do not, and both are measurable on a laptop CPU. This file takes about ninety seconds.
Volume rendering, and the encoding that made NeRF sharp
import numpy as np
import torch
import torch.nn as nn
print("PART 1 - volume rendering, the equation NeRF actually optimises\n")
def render_ray(density, colour, dt=0.05):
"""March along a ray: accumulate colour, weighted by how much light survives."""
alpha = 1.0 - np.exp(-density * dt) # chance of hitting something in this step
transmittance = np.cumprod(np.concatenate([[1.0], 1.0 - alpha[:-1]]))
weights = transmittance * alpha # how much this step contributes
return (weights * colour).sum(), weights, transmittance
steps = 40
z = np.arange(steps) * 0.05
cases = {
"empty space": (np.zeros(steps), np.ones(steps)),
"a hard surface at step 20": (np.where(np.arange(steps) == 20, 200.0, 0.0),
np.full(steps, 0.8)),
"thin fog everywhere": (np.full(steps, 1.0), np.full(steps, 0.4)),
"fog, then a surface": (np.where(np.arange(steps) < 20, 1.0, 0.0)
+ np.where(np.arange(steps) == 30, 200.0, 0.0),
np.where(np.arange(steps) < 20, 0.3, 0.9)),
}
for name, (d, c) in cases.items():
value, w, T = render_ray(d, c)
print(f" {name:<28} rendered value {value:.4f} weights sum to {w.sum():.4f}")
print(f" {'':<28} light left at the end: {T[-1]:.4f}")
print("\n the weights are a probability distribution over 'where the ray stopped'")
print(" an opaque surface concentrates all of it at one step; fog spreads it out\n")
print("PART 2 - why NeRF needs positional encoding\n")
torch.manual_seed(0)
H = W = 48
yy, xx = np.meshgrid(np.linspace(-1, 1, H), np.linspace(-1, 1, W), indexing="ij")
target = (0.5 + 0.5 * np.sin(9 * xx) * np.cos(7 * yy)
+ 0.3 * np.sin(23 * xx + 15 * yy)) / 1.6 # a deliberately detailed pattern
target = torch.tensor(np.clip(target, 0, 1), dtype=torch.float32).reshape(-1, 1)
coords = torch.tensor(np.stack([xx.ravel(), yy.ravel()], 1), dtype=torch.float32)
def encode(p, levels):
"""gamma(p) = [p, sin(2^0 pi p), cos(2^0 pi p), ... ] -- Fourier features."""
if levels == 0:
return p
out = [p]
for i in range(levels):
out += [torch.sin(2.0 ** i * np.pi * p), torch.cos(2.0 ** i * np.pi * p)]
return torch.cat(out, dim=-1)
def fit(levels, steps=2500):
torch.manual_seed(0)
x = encode(coords, levels)
net = nn.Sequential(nn.Linear(x.shape[1], 128), nn.ReLU(),
nn.Linear(128, 128), nn.ReLU(),
nn.Linear(128, 1), nn.Sigmoid())
opt = torch.optim.Adam(net.parameters(), lr=3e-3)
for _ in range(steps):
loss = ((net(x) - target) ** 2).mean()
opt.zero_grad()
loss.backward()
opt.step()
with torch.no_grad():
mse = ((net(x) - target) ** 2).mean().item()
return mse, x.shape[1], net
print(" fitting the same tiny network to the same picture, changing only the input")
print(f" {'encoding':<26}{'inputs':<10}{'final MSE':<14}{'PSNR (dB)'}")
for levels in [0, 2, 4, 6, 8]:
mse, dim, _ = fit(levels)
label = "raw (x, y)" if levels == 0 else f"{levels} frequency bands"
print(f" {label:<26}{dim:<10}{mse:<14.6f}{-10 * np.log10(mse):.2f}")
print("\n a plain MLP on raw coordinates is biased toward smooth functions")
print(" sine and cosine at rising frequencies hand it the detail directly\n")
mse, _, net = fit(6)
with torch.no_grad():
got = net(encode(coords, 6)).reshape(H, W).numpy()
truth = target.reshape(H, W).numpy()
print(" what 6 bands reproduces (left: target, right: fitted)")
chars = " .:-=+*#%@"
for r in range(0, H, 3):
a = "".join(chars[min(int(v * 9.99), 9)] for v in truth[r, ::2])
b = "".join(chars[min(int(v * 9.99), 9)] for v in got[r, ::2])
print(f" {a} {b}")
print("\nPART 3 - the cost of the real thing")
rays = 800 * 800
samples = 192
print(f" one 800x800 image = {rays:,} rays")
print(f" at {samples} samples per ray = {rays * samples:,} network evaluations")
print(" per image. A NeRF trains on 100 images for tens of thousands of steps.")
print(" that is why the original paper needed a day or two on a good GPU,")
print(" and why every follow-up paper is about making this number smaller.")PART 1 - volume rendering, the equation NeRF actually optimises
empty space rendered value 0.0000 weights sum to 0.0000
light left at the end: 1.0000
a hard surface at step 20 rendered value 0.8000 weights sum to 1.0000
light left at the end: 0.0000
thin fog everywhere rendered value 0.3459 weights sum to 0.8647
light left at the end: 0.1423
fog, then a surface rendered value 0.5207 weights sum to 1.0000
light left at the end: 0.0000
the weights are a probability distribution over 'where the ray stopped'
an opaque surface concentrates all of it at one step; fog spreads it out
PART 2 - why NeRF needs positional encoding
fitting the same tiny network to the same picture, changing only the input
encoding inputs final MSE PSNR (dB)
raw (x, y) 2 0.006359 21.97
2 frequency bands 10 0.000125 39.02
4 frequency bands 18 0.000083 40.83
6 frequency bands 26 0.000012 49.16
8 frequency bands 34 0.000113 39.49
a plain MLP on raw coordinates is biased toward smooth functions
sine and cosine at rising frequencies hand it the detail directly
what 6 bands reproduces (left: target, right: fitted)
.: ++-+- .-.+#-:- ==-* .: ++-+- .-.+#-:- ==-*
- -==#+ .. -#=== :--#* - -==#+ .. -#=== :--#*
. -:.++.-- :+--+: -:.++: . -:.++.-- :+--+: -:.++:
:+=.-: -=:=+..-..+=:=- : :+=.-: -=:=+..-..+=:=- :
**=+- -:+#-.: ++=*= : **=+- -:+#-.: ++=*= :
=-*+ .: -#=-= ---*+..: =-*+ .: -#=-= ---*+..:
.==.-=.:+::+:.=-.==.-=.: .==.-=.:+::+:.=-.==.-=.:
-: :+-=*. :..**:-: .=-=* -: :+-=*. :..**:-: .=-=*
- -:+#-.: +*=+= :.=%= - -:+#-.: +*=+= :.=%=
.-.:*-:=. -=:++..- :*=: .-.:*-:=. -=:++..- :*=:
-+::=. =-:+=.:- -+-:=. - -+::=. =-:+=.:- -+-:=. -
*=+*. ..:**:-: .+=+*: .. *=+*. ..:**:-: .+=+*: ..
-=#- : =*-== -:=#=.: -=#- : =*-== -:=#=.:
:+:.=-.==.-=.:=.:+-.=- - :+:.=- ==.-=.:=.:+-.=- -
= -=-*+ .: -*=-=. --:*+ = -=-*+ .: -*=-=. --:*+
. .::*#::: +==*: ...*#- . .::*#::: +==*: ...*#-
PART 3 - the cost of the real thing
one 800x800 image = 640,000 rays
at 192 samples per ray = 122,880,000 network evaluations
per image. A NeRF trains on 100 images for tens of thousands of steps.
that is why the original paper needed a day or two on a good GPU,
and why every follow-up paper is about making this number smaller.Reading the output carefully
empty space renders 0.0000 and leaves 1.0000 of the light. Nothing absorbed, nothing emitted. Correct, and worth checking in any implementation.
a hard surface at step 20 renders exactly 0.8000, the surface's colour, and the weights sum to 1.0000. With density 200 and step 0.05, alpha is $1 - e^{-10} = 0.99995$. The ray is stopped dead, and everything past step 20 contributes nothing. That is occlusion, and nobody wrote an occlusion rule — it emerged from the transmittance product.
thin fog everywhere renders 0.3459 with weights summing to 0.8647. Weights below 1 means the ray escaped without hitting anything: 0.1423 of the light passed straight through. In a real NeRF that residual is where the background colour goes.
fog, then a surface renders 0.5207, between the fog's 0.3 and the surface's 0.9. The fog contributed some colour and dimmed what came after. Semi-transparency, from the same three lines of code.
The encoding table is the headline: 21.97 dB to 49.16 dB. Same network, same data, same optimiser, same number of steps. The only change is the input representation. That is a factor of 500 in mean squared error, and it is what the ASCII comparison shows — the two panels are essentially indistinguishable at 6 bands.
8 bands is worse than 6: 39.49 against 49.16. This is real, and the NeRF paper reports the same effect. Frequencies above what the data contains give the network capacity to fit noise and make optimisation harder. The number of bands is a hyperparameter with a genuine optimum, not a "more is better" dial.
122,880,000 network evaluations for one 800x800 frame. Print that number once and every design decision in the follow-up literature becomes obvious.
What the full model looks like
The scene function takes a 3D position and a viewing direction, and returns a density and a colour:
$$ F_\Theta : (\mathbf{x}, \mathbf{d}) \mapsto (\sigma, \mathbf{c}) $$
Two details of the architecture carry real meaning.
Density depends on position only. The direction is injected late, after $\sigma$ has been produced. Physically, how solid a point is cannot depend on where you look from. Enforcing that in the architecture prevents the model from cheating by making a surface appear and disappear per view.
Colour depends on direction. That is how specular highlights and reflections are represented — the same point genuinely looks different from different angles.
The direction gets its own positional encoding, with far fewer bands (4 against 10 for position), because view-dependent effects vary slowly.
Running one for real
pip install nerfstudio
ns-process-data images --data ./photos --output-dir ./scene # runs COLMAP
ns-train nerfacto --data ./scene
ns-viewer --load-config ./outputs/.../config.ymlnerfstudio is the practical entry point. nerfacto is its default recipe, combining hash-grid encoding, proposal sampling and camera-pose refinement. A GPU with 8 GB or more is the realistic minimum.
The ns-process-data step is structure from motion — the poses have to come from somewhere, and a bad pose estimate produces a blurry NeRF that no amount of training fixes.
Common mistakes
Bad camera poses. The single most common cause of a blurry result. Check the COLMAP reconstruction before training.
Photographing from a straight line. A NeRF needs views that surround the subject. Walking past it in a line gives no constraint on the far side, and the network fills it with whatever renders the training views correctly.
Autofocus and auto-exposure on. Changing intrinsics between frames break the pose model, and changing exposure makes the same surface a different colour from different views, which the network resolves by inventing view-dependent effects.
Moving objects or people in the scene. A NeRF assumes a static world. Anything that moves becomes a floating cloud.
Reading PSNR without looking at renders. PSNR rewards blur. A slightly blurry render can outscore a sharper one with small misalignments. Report LPIPS and SSIM as well, and look at the images.
Try it yourself
Change levels from 6 to 12 and see the fit get worse still, and take longer. Then reduce the target's detail — change sin(23*xx + 15*yy) to sin(3*xx + 2*yy) — and rerun the sweep. The optimum shifts to fewer bands, because the signal now contains less high-frequency content. The right number of bands is a property of the scene, not the model.
What to learn next
- Gaussian splatting — the same rendering maths, made real-time.
- Structure from motion — where the camera poses come from.
- Diffusion models — the other family now generating 3D scenes.
Researcher — Mathematics and papers.
The formulation
Mildenhall et al. (ECCV 2020), NeRF: representing scenes as neural radiance fields for view synthesis. A scene is a continuous function
$$ F_\Theta : (\mathbf{x}, \mathbf{d}) \in \mathbb{R}^3 \times \mathbb{S}^2 \;\mapsto\; (\sigma, \mathbf{c}) \in \mathbb{R}^+ \times \mathbb{R}^3 $$
Where $\mathbf{x}$ is a 3D position, $\mathbf{d}$ a unit viewing direction, $\sigma$ the volume density, and $\mathbf{c}$ the emitted radiance.
The expected colour of a ray $\mathbf{r}(t) = \mathbf{o} + t\mathbf{d}$ between near and far bounds is the classical volume rendering integral (Kajiya and Von Herzen, 1984):
$$ C(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma(\mathbf{r}(t))\,\mathbf{c}(\mathbf{r}(t), \mathbf{d})\, dt, \qquad T(t) = \exp!\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(s))\, ds\right) $$
$T(t)$ is the transmittance: the probability that the ray travels from $t_n$ to $t$ without being absorbed.
The quadrature used in practice, and implemented in the code above, is:
$$ \hat{C}(\mathbf{r}) = \sum_{i=1}^{N} T_i \bigl(1 - e^{-\sigma_i \delta_i}\bigr) \mathbf{c}_i, \qquad T_i = \exp!\left(-\sum_{j<i}\sigma_j \delta_j\right) $$
with $\delta_i = t_{i+1} - t_i$. This is exactly alpha compositing with $\alpha_i = 1 - e^{-\sigma_i \delta_i}$, and it is differentiable in every term — which is the whole reason the scheme is trainable.
The weights $w_i = T_i \alpha_i$ form a distribution over "where along the ray the ray terminated", normalised only when the ray is fully opaque. Expected depth is $\sum_i w_i t_i$, and the weight entropy is a useful measure of how confident the model is about a surface.
Positional encoding
Tancik et al. (NeurIPS 2020), Fourier features let networks learn high frequency functions in low dimensional domains, gives the theory the empirical result above demonstrates.
An MLP with ReLU activations has a neural tangent kernel that, on low-dimensional inputs, decays rapidly with frequency. The network is therefore heavily biased toward low-frequency functions — spectral bias — and fitting fine detail requires impractically long training.
Mapping the input through
$$ \gamma(p) = \bigl[\sin(2^0\pi p), \cos(2^0\pi p), \dots, \sin(2^{L-1}\pi p), \cos(2^{L-1}\pi p)\bigr] $$
makes the composed kernel stationary and widens its spectrum, so high frequencies become learnable. Tancik et al. show that a random Gaussian Fourier mapping with a tuned bandwidth generalises this, and that the bandwidth — not the number of features — is the parameter that matters.
The 8-band regression in the output is the predicted over-bandwidth failure: frequencies above the signal's own content add variance without signal.
NeRF uses $L = 10$ for position and $L = 4$ for direction.
Hierarchical sampling
Uniform sampling wastes almost every sample in empty space. NeRF trains two networks: a coarse one sampled uniformly, then a fine one sampled from the piecewise-constant PDF given by the coarse network's weights $w_i$, via inverse transform sampling.
This is importance sampling of the rendering integral, and it typically improves quality more than doubling the sample count would. Its descendants — proposal networks in Mip-NeRF 360 and nerfacto — do the same job with a small dedicated density network.
The efficiency lineage
The 122-million-evaluation figure explains the whole research programme.
- Instant-NGP (Müller et al., SIGGRAPH 2022) replaces the frequency encoding with a multi-resolution hash grid of learnable feature vectors. Most of the capacity moves from network weights into the grid, so the MLP shrinks to two small layers. Training drops from days to seconds, with a fused CUDA implementation.
- Plenoxels (Fridovich-Keil et al., CVPR 2022) removes the network entirely: a sparse voxel grid of spherical-harmonic coefficients, optimised directly. It establishes that the neural network was never the essential ingredient — differentiable volume rendering was.
- TensoRF factorises the volume into low-rank tensor components, reducing memory sharply.
Quality, and the 360 problem
Mip-NeRF (Barron et al., ICCV 2021) replaces point samples with integrated positional encoding over a conical frustum, giving anti-aliasing and multi-scale consistency. Mip-NeRF 360 (CVPR 2022) adds a space contraction for unbounded scenes plus a distortion regulariser that suppresses the floating artefacts typical of outdoor captures. Zip-NeRF (ICCV 2023) combines that with hash grids.
Unbounded outdoor scenes remain the hard case, and the failure mode is characteristic: correct-looking training views plus "floaters", small clouds of density that satisfy the training views and appear from nowhere in new ones. They are an under-constrained-region symptom, not a bug.
Evaluation
PSNR, SSIM and LPIPS on held-out views, typically on Blender synthetic objects, LLFF forward-facing scenes, and Mip-NeRF 360 outdoor scenes.
Two cautions. PSNR is dominated by low frequencies and rewards blur, so a method can gain PSNR while looking worse; LPIPS (Zhang et al., CVPR 2018) tracks human judgement better and should be reported alongside. And held-out views in these datasets are interpolations between training views — genuine extrapolation away from the capture trajectory is much harder, and rarely measured.
Where NeRF stands after gaussian splatting
3D Gaussian splatting (2023) overtook NeRF on the metric that mattered commercially: real-time rendering. It is covered in the next lesson.
NeRF-style implicit fields remain preferable in several places, and it is worth being precise about which:
- Surface extraction. Density fields convert to meshes via marching cubes; SDF-based variants (VolSDF, NeuS, Neuralangelo) give clean geometry that point-based representations struggle with.
- Compactness. A hash-grid NeRF is often smaller on disk than the millions of Gaussians for the same scene.
- Continuity. A continuous field can be queried anywhere and regularised analytically, which matters for physics, inverse rendering and relighting.
The deeper point is that both share the same rendering equation. Gaussian splatting replaces "query a network along a ray" with "sort and blend primitives", and the alpha-compositing arithmetic in the code above is unchanged.
What to learn next
- Gaussian splatting — the same rendering maths, made real-time.
- Structure from motion — where the camera poses come from.
- Diffusion models — the other family now generating 3D scenes.