3D Vision and Depth

Neural radiance fields

A NeRF stores a scene as a function that answers "what colour and how solid is this point in space", and renders new views by marching rays through it.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How a picture gets made
  5. How it learns
  6. The trick that made it work
  7. The honest part
  8. Where you have already seen it
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A NeRF stores a whole 3D scene inside a small neural network, then shows it from angles nobody photographed.

NeRF stands for neural radiance field: a network holding a light value for every point in space.

The analogy you have already lived

Think of fog lit by a street lamp. You cannot point at a surface. What reaches your eye is the sum of light from every wisp of fog along the way.

Now imagine the fog is thick in some places and absent in others. Where it is thick, that patch is what you see, and everything behind it is hidden. Where it is absent, your gaze passes through.

A NeRF describes the whole world as that kind of fog. Solid objects are places where the fog is extremely thick.

Why it exists

The 3D models from structure from motion are clouds of points, or triangles stitched over them. Those struggle with real scenes.

Thin things like leaves and wires fall between the triangles. Glass and shiny metal do not have a single colour. Fur and smoke have no surface at all.

NeRF sidesteps the whole question. It never asks "where is the surface". It only asks two things about every point in space: how solid is it, and what colour is it.

How a picture gets made

   pick a pixel of the picture you want
        |
   shoot a ray from the camera through that pixel
        |
   take about a hundred sample points along the ray
        |
   ask the network at each one: how solid? what colour?
        |
   walk along the ray adding up colour, but stop
   contributing once the ray is blocked
        |
   that total is the pixel's colour

The last part is the important bit. If the ray hits something solid at step 20, whatever is at step 30 is behind it and contributes nothing. That is handled automatically by tracking how much light is still getting through.

How it learns

You have a hundred real photos and you know where each camera was, from structure from motion.

The network starts knowing nothing. Render what it thinks a view should look like, compare with the real photo, and nudge it.

Every step is differentiable, so this ordinary training loop teaches the network a whole scene. Nothing else is needed — no 3D scanner, no depth sensor.

The trick that made it work

The first version of this idea produced blurry mush. The scene came out smooth and lifeless.

The cause is subtle. A plain neural network fed raw positions has a strong bias toward smooth outputs, so fine detail never appears.

The fix was to stop feeding raw positions. Instead, feed a set of waves at rising frequencies computed from the position. Suddenly the network has sharp detail to build from.

In the run below, that one change takes the fit from 22 to 49 on a quality score. Higher is better, and this is the difference between a blur and a photograph.

The honest part

NeRFs are slow, and the numbers say why.

One 800 by 800 picture is 640,000 rays. About 192 samples each. That is over 122 million network calls, for one frame. Training runs through a hundred images, tens of thousands of times.

The original work needed a day or two on a strong GPU per scene. Nearly every follow-up paper since has been about making that number smaller.

Where you have already seen it

  • Property listings letting you fly through a room.
  • Museums publishing walkable captures of artefacts.
  • Film production replacing a set with a captured location.
  • Map apps with smooth 3D flythroughs of landmarks.

Remember this

  • A NeRF answers "how solid and what colour" for every point in space.
  • Rendering marches a ray and adds up colour until the light is blocked.
  • Feeding the network waves instead of raw positions is what made it sharp.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch==2.5.1 numpy==1.26.4

Training a real NeRF needs a GPU and a dataset. Its two central ideas do not, and both are measurable on a laptop CPU. This file takes about ninety seconds.

Volume rendering, and the encoding that made NeRF sharp

nerf_core.py
import numpy as np
import torch
import torch.nn as nn

print("PART 1 - volume rendering, the equation NeRF actually optimises\n")


def render_ray(density, colour, dt=0.05):
    """March along a ray: accumulate colour, weighted by how much light survives."""
    alpha = 1.0 - np.exp(-density * dt)          # chance of hitting something in this step
    transmittance = np.cumprod(np.concatenate([[1.0], 1.0 - alpha[:-1]]))
    weights = transmittance * alpha              # how much this step contributes
    return (weights * colour).sum(), weights, transmittance


steps = 40
z = np.arange(steps) * 0.05
cases = {
    "empty space": (np.zeros(steps), np.ones(steps)),
    "a hard surface at step 20": (np.where(np.arange(steps) == 20, 200.0, 0.0),
                                  np.full(steps, 0.8)),
    "thin fog everywhere": (np.full(steps, 1.0), np.full(steps, 0.4)),
    "fog, then a surface": (np.where(np.arange(steps) < 20, 1.0, 0.0)
                            + np.where(np.arange(steps) == 30, 200.0, 0.0),
                            np.where(np.arange(steps) < 20, 0.3, 0.9)),
}
for name, (d, c) in cases.items():
    value, w, T = render_ray(d, c)
    print(f"   {name:<28} rendered value {value:.4f}   weights sum to {w.sum():.4f}")
    print(f"   {'':<28} light left at the end: {T[-1]:.4f}")
print("\n   the weights are a probability distribution over 'where the ray stopped'")
print("   an opaque surface concentrates all of it at one step; fog spreads it out\n")

print("PART 2 - why NeRF needs positional encoding\n")
torch.manual_seed(0)
H = W = 48
yy, xx = np.meshgrid(np.linspace(-1, 1, H), np.linspace(-1, 1, W), indexing="ij")
target = (0.5 + 0.5 * np.sin(9 * xx) * np.cos(7 * yy)
          + 0.3 * np.sin(23 * xx + 15 * yy)) / 1.6      # a deliberately detailed pattern
target = torch.tensor(np.clip(target, 0, 1), dtype=torch.float32).reshape(-1, 1)
coords = torch.tensor(np.stack([xx.ravel(), yy.ravel()], 1), dtype=torch.float32)


def encode(p, levels):
    """gamma(p) = [p, sin(2^0 pi p), cos(2^0 pi p), ... ] -- Fourier features."""
    if levels == 0:
        return p
    out = [p]
    for i in range(levels):
        out += [torch.sin(2.0 ** i * np.pi * p), torch.cos(2.0 ** i * np.pi * p)]
    return torch.cat(out, dim=-1)


def fit(levels, steps=2500):
    torch.manual_seed(0)
    x = encode(coords, levels)
    net = nn.Sequential(nn.Linear(x.shape[1], 128), nn.ReLU(),
                        nn.Linear(128, 128), nn.ReLU(),
                        nn.Linear(128, 1), nn.Sigmoid())
    opt = torch.optim.Adam(net.parameters(), lr=3e-3)
    for _ in range(steps):
        loss = ((net(x) - target) ** 2).mean()
        opt.zero_grad()
        loss.backward()
        opt.step()
    with torch.no_grad():
        mse = ((net(x) - target) ** 2).mean().item()
    return mse, x.shape[1], net


print("   fitting the same tiny network to the same picture, changing only the input")
print(f"   {'encoding':<26}{'inputs':<10}{'final MSE':<14}{'PSNR (dB)'}")
for levels in [0, 2, 4, 6, 8]:
    mse, dim, _ = fit(levels)
    label = "raw (x, y)" if levels == 0 else f"{levels} frequency bands"
    print(f"   {label:<26}{dim:<10}{mse:<14.6f}{-10 * np.log10(mse):.2f}")
print("\n   a plain MLP on raw coordinates is biased toward smooth functions")
print("   sine and cosine at rising frequencies hand it the detail directly\n")

mse, _, net = fit(6)
with torch.no_grad():
    got = net(encode(coords, 6)).reshape(H, W).numpy()
truth = target.reshape(H, W).numpy()
print("   what 6 bands reproduces (left: target, right: fitted)")
chars = " .:-=+*#%@"
for r in range(0, H, 3):
    a = "".join(chars[min(int(v * 9.99), 9)] for v in truth[r, ::2])
    b = "".join(chars[min(int(v * 9.99), 9)] for v in got[r, ::2])
    print(f"   {a}   {b}")

print("\nPART 3 - the cost of the real thing")
rays = 800 * 800
samples = 192
print(f"   one 800x800 image = {rays:,} rays")
print(f"   at {samples} samples per ray = {rays * samples:,} network evaluations")
print("   per image. A NeRF trains on 100 images for tens of thousands of steps.")
print("   that is why the original paper needed a day or two on a good GPU,")
print("   and why every follow-up paper is about making this number smaller.")
Output
PART 1 - volume rendering, the equation NeRF actually optimises

   empty space                  rendered value 0.0000   weights sum to 0.0000
                                light left at the end: 1.0000
   a hard surface at step 20    rendered value 0.8000   weights sum to 1.0000
                                light left at the end: 0.0000
   thin fog everywhere          rendered value 0.3459   weights sum to 0.8647
                                light left at the end: 0.1423
   fog, then a surface          rendered value 0.5207   weights sum to 1.0000
                                light left at the end: 0.0000

   the weights are a probability distribution over 'where the ray stopped'
   an opaque surface concentrates all of it at one step; fog spreads it out

PART 2 - why NeRF needs positional encoding

   fitting the same tiny network to the same picture, changing only the input
   encoding                  inputs    final MSE     PSNR (dB)
   raw (x, y)                2         0.006359      21.97
   2 frequency bands         10        0.000125      39.02
   4 frequency bands         18        0.000083      40.83
   6 frequency bands         26        0.000012      49.16
   8 frequency bands         34        0.000113      39.49

   a plain MLP on raw coordinates is biased toward smooth functions
   sine and cosine at rising frequencies hand it the detail directly

   what 6 bands reproduces (left: target, right: fitted)
   .:  ++-+- .-.+#-:-  ==-*   .:  ++-+- .-.+#-:-  ==-*
   -  -==#+ .. -#===  :--#*   -  -==#+ .. -#===  :--#*
   . -:.++.-- :+--+: -:.++:   . -:.++.-- :+--+: -:.++:
   :+=.-: -=:=+..-..+=:=- :   :+=.-: -=:=+..-..+=:=- :
   **=+-  -:+#-.:  ++=*=  :   **=+-  -:+#-.:  ++=*=  :
   =-*+ .: -#=-=  ---*+..:    =-*+ .: -#=-=  ---*+..: 
   .==.-=.:+::+:.=-.==.-=.:   .==.-=.:+::+:.=-.==.-=.:
   -: :+-=*. :..**:-: .=-=*   -: :+-=*. :..**:-: .=-=*
   -  -:+#-.:  +*=+=  :.=%=   -  -:+#-.:  +*=+=  :.=%=
    .-.:*-:=. -=:++..- :*=:    .-.:*-:=. -=:++..- :*=:
   -+::=. =-:+=.:- -+-:=. -   -+::=. =-:+=.:- -+-:=. -
   *=+*. ..:**:-: .+=+*: ..   *=+*. ..:**:-: .+=+*: ..
   -=#- :  =*-==  -:=#=.:     -=#- :  =*-==  -:=#=.:  
   :+:.=-.==.-=.:=.:+-.=- -   :+:.=- ==.-=.:=.:+-.=- -
   =  -=-*+ .: -*=-=. --:*+   =  -=-*+ .: -*=-=. --:*+
   . .::*#:::  +==*: ...*#-   . .::*#:::  +==*: ...*#-

PART 3 - the cost of the real thing
   one 800x800 image = 640,000 rays
   at 192 samples per ray = 122,880,000 network evaluations
   per image. A NeRF trains on 100 images for tens of thousands of steps.
   that is why the original paper needed a day or two on a good GPU,
   and why every follow-up paper is about making this number smaller.

Reading the output carefully

empty space renders 0.0000 and leaves 1.0000 of the light. Nothing absorbed, nothing emitted. Correct, and worth checking in any implementation.

a hard surface at step 20 renders exactly 0.8000, the surface's colour, and the weights sum to 1.0000. With density 200 and step 0.05, alpha is $1 - e^{-10} = 0.99995$. The ray is stopped dead, and everything past step 20 contributes nothing. That is occlusion, and nobody wrote an occlusion rule — it emerged from the transmittance product.

thin fog everywhere renders 0.3459 with weights summing to 0.8647. Weights below 1 means the ray escaped without hitting anything: 0.1423 of the light passed straight through. In a real NeRF that residual is where the background colour goes.

fog, then a surface renders 0.5207, between the fog's 0.3 and the surface's 0.9. The fog contributed some colour and dimmed what came after. Semi-transparency, from the same three lines of code.

The encoding table is the headline: 21.97 dB to 49.16 dB. Same network, same data, same optimiser, same number of steps. The only change is the input representation. That is a factor of 500 in mean squared error, and it is what the ASCII comparison shows — the two panels are essentially indistinguishable at 6 bands.

8 bands is worse than 6: 39.49 against 49.16. This is real, and the NeRF paper reports the same effect. Frequencies above what the data contains give the network capacity to fit noise and make optimisation harder. The number of bands is a hyperparameter with a genuine optimum, not a "more is better" dial.

122,880,000 network evaluations for one 800x800 frame. Print that number once and every design decision in the follow-up literature becomes obvious.

What the full model looks like

The scene function takes a 3D position and a viewing direction, and returns a density and a colour:

$$ F_\Theta : (\mathbf{x}, \mathbf{d}) \mapsto (\sigma, \mathbf{c}) $$

Two details of the architecture carry real meaning.

Density depends on position only. The direction is injected late, after $\sigma$ has been produced. Physically, how solid a point is cannot depend on where you look from. Enforcing that in the architecture prevents the model from cheating by making a surface appear and disappear per view.

Colour depends on direction. That is how specular highlights and reflections are represented — the same point genuinely looks different from different angles.

The direction gets its own positional encoding, with far fewer bands (4 against 10 for position), because view-dependent effects vary slowly.

Running one for real

bash
pip install nerfstudio
ns-process-data images --data ./photos --output-dir ./scene    # runs COLMAP
ns-train nerfacto --data ./scene
ns-viewer --load-config ./outputs/.../config.yml

nerfstudio is the practical entry point. nerfacto is its default recipe, combining hash-grid encoding, proposal sampling and camera-pose refinement. A GPU with 8 GB or more is the realistic minimum.

The ns-process-data step is structure from motion — the poses have to come from somewhere, and a bad pose estimate produces a blurry NeRF that no amount of training fixes.

Common mistakes

Bad camera poses. The single most common cause of a blurry result. Check the COLMAP reconstruction before training.

Photographing from a straight line. A NeRF needs views that surround the subject. Walking past it in a line gives no constraint on the far side, and the network fills it with whatever renders the training views correctly.

Autofocus and auto-exposure on. Changing intrinsics between frames break the pose model, and changing exposure makes the same surface a different colour from different views, which the network resolves by inventing view-dependent effects.

Moving objects or people in the scene. A NeRF assumes a static world. Anything that moves becomes a floating cloud.

Reading PSNR without looking at renders. PSNR rewards blur. A slightly blurry render can outscore a sharper one with small misalignments. Report LPIPS and SSIM as well, and look at the images.

Try it yourself

Change levels from 6 to 12 and see the fit get worse still, and take longer. Then reduce the target's detail — change sin(23*xx + 15*yy) to sin(3*xx + 2*yy) — and rerun the sweep. The optimum shifts to fewer bands, because the signal now contains less high-frequency content. The right number of bands is a property of the scene, not the model.

What to learn next

Researcher — Mathematics and papers.

The formulation

Mildenhall et al. (ECCV 2020), NeRF: representing scenes as neural radiance fields for view synthesis. A scene is a continuous function

$$ F_\Theta : (\mathbf{x}, \mathbf{d}) \in \mathbb{R}^3 \times \mathbb{S}^2 \;\mapsto\; (\sigma, \mathbf{c}) \in \mathbb{R}^+ \times \mathbb{R}^3 $$

Where $\mathbf{x}$ is a 3D position, $\mathbf{d}$ a unit viewing direction, $\sigma$ the volume density, and $\mathbf{c}$ the emitted radiance.

The expected colour of a ray $\mathbf{r}(t) = \mathbf{o} + t\mathbf{d}$ between near and far bounds is the classical volume rendering integral (Kajiya and Von Herzen, 1984):

$$ C(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma(\mathbf{r}(t))\,\mathbf{c}(\mathbf{r}(t), \mathbf{d})\, dt, \qquad T(t) = \exp!\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(s))\, ds\right) $$

$T(t)$ is the transmittance: the probability that the ray travels from $t_n$ to $t$ without being absorbed.

The quadrature used in practice, and implemented in the code above, is:

$$ \hat{C}(\mathbf{r}) = \sum_{i=1}^{N} T_i \bigl(1 - e^{-\sigma_i \delta_i}\bigr) \mathbf{c}_i, \qquad T_i = \exp!\left(-\sum_{j<i}\sigma_j \delta_j\right) $$

with $\delta_i = t_{i+1} - t_i$. This is exactly alpha compositing with $\alpha_i = 1 - e^{-\sigma_i \delta_i}$, and it is differentiable in every term — which is the whole reason the scheme is trainable.

The weights $w_i = T_i \alpha_i$ form a distribution over "where along the ray the ray terminated", normalised only when the ray is fully opaque. Expected depth is $\sum_i w_i t_i$, and the weight entropy is a useful measure of how confident the model is about a surface.

Positional encoding

Tancik et al. (NeurIPS 2020), Fourier features let networks learn high frequency functions in low dimensional domains, gives the theory the empirical result above demonstrates.

An MLP with ReLU activations has a neural tangent kernel that, on low-dimensional inputs, decays rapidly with frequency. The network is therefore heavily biased toward low-frequency functions — spectral bias — and fitting fine detail requires impractically long training.

Mapping the input through

$$ \gamma(p) = \bigl[\sin(2^0\pi p), \cos(2^0\pi p), \dots, \sin(2^{L-1}\pi p), \cos(2^{L-1}\pi p)\bigr] $$

makes the composed kernel stationary and widens its spectrum, so high frequencies become learnable. Tancik et al. show that a random Gaussian Fourier mapping with a tuned bandwidth generalises this, and that the bandwidth — not the number of features — is the parameter that matters.

The 8-band regression in the output is the predicted over-bandwidth failure: frequencies above the signal's own content add variance without signal.

NeRF uses $L = 10$ for position and $L = 4$ for direction.

Hierarchical sampling

Uniform sampling wastes almost every sample in empty space. NeRF trains two networks: a coarse one sampled uniformly, then a fine one sampled from the piecewise-constant PDF given by the coarse network's weights $w_i$, via inverse transform sampling.

This is importance sampling of the rendering integral, and it typically improves quality more than doubling the sample count would. Its descendants — proposal networks in Mip-NeRF 360 and nerfacto — do the same job with a small dedicated density network.

The efficiency lineage

The 122-million-evaluation figure explains the whole research programme.

  • Instant-NGP (Müller et al., SIGGRAPH 2022) replaces the frequency encoding with a multi-resolution hash grid of learnable feature vectors. Most of the capacity moves from network weights into the grid, so the MLP shrinks to two small layers. Training drops from days to seconds, with a fused CUDA implementation.
  • Plenoxels (Fridovich-Keil et al., CVPR 2022) removes the network entirely: a sparse voxel grid of spherical-harmonic coefficients, optimised directly. It establishes that the neural network was never the essential ingredient — differentiable volume rendering was.
  • TensoRF factorises the volume into low-rank tensor components, reducing memory sharply.

Quality, and the 360 problem

Mip-NeRF (Barron et al., ICCV 2021) replaces point samples with integrated positional encoding over a conical frustum, giving anti-aliasing and multi-scale consistency. Mip-NeRF 360 (CVPR 2022) adds a space contraction for unbounded scenes plus a distortion regulariser that suppresses the floating artefacts typical of outdoor captures. Zip-NeRF (ICCV 2023) combines that with hash grids.

Unbounded outdoor scenes remain the hard case, and the failure mode is characteristic: correct-looking training views plus "floaters", small clouds of density that satisfy the training views and appear from nowhere in new ones. They are an under-constrained-region symptom, not a bug.

Evaluation

PSNR, SSIM and LPIPS on held-out views, typically on Blender synthetic objects, LLFF forward-facing scenes, and Mip-NeRF 360 outdoor scenes.

Two cautions. PSNR is dominated by low frequencies and rewards blur, so a method can gain PSNR while looking worse; LPIPS (Zhang et al., CVPR 2018) tracks human judgement better and should be reported alongside. And held-out views in these datasets are interpolations between training views — genuine extrapolation away from the capture trajectory is much harder, and rarely measured.

Where NeRF stands after gaussian splatting

3D Gaussian splatting (2023) overtook NeRF on the metric that mattered commercially: real-time rendering. It is covered in the next lesson.

NeRF-style implicit fields remain preferable in several places, and it is worth being precise about which:

  • Surface extraction. Density fields convert to meshes via marching cubes; SDF-based variants (VolSDF, NeuS, Neuralangelo) give clean geometry that point-based representations struggle with.
  • Compactness. A hash-grid NeRF is often smaller on disk than the millions of Gaussians for the same scene.
  • Continuity. A continuous field can be queried anywhere and regularised analytically, which matters for physics, inverse rendering and relighting.

The deeper point is that both share the same rendering equation. Gaussian splatting replaces "query a network along a ray" with "sort and blend primitives", and the alpha-compositing arithmetic in the code above is unchanged.

What to learn next