3D Vision and Depth

Gaussian splatting

Gaussian splatting stores a scene as millions of soft coloured blobs and draws them by sorting and stacking, which is why it renders in real time where a NeRF cannot.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. What one blob knows
  5. How a picture gets drawn
  6. How the blobs are found
  7. The honest part
  8. Where you have already seen it
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Gaussian splatting builds a scene out of millions of soft coloured blobs, then stacks them front to back.

The analogy you have already lived

Think of an airbrush artist making a portrait. Every pass lays down one soft oval of colour — some wide, some narrow, some faint, some solid.

No single oval means anything. Enough of them, in the right places, and you have a face that looks photographic.

Gaussian splatting is that, in three dimensions, with the ovals found automatically from photographs.

Why it exists

A NeRF is beautiful and slow. Every pixel needs a hundred network calls, so a single frame is over a hundred million calls.

The insight was that the network was never the point. What made NeRF work was soft, see-through pieces blended together, trained by comparing renders against photos.

So keep the blending and throw away the network. Store the scene as explicit blobs, each holding its own position, size, orientation, colour and see-through-ness. Now drawing a frame is sorting and stacking, which graphics hardware has done at high speed for thirty years.

The result renders at over a hundred frames a second, and trains in minutes instead of days.

What one blob knows

Each blob is described by a handful of numbers.

Where it is in space. How big it is along three directions. It can be a sphere, a flat disc, or a thin sliver. Which way it is turned. What colour it is, and how that colour changes depending on where you look from. How see-through it is.

In the 2D version below, seven numbers describe one blob. A real 3D one uses 59, mostly for the way its colour shifts with viewing angle.

How a picture gets drawn

   for the view you want:
      1. work out where each blob lands on the screen, and how big it appears
      2. sort them from nearest to furthest
      3. go through in order, laying each one on top,
         keeping track of how much light is still getting through
      4. stop early once a pixel is fully covered

Step 2 is the whole engineering problem. Millions of blobs, sorted, every frame.

And step 2 genuinely matters. In the run below, blending two overlapping blobs in the wrong order gives 0.29 where the right order gives 0.74. Same blobs, different answer.

How the blobs are found

Start with the sparse point cloud from structure from motion, and put a small blob at each point.

Then repeat: render a training view, compare with the real photo, and adjust every blob a little to reduce the difference. Everything is differentiable, so ordinary training works.

The clever extra step is that the number of blobs changes during training. Where the picture is still wrong, blobs are cloned or split into two. Blobs that faded to nothing get deleted. The model grows to fit the scene.

The run below shows why that matters: 20 blobs produce a poor fit, 400 produce an excellent one, with nothing else changed.

The honest part

Gaussian splatting is not a 3D model in the useful sense. It is a cloud of blobs that happens to look right from the outside.

You cannot easily 3D print it, drop it into a physics engine, or relight it. Getting a clean surface out of it needs extra work, and that work is an active research problem.

It is also large. A room can be millions of blobs and hundreds of megabytes. A NeRF of the same room is often far smaller.

Where you have already seen it

  • Property tours you can fly through smoothly on a phone.
  • Product captures on shopping sites that rotate realistically.
  • Film previsualisation with a captured real location.
  • Sports and concert replays from many cameras at once.

Remember this

  • The scene is millions of soft blobs, not a network and not a mesh.
  • Drawing means sorting by depth and stacking, which is fast.
  • More blobs where the picture is wrong; that growth is part of training.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch==2.5.1 numpy==1.26.4

The real thing needs a CUDA rasteriser and a GPU. Its two central ideas run on a laptop CPU, and both are worth touching directly: what one Gaussian is, and what fitting a scene out of them looks like. This file takes about fifty seconds.

Gaussians, blending order, and a fit

splatting.py
import numpy as np
import torch

print("PART 1 - one Gaussian, and why blending order matters\n")


def splat(cx, cy, sx, sy, theta=0.0, size=9):
    """The 2-D footprint of one Gaussian on a small pixel grid."""
    y, x = np.mgrid[0:size, 0:size]
    c, s = np.cos(theta), np.sin(theta)
    dx, dy = x - cx, y - cy
    u, v = c * dx + s * dy, -s * dx + c * dy      # rotate into the Gaussian's own axes
    return np.exp(-0.5 * ((u / sx) ** 2 + (v / sy) ** 2))


g = splat(4, 4, 2.2, 0.9, np.radians(30))
print("   a stretched, rotated Gaussian (values x100):")
for row in g:
    print("   " + " ".join(f"{int(round(v * 100)):3d}" for v in row))
print("\n   three numbers set the shape: two widths and one angle")
print("   in 3-D that becomes three scales and a quaternion\n")

# Two overlapping splats, one bright (0.9) in front, one dark (0.1) behind.
front = 0.85 * splat(3.5, 4, 1.6, 1.6)
back = 0.85 * splat(5.0, 4, 1.6, 1.6)
print("   two overlapping splats: a bright one (0.9) nearer, a dark one (0.1) behind")
print(f"   {'blend order':<24}{'front only':<14}{'overlap':<12}{'back only'}")
for name, order in [("front first (correct)", [(front, 0.9), (back, 0.1)]),
                    ("back first  (wrong)", [(back, 0.1), (front, 0.9)])]:
    colour, T = np.zeros_like(front), np.ones_like(front)
    for a, c in order:
        colour += T * a * c
        T *= (1 - a)
    print(f"   {name:<24}{colour[4, 2]:<14.4f}{colour[4, 4]:<12.4f}{colour[4, 7]:.4f}")
print("   the two agree away from the overlap and disagree inside it")
print("   so a splatting renderer must SORT by depth")
print("   sorting millions of primitives every frame is the engineering problem\n")

print("PART 2 - fitting a picture with Gaussians\n")
torch.manual_seed(0)
H = W = 64
yy, xx = torch.meshgrid(torch.linspace(0, H - 1, H), torch.linspace(0, W - 1, W),
                        indexing="ij")
ty, tx = yy / H * 2 - 1, xx / W * 2 - 1
target = torch.clamp(0.5 + 0.45 * torch.sin(5 * tx) * torch.cos(4 * ty)
                     + 0.25 * torch.exp(-8 * ((tx - 0.4) ** 2 + (ty + 0.3) ** 2)), 0, 1)


def build(n, seed=0):
    gen = torch.Generator().manual_seed(seed)
    return {
        "pos": torch.rand(n, 2, generator=gen).mul(torch.tensor([W, H])).requires_grad_(),
        "logs": torch.full((n, 2), 1.2).requires_grad_(),      # log of the two widths
        "rot": torch.zeros(n, requires_grad=True),
        "col": torch.full((n,), 0.5).requires_grad_(),
        "op": torch.zeros(n, requires_grad=True),              # logit of opacity
    }


def render(p):
    sx, sy = torch.exp(p["logs"]).clamp(0.6, 30).T
    c, s = torch.cos(p["rot"]), torch.sin(p["rot"])
    dx = xx[None] - p["pos"][:, 0, None, None]
    dy = yy[None] - p["pos"][:, 1, None, None]
    u = c[:, None, None] * dx + s[:, None, None] * dy
    v = -s[:, None, None] * dx + c[:, None, None] * dy
    g = torch.exp(-0.5 * ((u / sx[:, None, None]) ** 2 + (v / sy[:, None, None]) ** 2))
    alpha = torch.sigmoid(p["op"])[:, None, None] * g
    T = torch.cumprod(torch.cat([torch.ones(1, H, W), 1 - alpha[:-1]]), dim=0)
    weights = T * alpha                                        # front-to-back compositing
    return (weights * torch.sigmoid(p["col"])[:, None, None]).sum(0) + T[-1] * (1 - alpha[-1]) * 0.5


def fit(n, steps=700):
    p = build(n)
    opt = torch.optim.Adam(p.values(), lr=0.08)
    for _ in range(steps):
        loss = ((render(p) - target) ** 2).mean()
        opt.zero_grad()
        loss.backward()
        opt.step()
    with torch.no_grad():
        mse = ((render(p) - target) ** 2).mean().item()
    return mse, p


print(f"   {'gaussians':<12}{'parameters':<14}{'final MSE':<14}{'PSNR (dB)'}")
results = {}
for n in [20, 60, 150, 400]:
    mse, p = fit(n)
    results[n] = (mse, p)
    print(f"   {n:<12}{n * 7:<14}{mse:<14.6f}{-10 * np.log10(mse):.2f}")
print("\n   every Gaussian carries 7 numbers here: 2 position, 2 scale, 1 angle,")
print("   1 colour, 1 opacity. A 3-D one carries 59, mostly colour coefficients.\n")

mse, p = results[400]
with torch.no_grad():
    out = render(p).numpy()
    op = torch.sigmoid(p["op"]).numpy()
    scales = torch.exp(p["logs"]).clamp(0.6, 30).numpy()
print(f"   of 400 fitted gaussians, opacity ranges {op.min():.3f} to {op.max():.3f}")
print(f"   widths range from {scales.min():.2f} to {scales.max():.2f} pixels")
print(f"   {int((scales.max(1) > 20).sum())} grew to the width cap, "
      f"{int((scales.min(1) < 0.7).sum())} shrank to the floor")
print("   real implementations delete the faint ones and split the oversized ones\n")

chars = " .:-=+*#%@"
print("   target                             fitted with 400 gaussians")
t = target.numpy()
for r in range(0, H, 4):
    a = "".join(chars[min(int(v * 9.99), 9)] for v in t[r, ::2])
    b = "".join(chars[min(int(np.clip(v, 0, 1) * 9.99), 9)] for v in out[r, ::2])
    print(f"   {a}   {b}")
Output
PART 1 - one Gaussian, and why blending order matters

   a stretched, rotated Gaussian (values x100):
     1   1   1   0   0   0   0   0   0
     6   8   7   4   1   0   0   0   0
    12  25  33  27  14   5   1   0   0
     9  29  59  76  61  31  10   2   0
     2  12  40  79 100  79  40  12   2
     0   2  10  31  61  76  59  29   9
     0   0   1   5  14  27  33  25  12
     0   0   0   0   1   4   7   8   6
     0   0   0   0   0   0   1   1   1

   three numbers set the shape: two widths and one angle
   in 3-D that becomes three scales and a quaternion

   two overlapping splats: a bright one (0.9) nearer, a dark one (0.1) behind
   blend order             front only    overlap     back only
   front first (correct)   0.4996        0.7419      0.1058
   back first  (wrong)     0.4354        0.2891      0.0816
   the two agree away from the overlap and disagree inside it
   so a splatting renderer must SORT by depth
   sorting millions of primitives every frame is the engineering problem

PART 2 - fitting a picture with Gaussians

   gaussians   parameters    final MSE     PSNR (dB)
   20          140           0.005136      22.89
   60          420           0.000024      46.26
   150         1050          0.000013      48.92
   400         2800          0.000004      54.26

   every Gaussian carries 7 numbers here: 2 position, 2 scale, 1 angle,
   1 colour, 1 opacity. A 3-D one carries 59, mostly colour coefficients.

   of 400 fitted gaussians, opacity ranges 0.051 to 0.994
   widths range from 0.60 to 30.00 pixels
   15 grew to the width cap, 54 shrank to the floor
   real implementations delete the faint ones and split the oversized ones

   target                             fitted with 400 gaussians
   ::::-=++*#####*++=-:::::-==+*###   ::::-=++*#####*+==-:::::-==+*###
     ..:-+*#%@@%%#*+-:.. ..:-+*#%%@      .:-+*#%@@@%#*+-:.. ..:-+*#%@@
      .:-+*#%@@@%#*+-:.....:-+*#%@@      .:-+*#%@@@%#*+-:.....:-+*#%@@
   ...::-+*##%%%##*+=-::::--=+*#%%%   ...::-+*##%%%##*+=-::::--=+*#%%%
   -----=++*******+++====++++****##   -----=++*******+++====++++****##
   ++++++========+++***######**+++=   ++++++========+++***######**+++=
   ###**+==--:::--=+*#%@@@@%#*+=--:   ###**+==-::::--=+*#%@@@@%#*+=--:
   %%%%#*=-:.....:=+*%@@@@@@#*=-:..   %%%%#*=-:.....:=+*%@@@@@@#*=-:..
   @@@%#*=-:.   .:-+*%@@@@@%#+=:.     @@@%#*=-:.   .:-+*%@@@@@%#+=:.  
   %%%%#*=-:.....:-+*#%@@@%#*+=::..   %%%%#*=-:.....:-+*#%@@@%#*+=::..
   ###**+==--:::--=++*#####*++=--::   ###**+==--:::--=++*#####*++=--::
   ++++++==========+++++++++++=====   ++++++==========+++++++++++=====
   -----=++*******+==-------==+****   -----=++*******++=-------==+****
   ...::-+*##%%%##*=-::...::-=*##%%   ...::-+*##%%%##*=-::...::-=*##%%
      .:-+*#%@@@%#*=-:.   .:-=*#%@@      .:-+*#%@@@%#*=-:.   .:-=*#%@@
     ..:-+*#%@@%%#*=-:.   .:-=*#%%@     ..:-+*#%@@%%#*=-:.   .:-=*#%%@

Reading the output carefully

The printed Gaussian is anisotropic. It is wide along one axis and narrow along the other, and rotated 30 degrees. That is the whole reason Gaussians beat spheres for this job: a flat surface is well represented by a thin, wide, correctly-oriented disc, and badly represented by a stack of spheres.

0.7419 against 0.2891 in the overlap. The two blend orders agree in the regions where only one splat contributes and disagree sharply where both do. Alpha compositing is not commutative, which is why the real renderer sorts by depth every frame and why a fast sort is the core engineering achievement of the 2023 paper.

22.89 dB at 20 Gaussians, 46.26 dB at 60. A jump of 23 dB from tripling the primitive count. This is the argument for adaptive density control: rather than guessing a count, start small and add capacity wherever the render still disagrees with the photo.

Diminishing returns after that: 46.26, 48.92, 54.26. Twenty times the parameters between the first and last row for another 8 dB. Real scenes sit somewhere on this curve, and the practical question is always where to stop, since every Gaussian costs memory and sort time.

54 shrank to the floor and 15 grew to the cap. Both are failure signals a real implementation acts on. A Gaussian that collapses to sub-pixel size contributes almost nothing and would be pruned. One that grows enormous is a symptom of an under-reconstructed region, and the 3DGS paper's response is to split it into two smaller ones.

Opacity floor 0.051. Nothing went fully transparent because our toy has no pruning step. In the real algorithm, opacity is periodically reset toward zero and Gaussians that fail to recover are deleted — a deliberate mechanism for removing floaters.

What a real 3D Gaussian holds

PropertyNumbersNote
Position3the mean
Scale3stored as logs, so it stays positive
Rotation4a quaternion, normalised each step
Opacity1stored as a logit
Colour48spherical harmonics to degree 3, three channels

Fifty-nine numbers each, and a typical scene holds one to five million of them. That is why the files are large.

The 48 colour numbers are what makes a surface look different from different angles — reflections on a table, a highlight on a metal handle. Degree 0 alone would be a flat colour that looks the same everywhere.

Running one for real

bash
pip install nerfstudio
ns-process-data images --data ./photos --output-dir ./scene    # runs COLMAP
ns-train splatfacto --data ./scene

splatfacto is nerfstudio's Gaussian splatting method, built on the gsplat rasteriser. The original reference implementation from Inria is also available, and gsplat is the more actively maintained library.

The COLMAP step is not optional and is not a formality. Splatting is initialised from the SfM point cloud, so a poor reconstruction gives a poor start, and the capture advice in structure from motion applies unchanged.

Common mistakes

Expecting a mesh. The output is a point cloud with covariances. Extracting a surface needs 2D Gaussian splatting, SuGaR, or a similar method, and none of them is as clean as marching cubes on a density field.

Ignoring the size of the result. Millions of Gaussians at 59 floats each runs to hundreds of megabytes. Compression methods exist — vector quantisation, pruning, SH truncation — and are worth applying before shipping to a browser.

Training on views that all look the same way. Under-observed regions fill with floaters that satisfy the training views and appear from nowhere elsewhere. This is the same failure as with NeRF and the same fix: surround the subject.

Assuming it can be relit. Colour and lighting are baked together. Changing the lighting requires inverse rendering, which is a separate and harder problem.

Comparing against NeRF on PSNR alone. The interesting comparison is quality at a given render time, and on that axis splatting wins by two orders of magnitude.

Try it yourself

Set steps=200 in fit and rerun the sweep. Every row gets worse, and the 400-Gaussian row loses more than the 20-Gaussian row. More primitives need more optimisation to place, which is why real training schedules add Gaussians gradually rather than all at once.

What to learn next

Researcher — Mathematics and papers.

The representation

Kerbl et al. (SIGGRAPH 2023), 3D Gaussian splatting for real-time radiance field rendering, ACM TOG 42(4). A scene is a set of anisotropic 3D Gaussians

$$ G_i(\mathbf{x}) = \exp!\left(-\tfrac{1}{2}(\mathbf{x} - \boldsymbol{\mu}_i)^\top \Sigma_i^{-1} (\mathbf{x} - \boldsymbol{\mu}_i)\right) $$

Optimising $\Sigma$ directly would require keeping it positive semi-definite, so it is parameterised as

$$ \Sigma = R S S^\top R^\top $$

with $S$ a diagonal scale matrix and $R$ a rotation from a normalised quaternion. Any gradient step then produces a valid covariance by construction. The parameters stored are the scale logs and the quaternion, not $\Sigma$.

Projection

A 3D Gaussian projects to a 2D Gaussian on the image plane. The projected covariance is, following Zwicker et al. (2001), EWA volume splatting:

$$ \Sigma' = J W \Sigma W^\top J^\top $$

Where $W$ is the world-to-camera rotation and $J$ the Jacobian of the perspective projection at the Gaussian's centre. $J$ is a local affine approximation to a nonlinear map, so the projection is exact only for small Gaussians near the optical axis. Large Gaussians toward the image edges are drawn slightly wrong, and this is a real, known artefact rather than an implementation detail.

The 2D covariance is then dilated slightly, adding a small multiple of the identity, to guarantee a minimum screen-space footprint of roughly one pixel. Without it, sub-pixel Gaussians alias badly.

Rendering

Pixel colour is exactly the alpha compositing in the code above:

$$ C = \sum_{i \in \mathcal{N}} c_i \alpha_i \prod_{j=1}^{i-1}(1 - \alpha_j), \qquad \alpha_i = o_i \, G'_i(\mathbf{p}) $$

with $\mathcal{N}$ the depth-sorted Gaussians overlapping the pixel, $o_i$ the learned opacity, and $G'_i$ the projected 2D Gaussian. This is the same equation as the NeRF quadrature. The difference is entirely in how the samples along the ray are obtained: queried from a network, or read from sorted explicit primitives.

The rasteriser is what makes it fast:

  1. Split the screen into 16x16 tiles. Assign each Gaussian to every tile its 99% confidence ellipse touches.
  2. Build (tile, depth) keys and radix sort globally once per frame, giving a per-tile depth order without per-pixel sorting.
  3. Each tile is one CUDA block; threads cooperatively load Gaussians into shared memory and blend front to back, terminating a pixel once transmittance falls below a threshold.
  4. The backward pass traverses back to front, recovering the required transmittance by division rather than storing per-pixel lists.

The global sort is an approximation: Gaussians are sorted by centre depth, not per-pixel, so two intersecting Gaussians can be blended in the wrong order for some pixels. It is visually acceptable and is the source of some popping artefacts under camera motion.

Adaptive density control

Initialised from the SfM point cloud, then every 100 iterations:

  • Densify Gaussians whose average positional gradient magnitude exceeds a threshold — a signal that the region is not yet well reconstructed. Small ones are cloned and shifted along the gradient; large ones are split into two with scale divided by roughly 1.6.
  • Prune Gaussians with opacity below a threshold, and those with excessive screen-space extent.
  • Reset opacity toward zero periodically, forcing everything to re-earn its place. This is the main mechanism against floaters near the cameras.

This is the piece that lets the model choose its own capacity, and it has no equivalent in a fixed-size MLP.

Trade-offs against NeRF

3DGSNeRF (hash-grid)
Render speed100+ FPS at 1080pInteractive to slow
Train timeMinutes to tens of minutesMinutes to hours
StorageHundreds of MB typicalTens of MB typical
Surface extractionAwkward; needs variantsMarching cubes on density
EditingDirect, primitives are explicitIndirect
Unbounded scenesGoodGood with contraction
Anti-aliasing across scalesNeeds Mip-SplattingNative with Mip-NeRF

Notable follow-ups

  • Mip-Splatting (Yu et al., CVPR 2024) adds a 3D smoothing filter bounding each Gaussian's maximum frequency by the training sampling rate, plus a 2D Mip filter replacing the dilation. It fixes the erosion and dilation artefacts that appear when rendering at a different resolution or distance from training.
  • 2D Gaussian splatting (Huang et al., SIGGRAPH 2024) replaces 3D ellipsoids with oriented planar discs, giving view-consistent geometry and considerably better surface extraction. The right choice when you need a mesh.
  • Deformable and 4D variants extend the representation to dynamic scenes by adding a time-conditioned deformation field or time-varying primitives.
  • gsplat (Ye et al., 2024, arXiv 2409.06765) is the open reference rasteriser, used by nerfstudio's Splatfacto, with continuing work through 2026 on faster rasterisation paths and broader camera and sensor models including fisheye and lidar.
  • Compression. Vector quantisation of SH coefficients, opacity-based pruning and entropy coding routinely cut file sizes by an order of magnitude with small quality loss.

The honest limits

It is not geometry. Photometric loss constrains what the scene looks like, not what it is. Gaussians drift off surfaces, become large and semi-transparent in ambiguous regions, and adopt shapes that render correctly and mean nothing. Any downstream use assuming a surface — collision, physics, printing, measurement — needs a variant designed for it.

Baked appearance. Lighting is inseparable from albedo. Relighting requires inverse rendering and remains open.

The sorting approximation. Per-Gaussian rather than per-fragment depth ordering causes popping. Exact solutions cost the performance that motivated the method.

Benchmarks are narrow. Mip-NeRF 360, Tanks and Temples, and Deep Blending are small, static, well-lit and densely captured. Sparse views, motion, changing lighting and reflective surfaces are all substantially harder, and headline PSNR figures say little about them.

What to learn next