Gaussian splatting
Gaussian splatting stores a scene as millions of soft coloured blobs and draws them by sorting and stacking, which is why it renders in real time where a NeRF cannot.
- 18 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Gaussian splatting builds a scene out of millions of soft coloured blobs, then stacks them front to back.
The analogy you have already lived
Think of an airbrush artist making a portrait. Every pass lays down one soft oval of colour — some wide, some narrow, some faint, some solid.
No single oval means anything. Enough of them, in the right places, and you have a face that looks photographic.
Gaussian splatting is that, in three dimensions, with the ovals found automatically from photographs.
Why it exists
A NeRF is beautiful and slow. Every pixel needs a hundred network calls, so a single frame is over a hundred million calls.
The insight was that the network was never the point. What made NeRF work was soft, see-through pieces blended together, trained by comparing renders against photos.
So keep the blending and throw away the network. Store the scene as explicit blobs, each holding its own position, size, orientation, colour and see-through-ness. Now drawing a frame is sorting and stacking, which graphics hardware has done at high speed for thirty years.
The result renders at over a hundred frames a second, and trains in minutes instead of days.
What one blob knows
Each blob is described by a handful of numbers.
Where it is in space. How big it is along three directions. It can be a sphere, a flat disc, or a thin sliver. Which way it is turned. What colour it is, and how that colour changes depending on where you look from. How see-through it is.
In the 2D version below, seven numbers describe one blob. A real 3D one uses 59, mostly for the way its colour shifts with viewing angle.
How a picture gets drawn
for the view you want:
1. work out where each blob lands on the screen, and how big it appears
2. sort them from nearest to furthest
3. go through in order, laying each one on top,
keeping track of how much light is still getting through
4. stop early once a pixel is fully coveredStep 2 is the whole engineering problem. Millions of blobs, sorted, every frame.
And step 2 genuinely matters. In the run below, blending two overlapping blobs in the wrong order gives 0.29 where the right order gives 0.74. Same blobs, different answer.
How the blobs are found
Start with the sparse point cloud from structure from motion, and put a small blob at each point.
Then repeat: render a training view, compare with the real photo, and adjust every blob a little to reduce the difference. Everything is differentiable, so ordinary training works.
The clever extra step is that the number of blobs changes during training. Where the picture is still wrong, blobs are cloned or split into two. Blobs that faded to nothing get deleted. The model grows to fit the scene.
The run below shows why that matters: 20 blobs produce a poor fit, 400 produce an excellent one, with nothing else changed.
The honest part
Gaussian splatting is not a 3D model in the useful sense. It is a cloud of blobs that happens to look right from the outside.
You cannot easily 3D print it, drop it into a physics engine, or relight it. Getting a clean surface out of it needs extra work, and that work is an active research problem.
It is also large. A room can be millions of blobs and hundreds of megabytes. A NeRF of the same room is often far smaller.
Where you have already seen it
- Property tours you can fly through smoothly on a phone.
- Product captures on shopping sites that rotate realistically.
- Film previsualisation with a captured real location.
- Sports and concert replays from many cameras at once.
Remember this
- The scene is millions of soft blobs, not a network and not a mesh.
- Drawing means sorting by depth and stacking, which is fast.
- More blobs where the picture is wrong; that growth is part of training.
What to learn next
- Neural radiance fields — the same rendering equation, solved the other way.
- Structure from motion — the step every splatting capture starts with.
- Diffusion models — the generative side now feeding 3D scene creation.
Developer — Code and libraries.
Setup
pip install torch==2.5.1 numpy==1.26.4The real thing needs a CUDA rasteriser and a GPU. Its two central ideas run on a laptop CPU, and both are worth touching directly: what one Gaussian is, and what fitting a scene out of them looks like. This file takes about fifty seconds.
Gaussians, blending order, and a fit
import numpy as np
import torch
print("PART 1 - one Gaussian, and why blending order matters\n")
def splat(cx, cy, sx, sy, theta=0.0, size=9):
"""The 2-D footprint of one Gaussian on a small pixel grid."""
y, x = np.mgrid[0:size, 0:size]
c, s = np.cos(theta), np.sin(theta)
dx, dy = x - cx, y - cy
u, v = c * dx + s * dy, -s * dx + c * dy # rotate into the Gaussian's own axes
return np.exp(-0.5 * ((u / sx) ** 2 + (v / sy) ** 2))
g = splat(4, 4, 2.2, 0.9, np.radians(30))
print(" a stretched, rotated Gaussian (values x100):")
for row in g:
print(" " + " ".join(f"{int(round(v * 100)):3d}" for v in row))
print("\n three numbers set the shape: two widths and one angle")
print(" in 3-D that becomes three scales and a quaternion\n")
# Two overlapping splats, one bright (0.9) in front, one dark (0.1) behind.
front = 0.85 * splat(3.5, 4, 1.6, 1.6)
back = 0.85 * splat(5.0, 4, 1.6, 1.6)
print(" two overlapping splats: a bright one (0.9) nearer, a dark one (0.1) behind")
print(f" {'blend order':<24}{'front only':<14}{'overlap':<12}{'back only'}")
for name, order in [("front first (correct)", [(front, 0.9), (back, 0.1)]),
("back first (wrong)", [(back, 0.1), (front, 0.9)])]:
colour, T = np.zeros_like(front), np.ones_like(front)
for a, c in order:
colour += T * a * c
T *= (1 - a)
print(f" {name:<24}{colour[4, 2]:<14.4f}{colour[4, 4]:<12.4f}{colour[4, 7]:.4f}")
print(" the two agree away from the overlap and disagree inside it")
print(" so a splatting renderer must SORT by depth")
print(" sorting millions of primitives every frame is the engineering problem\n")
print("PART 2 - fitting a picture with Gaussians\n")
torch.manual_seed(0)
H = W = 64
yy, xx = torch.meshgrid(torch.linspace(0, H - 1, H), torch.linspace(0, W - 1, W),
indexing="ij")
ty, tx = yy / H * 2 - 1, xx / W * 2 - 1
target = torch.clamp(0.5 + 0.45 * torch.sin(5 * tx) * torch.cos(4 * ty)
+ 0.25 * torch.exp(-8 * ((tx - 0.4) ** 2 + (ty + 0.3) ** 2)), 0, 1)
def build(n, seed=0):
gen = torch.Generator().manual_seed(seed)
return {
"pos": torch.rand(n, 2, generator=gen).mul(torch.tensor([W, H])).requires_grad_(),
"logs": torch.full((n, 2), 1.2).requires_grad_(), # log of the two widths
"rot": torch.zeros(n, requires_grad=True),
"col": torch.full((n,), 0.5).requires_grad_(),
"op": torch.zeros(n, requires_grad=True), # logit of opacity
}
def render(p):
sx, sy = torch.exp(p["logs"]).clamp(0.6, 30).T
c, s = torch.cos(p["rot"]), torch.sin(p["rot"])
dx = xx[None] - p["pos"][:, 0, None, None]
dy = yy[None] - p["pos"][:, 1, None, None]
u = c[:, None, None] * dx + s[:, None, None] * dy
v = -s[:, None, None] * dx + c[:, None, None] * dy
g = torch.exp(-0.5 * ((u / sx[:, None, None]) ** 2 + (v / sy[:, None, None]) ** 2))
alpha = torch.sigmoid(p["op"])[:, None, None] * g
T = torch.cumprod(torch.cat([torch.ones(1, H, W), 1 - alpha[:-1]]), dim=0)
weights = T * alpha # front-to-back compositing
return (weights * torch.sigmoid(p["col"])[:, None, None]).sum(0) + T[-1] * (1 - alpha[-1]) * 0.5
def fit(n, steps=700):
p = build(n)
opt = torch.optim.Adam(p.values(), lr=0.08)
for _ in range(steps):
loss = ((render(p) - target) ** 2).mean()
opt.zero_grad()
loss.backward()
opt.step()
with torch.no_grad():
mse = ((render(p) - target) ** 2).mean().item()
return mse, p
print(f" {'gaussians':<12}{'parameters':<14}{'final MSE':<14}{'PSNR (dB)'}")
results = {}
for n in [20, 60, 150, 400]:
mse, p = fit(n)
results[n] = (mse, p)
print(f" {n:<12}{n * 7:<14}{mse:<14.6f}{-10 * np.log10(mse):.2f}")
print("\n every Gaussian carries 7 numbers here: 2 position, 2 scale, 1 angle,")
print(" 1 colour, 1 opacity. A 3-D one carries 59, mostly colour coefficients.\n")
mse, p = results[400]
with torch.no_grad():
out = render(p).numpy()
op = torch.sigmoid(p["op"]).numpy()
scales = torch.exp(p["logs"]).clamp(0.6, 30).numpy()
print(f" of 400 fitted gaussians, opacity ranges {op.min():.3f} to {op.max():.3f}")
print(f" widths range from {scales.min():.2f} to {scales.max():.2f} pixels")
print(f" {int((scales.max(1) > 20).sum())} grew to the width cap, "
f"{int((scales.min(1) < 0.7).sum())} shrank to the floor")
print(" real implementations delete the faint ones and split the oversized ones\n")
chars = " .:-=+*#%@"
print(" target fitted with 400 gaussians")
t = target.numpy()
for r in range(0, H, 4):
a = "".join(chars[min(int(v * 9.99), 9)] for v in t[r, ::2])
b = "".join(chars[min(int(np.clip(v, 0, 1) * 9.99), 9)] for v in out[r, ::2])
print(f" {a} {b}")PART 1 - one Gaussian, and why blending order matters
a stretched, rotated Gaussian (values x100):
1 1 1 0 0 0 0 0 0
6 8 7 4 1 0 0 0 0
12 25 33 27 14 5 1 0 0
9 29 59 76 61 31 10 2 0
2 12 40 79 100 79 40 12 2
0 2 10 31 61 76 59 29 9
0 0 1 5 14 27 33 25 12
0 0 0 0 1 4 7 8 6
0 0 0 0 0 0 1 1 1
three numbers set the shape: two widths and one angle
in 3-D that becomes three scales and a quaternion
two overlapping splats: a bright one (0.9) nearer, a dark one (0.1) behind
blend order front only overlap back only
front first (correct) 0.4996 0.7419 0.1058
back first (wrong) 0.4354 0.2891 0.0816
the two agree away from the overlap and disagree inside it
so a splatting renderer must SORT by depth
sorting millions of primitives every frame is the engineering problem
PART 2 - fitting a picture with Gaussians
gaussians parameters final MSE PSNR (dB)
20 140 0.005136 22.89
60 420 0.000024 46.26
150 1050 0.000013 48.92
400 2800 0.000004 54.26
every Gaussian carries 7 numbers here: 2 position, 2 scale, 1 angle,
1 colour, 1 opacity. A 3-D one carries 59, mostly colour coefficients.
of 400 fitted gaussians, opacity ranges 0.051 to 0.994
widths range from 0.60 to 30.00 pixels
15 grew to the width cap, 54 shrank to the floor
real implementations delete the faint ones and split the oversized ones
target fitted with 400 gaussians
::::-=++*#####*++=-:::::-==+*### ::::-=++*#####*+==-:::::-==+*###
..:-+*#%@@%%#*+-:.. ..:-+*#%%@ .:-+*#%@@@%#*+-:.. ..:-+*#%@@
.:-+*#%@@@%#*+-:.....:-+*#%@@ .:-+*#%@@@%#*+-:.....:-+*#%@@
...::-+*##%%%##*+=-::::--=+*#%%% ...::-+*##%%%##*+=-::::--=+*#%%%
-----=++*******+++====++++****## -----=++*******+++====++++****##
++++++========+++***######**+++= ++++++========+++***######**+++=
###**+==--:::--=+*#%@@@@%#*+=--: ###**+==-::::--=+*#%@@@@%#*+=--:
%%%%#*=-:.....:=+*%@@@@@@#*=-:.. %%%%#*=-:.....:=+*%@@@@@@#*=-:..
@@@%#*=-:. .:-+*%@@@@@%#+=:. @@@%#*=-:. .:-+*%@@@@@%#+=:.
%%%%#*=-:.....:-+*#%@@@%#*+=::.. %%%%#*=-:.....:-+*#%@@@%#*+=::..
###**+==--:::--=++*#####*++=--:: ###**+==--:::--=++*#####*++=--::
++++++==========+++++++++++===== ++++++==========+++++++++++=====
-----=++*******+==-------==+**** -----=++*******++=-------==+****
...::-+*##%%%##*=-::...::-=*##%% ...::-+*##%%%##*=-::...::-=*##%%
.:-+*#%@@@%#*=-:. .:-=*#%@@ .:-+*#%@@@%#*=-:. .:-=*#%@@
..:-+*#%@@%%#*=-:. .:-=*#%%@ ..:-+*#%@@%%#*=-:. .:-=*#%%@Reading the output carefully
The printed Gaussian is anisotropic. It is wide along one axis and narrow along the other, and rotated 30 degrees. That is the whole reason Gaussians beat spheres for this job: a flat surface is well represented by a thin, wide, correctly-oriented disc, and badly represented by a stack of spheres.
0.7419 against 0.2891 in the overlap. The two blend orders agree in the regions where only one splat contributes and disagree sharply where both do. Alpha compositing is not commutative, which is why the real renderer sorts by depth every frame and why a fast sort is the core engineering achievement of the 2023 paper.
22.89 dB at 20 Gaussians, 46.26 dB at 60. A jump of 23 dB from tripling the primitive count. This is the argument for adaptive density control: rather than guessing a count, start small and add capacity wherever the render still disagrees with the photo.
Diminishing returns after that: 46.26, 48.92, 54.26. Twenty times the parameters between the first and last row for another 8 dB. Real scenes sit somewhere on this curve, and the practical question is always where to stop, since every Gaussian costs memory and sort time.
54 shrank to the floor and 15 grew to the cap. Both are failure signals a real implementation acts on. A Gaussian that collapses to sub-pixel size contributes almost nothing and would be pruned. One that grows enormous is a symptom of an under-reconstructed region, and the 3DGS paper's response is to split it into two smaller ones.
Opacity floor 0.051. Nothing went fully transparent because our toy has no pruning step. In the real algorithm, opacity is periodically reset toward zero and Gaussians that fail to recover are deleted — a deliberate mechanism for removing floaters.
What a real 3D Gaussian holds
| Property | Numbers | Note |
|---|---|---|
| Position | 3 | the mean |
| Scale | 3 | stored as logs, so it stays positive |
| Rotation | 4 | a quaternion, normalised each step |
| Opacity | 1 | stored as a logit |
| Colour | 48 | spherical harmonics to degree 3, three channels |
Fifty-nine numbers each, and a typical scene holds one to five million of them. That is why the files are large.
The 48 colour numbers are what makes a surface look different from different angles — reflections on a table, a highlight on a metal handle. Degree 0 alone would be a flat colour that looks the same everywhere.
Running one for real
pip install nerfstudio
ns-process-data images --data ./photos --output-dir ./scene # runs COLMAP
ns-train splatfacto --data ./scenesplatfacto is nerfstudio's Gaussian splatting method, built on the gsplat rasteriser. The original reference implementation from Inria is also available, and gsplat is the more actively maintained library.
The COLMAP step is not optional and is not a formality. Splatting is initialised from the SfM point cloud, so a poor reconstruction gives a poor start, and the capture advice in structure from motion applies unchanged.
Common mistakes
Expecting a mesh. The output is a point cloud with covariances. Extracting a surface needs 2D Gaussian splatting, SuGaR, or a similar method, and none of them is as clean as marching cubes on a density field.
Ignoring the size of the result. Millions of Gaussians at 59 floats each runs to hundreds of megabytes. Compression methods exist — vector quantisation, pruning, SH truncation — and are worth applying before shipping to a browser.
Training on views that all look the same way. Under-observed regions fill with floaters that satisfy the training views and appear from nowhere elsewhere. This is the same failure as with NeRF and the same fix: surround the subject.
Assuming it can be relit. Colour and lighting are baked together. Changing the lighting requires inverse rendering, which is a separate and harder problem.
Comparing against NeRF on PSNR alone. The interesting comparison is quality at a given render time, and on that axis splatting wins by two orders of magnitude.
Try it yourself
Set steps=200 in fit and rerun the sweep. Every row gets worse, and the 400-Gaussian row loses more than the 20-Gaussian row. More primitives need more optimisation to place, which is why real training schedules add Gaussians gradually rather than all at once.
What to learn next
- Neural radiance fields — the same rendering equation, solved the other way.
- Structure from motion — the step every splatting capture starts with.
- Diffusion models — the generative side now feeding 3D scene creation.
Researcher — Mathematics and papers.
The representation
Kerbl et al. (SIGGRAPH 2023), 3D Gaussian splatting for real-time radiance field rendering, ACM TOG 42(4). A scene is a set of anisotropic 3D Gaussians
$$ G_i(\mathbf{x}) = \exp!\left(-\tfrac{1}{2}(\mathbf{x} - \boldsymbol{\mu}_i)^\top \Sigma_i^{-1} (\mathbf{x} - \boldsymbol{\mu}_i)\right) $$
Optimising $\Sigma$ directly would require keeping it positive semi-definite, so it is parameterised as
$$ \Sigma = R S S^\top R^\top $$
with $S$ a diagonal scale matrix and $R$ a rotation from a normalised quaternion. Any gradient step then produces a valid covariance by construction. The parameters stored are the scale logs and the quaternion, not $\Sigma$.
Projection
A 3D Gaussian projects to a 2D Gaussian on the image plane. The projected covariance is, following Zwicker et al. (2001), EWA volume splatting:
$$ \Sigma' = J W \Sigma W^\top J^\top $$
Where $W$ is the world-to-camera rotation and $J$ the Jacobian of the perspective projection at the Gaussian's centre. $J$ is a local affine approximation to a nonlinear map, so the projection is exact only for small Gaussians near the optical axis. Large Gaussians toward the image edges are drawn slightly wrong, and this is a real, known artefact rather than an implementation detail.
The 2D covariance is then dilated slightly, adding a small multiple of the identity, to guarantee a minimum screen-space footprint of roughly one pixel. Without it, sub-pixel Gaussians alias badly.
Rendering
Pixel colour is exactly the alpha compositing in the code above:
$$ C = \sum_{i \in \mathcal{N}} c_i \alpha_i \prod_{j=1}^{i-1}(1 - \alpha_j), \qquad \alpha_i = o_i \, G'_i(\mathbf{p}) $$
with $\mathcal{N}$ the depth-sorted Gaussians overlapping the pixel, $o_i$ the learned opacity, and $G'_i$ the projected 2D Gaussian. This is the same equation as the NeRF quadrature. The difference is entirely in how the samples along the ray are obtained: queried from a network, or read from sorted explicit primitives.
The rasteriser is what makes it fast:
- Split the screen into 16x16 tiles. Assign each Gaussian to every tile its 99% confidence ellipse touches.
- Build (tile, depth) keys and radix sort globally once per frame, giving a per-tile depth order without per-pixel sorting.
- Each tile is one CUDA block; threads cooperatively load Gaussians into shared memory and blend front to back, terminating a pixel once transmittance falls below a threshold.
- The backward pass traverses back to front, recovering the required transmittance by division rather than storing per-pixel lists.
The global sort is an approximation: Gaussians are sorted by centre depth, not per-pixel, so two intersecting Gaussians can be blended in the wrong order for some pixels. It is visually acceptable and is the source of some popping artefacts under camera motion.
Adaptive density control
Initialised from the SfM point cloud, then every 100 iterations:
- Densify Gaussians whose average positional gradient magnitude exceeds a threshold — a signal that the region is not yet well reconstructed. Small ones are cloned and shifted along the gradient; large ones are split into two with scale divided by roughly 1.6.
- Prune Gaussians with opacity below a threshold, and those with excessive screen-space extent.
- Reset opacity toward zero periodically, forcing everything to re-earn its place. This is the main mechanism against floaters near the cameras.
This is the piece that lets the model choose its own capacity, and it has no equivalent in a fixed-size MLP.
Trade-offs against NeRF
| 3DGS | NeRF (hash-grid) | |
|---|---|---|
| Render speed | 100+ FPS at 1080p | Interactive to slow |
| Train time | Minutes to tens of minutes | Minutes to hours |
| Storage | Hundreds of MB typical | Tens of MB typical |
| Surface extraction | Awkward; needs variants | Marching cubes on density |
| Editing | Direct, primitives are explicit | Indirect |
| Unbounded scenes | Good | Good with contraction |
| Anti-aliasing across scales | Needs Mip-Splatting | Native with Mip-NeRF |
Notable follow-ups
- Mip-Splatting (Yu et al., CVPR 2024) adds a 3D smoothing filter bounding each Gaussian's maximum frequency by the training sampling rate, plus a 2D Mip filter replacing the dilation. It fixes the erosion and dilation artefacts that appear when rendering at a different resolution or distance from training.
- 2D Gaussian splatting (Huang et al., SIGGRAPH 2024) replaces 3D ellipsoids with oriented planar discs, giving view-consistent geometry and considerably better surface extraction. The right choice when you need a mesh.
- Deformable and 4D variants extend the representation to dynamic scenes by adding a time-conditioned deformation field or time-varying primitives.
- gsplat (Ye et al., 2024, arXiv 2409.06765) is the open reference rasteriser, used by nerfstudio's Splatfacto, with continuing work through 2026 on faster rasterisation paths and broader camera and sensor models including fisheye and lidar.
- Compression. Vector quantisation of SH coefficients, opacity-based pruning and entropy coding routinely cut file sizes by an order of magnitude with small quality loss.
The honest limits
It is not geometry. Photometric loss constrains what the scene looks like, not what it is. Gaussians drift off surfaces, become large and semi-transparent in ambiguous regions, and adopt shapes that render correctly and mean nothing. Any downstream use assuming a surface — collision, physics, printing, measurement — needs a variant designed for it.
Baked appearance. Lighting is inseparable from albedo. Relighting requires inverse rendering and remains open.
The sorting approximation. Per-Gaussian rather than per-fragment depth ordering causes popping. Exact solutions cost the performance that motivated the method.
Benchmarks are narrow. Mip-NeRF 360, Tanks and Temples, and Deep Blending are small, static, well-lit and densely captured. Sparse views, motion, changing lighting and reflective surfaces are all substantially harder, and headline PSNR figures say little about them.
What to learn next
- Neural radiance fields — the same rendering equation, solved the other way.
- Structure from motion — the step every splatting capture starts with.
- Diffusion models — the generative side now feeding 3D scene creation.