Medical Imaging AI

Segmenting 3D scans

A CT or MRI scan is a stack of slices, not one picture, so labelling it means outlining a shape through three dimensions at once.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Segmenting a 3D scan means outlining a shape through many stacked slices, not one flat picture.

Think about slicing a loaf of bread. Each slice looks a little different from the one before it. To describe the whole loaf, you need every slice, not one alone.

A CT or MRI scan is exactly that: a stack of thin slices through the body. A tumor does not live on one slice. It passes through many, changing shape as it goes.

Why it exists

An ordinary photo has height and width. A CT or MRI scan adds a third dimension — depth, one slice per position through the body.

A tumor, an organ, or a bleed is a 3D shape sitting inside that stack. Outlining it on a single slice tells you almost nothing about its true size. A doctor planning surgery needs the whole shape. How deep does it go? How does it touch nearby structures? How much volume does it occupy?

Marking every slice by hand takes real time from a trained radiologist. That cost is exactly why automatic 3D segmentation is worth building.

How it works

Slice 1     Slice 2     Slice 3     Slice 4     ...
  o           O           O           o
(small)    (bigger)    (bigger)    (shrinking)

          Stack every slice's outline together
                        |
                        v
              one 3D shape, through the body

A single 2D outline is a cross-section. Stack the outlines from every slice, and the true 3D shape appears.

Where you have already seen it

  • A CT scan viewer lets you scroll through slices one at a time. You are looking at a 3D volume, through one 2D window.
  • A weather map showing storm layers at different altitudes uses the same idea: many flat slices, describing one 3D shape.

An honest warning

A 3D scan is enormous compared to an ordinary photo. A single chest CT can be hundreds of times larger than a typical photo file.

That size means a model cannot look at the whole scan at once. It cannot treat it the way it treats a small photo. Real systems break the volume into smaller pieces first — a workaround covered in the developer section below.

A segmentation model's output still needs a radiologist's sign-off. Regulatory clearance matters too, before it guides a real surgery.

Remember this

  • A CT or MRI scan is a 3D volume, built from many 2D slices stacked together.
  • A shape's true size and depth only appear once every slice is considered together.
  • 3D scans are far larger than ordinary photos, which shapes how any model has to handle them.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

volumetric_dice.py
import numpy as np

rng = np.random.default_rng(0)

# A tiny stand-in for a CT volume: depth x height x width. A real chest CT
# is closer to 200 x 512 x 512 -- about 1,600 times bigger than this toy.
depth, height, width = 24, 48, 48
volume = rng.normal(0, 1, (depth, height, width))

# Embed a bright spherical "lesion" -- the 3D version of a blob on one slice.
zz, yy, xx = np.mgrid[0:depth, 0:height, 0:width]
center = (12, 24, 24)
distance = np.sqrt((zz - center[0])**2 + (yy - center[1])**2 + (xx - center[2])**2)
true_mask = distance <= 6
volume[true_mask] += 3.0  # the lesion is brighter than surrounding tissue

# The simplest possible 3D segmentation: threshold every voxel by intensity.
# ("Voxel" is a 3D pixel -- a small cube of volume instead of a flat square.)
predicted_mask = volume > 1.5

def dice_score(a, b):
    intersection = np.logical_and(a, b).sum()
    return 2 * intersection / (a.sum() + b.sum())

print(f"volume shape: {volume.shape}, total voxels: {volume.size:,}")
print(f"true lesion voxels: {true_mask.sum()}, predicted lesion voxels: {predicted_mask.sum()}")
print(f"3D Dice score: {dice_score(true_mask, predicted_mask):.3f}")

# What a full-size real volume would cost, at 4 bytes per voxel (float32).
real_shape = (200, 512, 512)
real_bytes = np.prod(real_shape) * 4
print(f"a real {real_shape} volume in float32 needs {real_bytes / 1e6:.0f} MB -- for ONE scan")
Output
volume shape: (24, 48, 48), total voxels: 55,296
true lesion voxels: 925, predicted lesion voxels: 4473
3D Dice score: 0.318
a real (200, 512, 512) volume in float32 needs 210 MB -- for ONE scan

What actually happened

A Dice score of 0.318 looks poor, on purpose. Global intensity thresholding — "call every bright voxel a lesion" — also catches ordinary random noise across the entire 55,296-voxel volume. About 6.7% of a standard normal distribution sits above 1.5, so thousands of noise voxels get falsely flagged, on top of the real lesion.

This is the honest lesson: naive thresholding does not scale to 3D. It has no sense of shape, no sense of neighbourhood, and no learned notion of what a real lesion looks like. That gap is exactly why trained 3D segmentation networks exist.

Line by line, the parts that are not obvious:

  • np.mgrid[0:depth, 0:height, 0:width] builds three coordinate grids, one per axis — the standard way to compute a distance from a fixed 3D point across an entire volume at once.
  • dice_score uses logical_and(a, b).sum() for the overlap and a.sum() + b.sum() for the denominator — this is the exact same Dice formula used in real segmentation papers, applied here to a toy volume.
  • The final print shows the real cost of 3D data: a single realistic scan needs 210 MB in float32, purely for the pixel data, before a model has even been loaded.

Common mistakes

Loading a full-size real volume straight into memory without checking size first. A batch of even a few real 200x512x512 volumes can exhaust ordinary GPU memory before training begins.

Judging a segmentation model by 2D Dice on a few slices. A model can look excellent on the slices someone happened to check, and still miss the lesion entirely on adjacent slices — 3D Dice, across the whole volume, is what actually matters.

Assuming intensity alone separates a lesion from noise. As this example shows, it usually does not — real 3D segmentation networks learn shape and context, not a single brightness cutoff.

Try it yourself

Lower the noise: change rng.normal(0, 1, ...) to rng.normal(0, 0.5, ...), halving the noise level. Rerun and watch the Dice score improve — a concrete demonstration of how much simple thresholding depends on a clean, low-noise signal.

What to learn next

Researcher — Mathematics and papers.

3D U-Net

Çiçek et al. (2016) extended the 2D U-Net (Ronneberger et al., 2015) by replacing every 2D convolution, pooling, and upsampling operation with its 3D equivalent, preserving the same encoder-decoder-with-skip-connections structure. The output is a per-voxel class probability volume, trained with a voxel-wise loss such as Dice loss or cross-entropy.

Patch-based training and sliding-window inference

Full-resolution 3D volumes rarely fit in GPU memory at once. The standard workaround trains on randomly cropped patches — small 3D sub-volumes, commonly 128^3 or smaller — and performs inference with a sliding window: the trained network is applied to overlapping patches across the full volume, with overlapping regions averaged or Gaussian-weighted to reduce boundary artefacts.

GPU memory for one 3D conv activation, roughly:

  memory ≈ batch_size * channels * depth * height * width * bytes_per_element
  • bytes_per_element — 4 for float32, 2 for float16/bfloat16

For a typical intermediate activation with 32 channels on a 128^3 patch in float32, this is already on the order of several hundred megabytes — per activation, before accounting for gradients and optimiser state, which is why patch size is one of the first hyperparameters a practitioner has to compromise on when working with limited GPU memory.

nnU-Net

Isensee et al. (2021, Nature Methods) introduced nnU-Net ("no new U-Net"), which does not propose a new architecture. Instead, it automates the configuration decisions around a standard U-Net — patch size, normalisation, resampling to a consistent voxel spacing, and post-processing — based on dataset properties, computed automatically from the training data itself. Across a large range of published medical segmentation benchmarks, nnU-Net's automated configuration has repeatedly outperformed hand-tuned, architecturally novel competitors, making it a standard baseline any new 3D segmentation method is expected to beat.

Cost

For a volume of D x H x W voxels processed with a 3D CNN of k^3 kernels and C channels per layer, one convolutional layer costs approximately:

FLOPs ≈ D * H * W * C_in * C_out * k^3

The k^3 term is the key difference from 2D convolution's k^2 — a 3D kernel costs a full factor of k more per output element than its 2D equivalent, which compounds significantly across a deep network and is the core reason 3D medical imaging models are markedly more compute-hungry than their 2D counterparts of similar depth.

Key references

  • Çiçek, Ö. et al. (2016). 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation. MICCAI.
  • Ronneberger, O., Fischer, P. & Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. MICCAI.
  • Isensee, F. et al. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature Methods 18.
  • Milletari, F., Navab, N. & Ahmadi, S. (2016). V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. 3DV. Introduces Dice loss for volumetric segmentation.

Current state

nnU-Net and its descendants remain the dominant, most-cited baseline across published 3D medical segmentation benchmarks as of this writing. Transformer-based 3D architectures (e.g. UNETR, SwinUNETR) are competitive on specific benchmarks but have not displaced the convolutional nnU-Net baseline as the default first choice in most practical settings, largely due to nnU-Net's strong automated configuration and comparatively lower data requirements.

What to learn next