Gigapixel pathology slides
A digitised pathology slide can be too large to open in memory at once, forcing every model to work on small tiles instead of the whole picture.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A digitised pathology slide can be too large for a computer to open at once.
Think about a satellite photo of an entire city, zoomed in enough to see individual cars. That single image would be enormous — far too large to open normally on a phone.
A digital pathology slide is that same idea, applied to tissue instead of a city. A pathologist scans a thin tissue slice on a glass slide, fine enough to see individual cells. The result is a single image that can be tens of thousands of pixels wide.
Why it exists
A biopsy sample gets sliced, stained, and mounted on a glass slide, then examined under a microscope. Digitising that slide means scanning it at high resolution. It has to show the same cell-level detail a pathologist would see by eye.
Cells are small. Slides are physically large. Combine those two facts, and the digital file explodes in size. A single slide can reach tens of gigapixels — sometimes larger than any other image type used in medicine.
No ordinary neural network can load an image that large directly. It would exceed the memory of even a large server, let alone a single graphics card.
How it works
One giant slide (say, 80,000 x 60,000 pixels)
|
v
cut into thousands of small tiles
(each tile maybe 256 x 256 pixels)
|
v
a model looks at tiles, not the whole slide at onceThe slide gets broken into small, ordinary-sized pieces first. Only then can a standard CNN look at any of it.
Where you have already seen it
- Google Maps zooming in loads small map tiles as you zoom, rather than one enormous image of the whole world.
- A very high-resolution museum scan of a large painting is often browsed the same tiled way.
An honest warning
Most of a pathology slide is empty glass, or ordinary tissue with nothing wrong. The actual diagnostic detail might sit in a tiny handful of tiles out of thousands. That imbalance — one needle, an enormous haystack — shapes almost every technique covered in the rest of this section.
Any system built on this pipeline still needs a pathologist's validation. Regulatory clearance matters too, before it touches a real diagnosis.
Remember this
- A digitised pathology slide can be far too large to load into memory at once.
- Slides are broken into small tiles before any standard model can process them.
- The diagnostic detail often sits in only a small fraction of a slide's total tiles.
What to learn next
- Multiple instance learning — how a model learns from a slide when only the tiles, not the diagnosis location, are known.
- Image classification — the technique applied to each individual tile.
- Choosing a backbone — picking the CNN that actually processes each tile.
Developer — Code and libraries.
Setup
pip install numpyMinimal runnable code
import numpy as np
rng = np.random.default_rng(0)
# A small stand-in for a pathology slide image. A real digitised slide is
# closer to 100,000 x 100,000 pixels -- far too large to load into memory
# or feed into a CNN in one piece.
slide = rng.integers(200, 255, size=(4096, 4096, 3), dtype=np.uint8)
def tile_slide(image, tile_size=256, stride=256):
"""Cut a large image into a grid of small tiles a CNN can actually load."""
h, w = image.shape[:2]
tiles = []
for y in range(0, h - tile_size + 1, stride):
for x in range(0, w - tile_size + 1, stride):
tiles.append(image[y:y + tile_size, x:x + tile_size])
return tiles
tiles = tile_slide(slide, tile_size=256)
print(f"slide shape: {slide.shape}")
print(f"number of 256x256 tiles from this small slide: {len(tiles)}")
print(f"one tile shape: {tiles[0].shape}")
# What the same maths gives for a real, full-size slide.
real_h, real_w = 100_000, 100_000
real_tiles = (real_h // 256) * (real_w // 256)
print(f"the same tiling on a real {real_h}x{real_w} slide: {real_tiles:,} tiles")
print("most of those tiles are background glass, not tissue --")
print("a model trained on all of them wastes almost all its capacity")slide shape: (4096, 4096, 3) number of 256x256 tiles from this small slide: 256 one tile shape: (256, 256, 3) the same tiling on a real 100000x100000 slide: 152,100 tiles most of those tiles are background glass, not tissue -- a model trained on all of them wastes almost all its capacity
What actually happened
The small 4096x4096 toy slide produced 256 tiles — manageable. Scale the same tiling formula up to a realistic 100,000 x 100,000 slide, and the count jumps to 152,100 tiles, from a single slide. A real pathology dataset might contain thousands of slides, each producing this many tiles.
That number sets up the rest of this section. No pathologist labels 152,100 individual tiles by hand. The label that actually exists is usually one diagnosis for the whole slide — which is exactly the setting the next lesson, multiple instance learning, is built to handle.
Line by line, the parts that are not obvious:
tile_slideuses non-overlapping tiles here (stride == tile_size), the simplest case. Real pipelines often overlap tiles slightly, to avoid missing a finding that straddles a tile boundary.- The loop bound
h - tile_size + 1avoids generating a partial tile at the image's edge — a common off-by-one bug when tiling real images. real_tilesuses simple integer division, ignoring the small remainder at each edge — close enough for this illustration, though a real pipeline needs to decide explicitly how to handle a slide's ragged edges.
Common mistakes
Tiling with a fixed grid and ignoring which tiles contain actual tissue. Most of a slide is empty glass — a basic tissue-detection step (thresholding on color or saturation) before tiling saves enormous wasted computation.
Assuming a slide's magnification is standardised. Different scanners save at different base resolutions — tile size in pixels does not automatically correspond to the same physical tissue size across two different slides.
Loading every tile from a slide into memory at once for training. With well over a hundred thousand tiles per slide, this is rarely feasible — pipelines stream tiles from disk, or from the slide file directly, on demand.
Try it yourself
Change tile_size to 512 and rerun. Notice the tile count drops by roughly a factor of four, not two — tiling area scales with the square of the tile size, not linearly.
What to learn next
- Multiple instance learning — learning from a slide's tiles when only one label exists for the whole slide.
- Choosing a backbone — which CNN architecture processes each individual tile.
- Datasets that do not fit in memory — the general engineering pattern this lesson's scale problem requires.
Researcher — Mathematics and papers.
Whole-slide image formats
Digitised pathology slides are typically stored as pyramidal, tiled TIFF variants (Aperio SVS, Hamamatsu NDPI, generic pyramidal TIFF), storing the same image at multiple resolution levels — full resolution down to a low-resolution thumbnail — so software can efficiently read only the resolution level and region actually needed, without decoding the entire file. OpenSlide is the standard open-source library for reading these formats across vendors.
Scale, concretely
A slide scanned at 0.25 microns per pixel (a common resolution for 40x-equivalent scanning) covering a 15mm x 15mm tissue section produces an image roughly:
15mm / 0.25um = 60,000 pixels per sideAt three color channels and 8 bits per channel, that is 60,000^2 * 3 ≈ 10.8 gigabytes uncompressed for a single slide — before compression, and before considering that a case may include multiple slides.
Weakly-supervised approaches
Because pixel-level or even tile-level labels are rarely available, whole-slide classification is typically framed as multiple instance learning (covered in depth in the next lesson), with a slide-level label and no ground truth for which specific tiles drove that label. CLAM (Lu et al., 2021, Nature Biomedical Engineering) is a widely adopted attention-based MIL framework specifically designed for whole-slide classification with only slide-level labels, using clustering-constrained attention to additionally identify which tiles the model considered most diagnostically relevant.
Cost
For a slide with N tissue-containing tiles, extracting a feature vector per tile with a CNN backbone costs O(N) forward passes — for a realistic slide with N on the order of 10^4 to 10^5 tiles after background removal, this dominates the computational cost of the entire pipeline, far exceeding the cost of the downstream MIL aggregation step itself, which typically operates on the resulting small set of tile-level feature vectors rather than raw pixels.
Key references
- Lu, M. et al. (2021). Data-efficient and weakly supervised computational pathology on whole-slide images. Nature Biomedical Engineering 5. Introduces CLAM.
- Campanella, G. et al. (2019). Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nature Medicine 25. A large-scale clinical deployment study using MIL on real pathology slides.
- Goode, A. et al. (2013). OpenSlide: A vendor-neutral software foundation for digital pathology. Journal of Pathology Informatics 4.
Current state
Foundation models pretrained specifically on large collections of pathology tiles (e.g. via self-supervised contrastive or masked-image objectives on millions of tiles) are an active and rapidly evolving area, aiming to produce general-purpose tile embeddings that transfer across cancer types and institutions with less task-specific labelled data. This remains a fast-moving research area, and any specific tool or benchmark result should be checked against current literature rather than treated as settled.
What to learn next
- Multiple instance learning — the weakly-supervised learning framework this scale problem requires.
- Datasets that do not fit in memory — general engineering patterns for data at this scale.
- Why your model fails at the next hospital — how differences between scanners affect pathology models specifically.