Reading an embedding plot honestly
t-SNE and UMAP pictures answer only one question — which points belong together — and every other thing your eye reads from them can be an artefact of a knob setting.
- 8 min read
- 3 reading levels
- Published
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An embedding plot is a metro map, not a satellite photo — trust the connections, never the distances.
Look at any city's metro map. It is superb at one question: which stations link together, and which line takes you there. But measure the map with a ruler and you will plan nonsense. It stretches crowded central areas and shrinks empty stretches, so that everything stays readable. Two stations a thumb-width apart may be forty minutes distant.
t-SNE and UMAP plots are metro maps of your data. They exist to answer: which points belong together? Everything else your eye wants to read from them — sizes, gaps, directions — the map never promised.
Why this lesson exists
Embedding plots are the most shared pictures in machine learning, and the most over-read. Real decisions get made from them: "these two customer groups are basically one", "this cluster is tiny, ignore it", "the model separates classes, ship it". Each of those claims reads the ruler, not the metro lines — and each can flip when one of the drawing settings changes. A picture that changes its story with a setting was never evidence for that story.
What you may and may not read
SAFE to read NOT safe to read
──────────────── ────────────────
these points form an island this island is bigger, so more varied
these two islands are separate these islands are close, so similar
this point sits in the wrong the x-axis means something
island (check its label!) the empty space between means distanceThe three honest habits
- Wiggle the knobs. Rerun at several perplexity or neighbour settings. Believe only what survives all of them.
- Colour by everything. Colour by label, then by collection date, then by data source. If islands align with source instead of label, you found a data problem, not a discovery. This is how leakage and batch effects get caught by eye.
- Confirm in the original space. An island is a hypothesis. Test it with numbers computed on the real data — a clustering score, a classifier, an average difference — never with the picture alone.
A real example you have seen
Election maps distort honestly on purpose: a cartogram inflates tiny dense constituencies so votes stay visible. Useful map, terrible ruler. Whoever reads land area as vote count misreads the country. Same picture-versus-measurement trap, same cure: know what the map preserves.
Remember this
- Embedding plots answer "what groups together?" — nothing else.
- Sizes, gaps and axes are layout artefacts; knobs can change them freely.
- Islands are hypotheses to verify in the original data, not conclusions.
What to learn next
- t-SNE — the mechanics behind the artefacts measured here.
- UMAP — the faster map, same reading rules.
- Clustering — turning "these look grouped" into a testable result.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2. Exact numbers vary by version and platform; the instability itself is the reproducible finding.
Catching the map lying, with numbers
Fix the data. Fix the seed. Change only perplexity — the "how many friends count" knob — and measure the gap between the same two digit-clusters on each map.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
X, y = load_digits(return_X_y=True)
def gap(emb, a, b):
"""Distance between two digit-cluster centres, as drawn on the map."""
return np.linalg.norm(emb[y == a].mean(0) - emb[y == b].mean(0))
print("perplexity 0-to-6 gap 1-to-7 gap ratio")
for perp in [5, 30, 100]:
emb = TSNE(n_components=2, perplexity=perp, random_state=0).fit_transform(X)
g06, g17 = gap(emb, 0, 6), gap(emb, 1, 7)
print(f"{perp:10d} {g06:10.1f} {g17:10.1f} {g06/g17:5.2f}")perplexity 0-to-6 gap 1-to-7 gap ratio
5 55.6 56.4 0.99
30 46.6 28.9 1.61
100 21.2 14.7 1.44The walkthrough
Read the ratio column and pick your story. At perplexity 5, the map says zeros sit as far from sixes as ones do from sevens (ratio 0.99). At perplexity 30, the same map style says the 0-6 gap is 61% larger (1.61). Same data, same seed, one knob — and "which pair is more different" reverses. Any conclusion of the form "cluster A is closer to B than to C" is at the mercy of a setting the reader of your plot will never see.
The islands themselves survive. At every setting, all ten digits still form ten identifiable islands — that is the part t-SNE actually promises, and it holds. The lie lives entirely in the geometry between islands.
The same test works for UMAP — swap in umap.UMAP(n_neighbors=...) and sweep 5, 15, 50. UMAP's between-island layout is steadier than t-SNE's but still moves; measure it rather than assuming.
A pre-publication checklist
Before an embedding plot goes into a report or a paper:
- Ran at 3+ settings of perplexity /
n_neighbors; conclusions survived. - Ran at 2+ seeds; islands stable in membership (not in position — position is free).
- Coloured by label and by batch, date, source, annotator. No suspicious alignment.
- Any distance-flavoured claim rechecked with numbers in the original space.
- Caption states the tool, the settings, and "distances between clusters are not meaningful".
Common mistakes
Clustering on the 2-D coordinates and reporting those clusters. Density on the page was manufactured — t-SNE equalises cluster sizes, UMAP's min_dist sets packing. Run clustering in the original or PCA space, and use the map only to display the result.
"The classes overlap in the plot, so they are inseparable." The plot is 2-D; your model works in the full space. Overlapping islands often separate perfectly with all features — check a classifier's evaluation before declaring defeat. The reverse error — "islands separate, ship the model" — is equally common: neighbour purity on a map is not accuracy on new data.
Publishing a single beautiful run. If nine runs looked messy and one looked clean, showing the clean one is seed-shopping. The honest figure shows a representative run, with settings stated.
Believing empty space. The blank canvas between islands is not "no man's land" in feature space. Points landing there in a new batch are not automatically anomalies; verify with a real outlier score.
Try it yourself
Extend the script's inner loop over seeds 0, 1, 2 at fixed perplexity 30 and print the same table. You will find the seed changes gaps less than perplexity did — then flip a coin: rotate the perplexity-5 embedding 90 degrees and note that even axis direction was never stable.
What to learn next
- t-SNE — the mechanics behind the artefacts measured here.
- UMAP — the faster map, same reading rules.
- Clustering — turning "these look grouped" into a testable result.
Researcher — Mathematics and papers.
Why the geometry cannot be trusted, formally
t-SNE's loss KL(P || Q) is dominated by terms with large p_ij — near neighbours. Pairs in different clusters have p_ij near zero, and the objective is almost indifferent to where such pairs land, subject to aggregate repulsion. Inter-cluster distances therefore occupy a near-flat region of the loss: many arrangements are equally optimal, and which one you get is decided by initialisation, early exaggeration and the annealing path. The per-point bandwidth sigma_i further equalises effective density, erasing cluster-scale information by design — the direct cause of the island-size lie. UMAP's cross-entropy adds a stronger repulsive structure, tying layouts somewhat closer to the k-NN graph's coarse topology, but between-component placement remains under-constrained.
The literature of caution
- Wattenberg, Viegas and Johnson (2016), How to use t-SNE effectively, Distill — the canonical interactive catalogue of artefacts: cluster sizes, distances, and shapes as functions of perplexity, including structure conjured from pure noise.
- Kobak and Berens (2019), The art of using t-SNE for single-cell transcriptomics, Nature Communications — PCA initialisation, multi-scale similarities, and exaggeration schedules as partial remedies for global structure.
- Kobak and Linderman (2021), Nature Biotechnology — initialisation, not algorithm, drives much of the t-SNE-vs-UMAP global-structure folklore.
- Chari and Pachter (2023), The specious art of single-cell genomics, PLoS Computational Biology — the sharpest indictment: 2-D embeddings necessarily distort high-dimensional geometry (dimensionality arguments), and an autoencoder trained to produce any target shape (including an elephant outline) can match standard pipelines on the usual quality metrics. The paper's constructive point: task-specific quantitative analyses should replace geometric readings of 2-D maps.
Quantitative reading aids
If a claim must rest on an embedding, attach a metric: trustworthiness (sklearn's trustworthiness; Venna and Kaski, 2001) penalises new false neighbours; continuity penalises lost true neighbours; k-NN label agreement (used across these lessons) audits neighbourhood purity; Procrustes distance between runs quantifies stability. Report the metric at the k that matters to the claim — neighbourhood metrics at k = 10 say nothing about a statement comparing cluster centroids.
The deeper limit
A theorem-shaped intuition worth keeping: by the Johnson-Lindenstrauss lower bounds, preserving all pairwise distances of n points even approximately requires Omega(log n / eps^2) dimensions — 2 is not that, for any interesting n (see random projections). Every 2-D map of high-dimensional data is therefore lossy by arithmetic, before any algorithm choices. The only free choice is which information to sacrifice, and t-SNE/UMAP's choice — keep neighbourhoods, sacrifice geometry — is a good one precisely when readers respect it.
What to learn next
- t-SNE — the mechanics behind the artefacts measured here.
- UMAP — the faster map, same reading rules.
- Clustering — turning "these look grouped" into a testable result.