Autoencoders
An autoencoder squeezes its input through a narrow middle and rebuilds it, and the squeezed middle turns out to be a more useful description of the data than the input was.
- 23 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An autoencoder squeezes its input down to a few numbers, then rebuilds the original from them.
Telling a friend about a film
You watched a three-hour film. Your friend missed it and asks what happened. You tell them in five sentences.
They now know the story. Not every shot, not the exact dialogue, but the shape of it. Enough that they could tell someone else and get it roughly right.
To do that, you had to work out which parts carried the story and which parts did not. Nobody handed you a list. Compressing forced you to decide what mattered.
That forced decision is the whole point of an autoencoder, and the five sentences are worth more than the retelling.
Why it exists
Almost all data comes with no labels. Nobody has written the answers next to it, and paying people to do so is slow and expensive.
An autoencoder needs no labels at all, because the input is its own answer. Show it a photo, ask it to produce that photo, and grade it on how close it got. The data teaches the network using nothing but itself.
That sounds pointless at first. A network that copies its input has learned nothing, and copying is easy.
So you block the easy route. Force everything through a middle layer far too narrow to hold a copy. Say six numbers, where the input had four hundred.
Now copying is impossible. The only way to score well is to work out what those four hundred numbers were really about.
The rebuilt output is not the product. The narrow middle is.
How it works
input encoder code decoder output
400 numbers → [ squeeze ] → 6 numbers → [ expand ] → 400 numbers
▲ │
│ ▼
this is what you wanted compared with the input,
and the difference is the errorThe encoder is the half that squeezes. The decoder is the half that expands. The narrow point in the middle is the bottleneck. The numbers there are the code, sometimes called the latent representation, meaning a hidden description.
Training is ordinary. Push data through, compare the output with the input, adjust the weights to reduce the difference. There is nothing exotic in the loop.
Once trained, most people throw the decoder away and keep the encoder.
Where you have already seen this
- Machine fault alarms. Train on healthy sensor readings. Anything the network rebuilds badly is unlike anything healthy, so it gets flagged.
- Photo cleanup that removes grain from a dark photo, by training the network to rebuild the clean version from a noisy one.
- Recommendations. Your viewing history is thousands of numbers; the code is a few dozen; people with similar codes like similar things.
- Image generators. Tools like Stable Diffusion do not work on pixels. They work inside the code space of an autoencoder, because it is far smaller and cheaper.
- Search by meaning. The code becomes a fingerprint, and similar things get similar fingerprints.
That fourth one is the biggest use of autoencoders today, and it is invisible. Every modern image generator has one inside it doing the compression.
What is honestly hard here
It rebuilds things blurrily. An autoencoder is trained to minimise average difference. The safest way to be a little bit wrong everywhere is to smudge. Faces come out soft. This is why the image systems that use autoencoders attach a GAN-style critic to sharpen the output.
Making new things with it does not go well. It is tempting to invent a code and decode it. The problem is that the network only ever learned the codes it actually produced. The space between them can contain nothing meaningful. Feed it an unused code and you get output that matches nothing. The Developer section measures this.
There is a repair, called a variational autoencoder or VAE, which trains the code space to be filled in rather than patchy. It works, and it blurs even more.
It is not a compression program. Zip and JPEG work on any file. An autoencoder works on data that resembles what it trained on, and produces confident nonsense on anything else.
Its moment as a representation learner has passed. For learning useful descriptions of images and text, the field moved on to other self-supervised methods. Autoencoders remain in use where they earn it: as the compression stage of image generators, and for anomaly detection.
Remember this
- An autoencoder squeezes input through a narrow middle and rebuilds it, using no labels at all.
- The middle is the product. The rebuilt output is a training device.
- It blurs, it does not invent well on its own, and it is alive today mostly inside image generators.
What to learn next
- Transformers — the architecture that now sits on top of these code spaces.
- Diffusion models — the generator that runs inside an autoencoder's latent space.
- Clustering — what to do with the codes once you have them.
Developer — Code and libraries.
Setup
pip install torch scikit-learn numpyEverything runs on CPU in a few seconds. All data is generated in memory.
What an autoencoder buys you over PCA
PCA, principal component analysis, is the classical way to squeeze data into fewer numbers using straight lines only. Comparing against it is the fastest way to see what a network adds.
The data is a spiral: 800 points winding around a coil in three dimensions. One number, the position along the coil, generates all three coordinates. So the data is really one-dimensional, and the shape it lies on is curved.
import numpy as np
import torch
import torch.nn as nn
from sklearn.decomposition import PCA
torch.manual_seed(0)
rng = np.random.default_rng(0)
# 800 points on a spiral in 3-D. One number, t, generates all three coordinates,
# so the data is really one-dimensional -- but the shape it lies on is curved.
t = rng.uniform(0, 3 * np.pi, 800)
X = np.column_stack([np.sin(t), np.cos(t), t / 6.0]).astype(np.float32)
X += rng.normal(0, 0.02, X.shape).astype(np.float32)
Xt = torch.from_numpy(X)
pca = PCA(n_components=1).fit(X)
pca_err = float(((X - pca.inverse_transform(pca.transform(X))) ** 2).mean())
def train(model, steps=4000):
opt = torch.optim.Adam(model.parameters(), lr=0.01)
for _ in range(steps):
opt.zero_grad()
loss = ((model(Xt) - Xt) ** 2).mean() # the input is its own target
loss.backward()
opt.step()
with torch.no_grad():
return ((model(Xt) - Xt) ** 2).mean().item()
linear = nn.Sequential(nn.Linear(3, 1), nn.Linear(1, 3)) # no bend anywhere
curved = nn.Sequential(nn.Linear(3, 16), nn.Tanh(), nn.Linear(16, 1),
nn.Linear(1, 16), nn.Tanh(), nn.Linear(16, 3))
print("squeezing 3 numbers down to 1, then rebuilding them")
print(f" PCA, 1 component error {pca_err:.4f}")
print(f" linear autoencoder error {train(linear):.4f}")
print(f" autoencoder with a bend error {train(curved):.4f}")squeezing 3 numbers down to 1, then rebuilding them PCA, 1 component error 0.2214 linear autoencoder error 0.2214 autoencoder with a bend error 0.0216
The first two numbers will match yours exactly. The third one will not. Adam on a
CPU sums in an order that depends on how many threads your machine uses, and 4000 steps
of training amplify that into the third decimal place. Expect something between about
0.014 and 0.022. The gap to the two PCA lines is the result; the exact digits are not.
The first two lines are identical to four decimal places, and that is not a coincidence.
A linear autoencoder cannot beat PCA. This was proved by Baldi and Hornik in 1989. With no activation function anywhere, the best squeeze-and-rebuild a network can perform is a projection onto the same subspace PCA finds. It reaches the same error by a slower route.
The weights are not the same as PCA's — any rescaling of the code gives the same reconstruction — but the plane it flattens onto is.
Adding one Tanh on each side drops the error by a factor of ten. The network can now follow a curve. It learned that a point on a spiral is described by how far along the spiral it sits, which is the true one-dimensional description of this data, and which no straight line can express.
That gap is the entire reason to reach for a network here. An autoencoder with no activation function is PCA with extra steps. If someone shows you one, it has a bug.
Finding faults nobody labelled
This is where autoencoders still earn their place in production. Train on normal data only. Anything the model rebuilds badly is unlike normal.
import numpy as np
import torch
import torch.nn as nn
torch.manual_seed(0)
rng = np.random.default_rng(1)
# 600 healthy machines. Four sensors, but only two independent things are happening:
# a load level and a temperature offset. The other two sensors are mixtures of them.
load = rng.uniform(0, 1, 600)
temp = rng.uniform(0, 1, 600)
normal = np.column_stack([load, temp, 0.8 * load + 0.2 * temp, 0.5 * load + 0.5 * temp])
normal = (normal + rng.normal(0, 0.02, normal.shape)).astype(np.float32)
model = nn.Sequential(nn.Linear(4, 8), nn.Tanh(), nn.Linear(8, 2),
nn.Linear(2, 8), nn.Tanh(), nn.Linear(8, 4))
opt = torch.optim.Adam(model.parameters(), lr=0.01)
Xn = torch.from_numpy(normal)
for _ in range(3000):
opt.zero_grad()
((model(Xn) - Xn) ** 2).mean().backward()
opt.step()
def error(rows):
x = torch.tensor(np.array(rows, dtype=np.float32))
with torch.no_grad():
return ((model(x) - x) ** 2).mean(dim=1).numpy()
healthy = error(normal)
threshold = float(np.percentile(healthy, 99)) # allow 1 in 100 healthy machines to alarm
print(f"typical rebuild error on healthy machines: {healthy.mean():.5f}")
print(f"alarm threshold (99th percentile) : {threshold:.5f}")
print()
checks = [
("healthy machine", [0.60, 0.40, 0.56, 0.50]),
("sensor 3 stuck high", [0.60, 0.40, 0.99, 0.50]),
("sensor 4 disconnected", [0.60, 0.40, 0.56, 0.00]),
("all four off the scale", [2.00, 2.00, 2.00, 2.00]),
("odd, but self-consistent", [0.02, 0.02, 0.02, 0.02]),
]
for name, row in checks:
e = error([row])[0]
print(f"{name:26s} error {e:.5f} {'ALARM' if e > threshold else 'ok'}")typical rebuild error on healthy machines: 0.00021 alarm threshold (99th percentile) : 0.00091 healthy machine error 0.00001 ok sensor 3 stuck high error 0.03031 ALARM sensor 4 disconnected error 0.04551 ALARM all four off the scale error 0.20022 ALARM odd, but self-consistent error 0.00008 ok
Three faults caught, at 33, 50 and 220 times the alarm threshold. No fault was ever labelled and no fault was ever shown to the model during training. It learned what "healthy" looks like and everything else stands out.
The two-number bottleneck is doing the work. Four sensors, but only two independent things are happening. Squeezed to two numbers and rebuilt, a healthy machine survives the trip almost unchanged.
Now read the last line, because it is the honest limit of the method. The reading [0.02, 0.02, 0.02, 0.02] is extreme: every sensor at the very bottom of its range. The model waved it through.
It did so correctly. Check the arithmetic: 0.8 × 0.02 + 0.2 × 0.02 = 0.02, and 0.5 × 0.02 + 0.5 × 0.02 = 0.02. The reading obeys every relationship the model learned. It is not inconsistent — it is only unusual.
An autoencoder detects readings that break the pattern, not readings that are rare. For rare-but-consistent readings you need a density estimate, or an isolation forest, or a plain range check on each sensor. Running a simple range check alongside is standard practice and costs nothing.
The holes in the middle
The last thing to check is whether you can invent a code and decode it. Here is the spiral again, with a deliberate gap: no training data for one stretch of the coil.
import numpy as np
import torch
import torch.nn as nn
torch.manual_seed(0)
rng = np.random.default_rng(0)
# The same spiral, but with a deliberate gap: no data for t between 4 and 6.
t = rng.uniform(0, 3 * np.pi, 1200)
t = t[(t < 4) | (t > 6)]
X = np.column_stack([np.sin(t), np.cos(t), t / 6.0]).astype(np.float32)
Xt = torch.from_numpy(X)
enc = nn.Sequential(nn.Linear(3, 16), nn.Tanh(), nn.Linear(16, 1))
dec = nn.Sequential(nn.Linear(1, 16), nn.Tanh(), nn.Linear(16, 3))
model = nn.Sequential(enc, dec)
opt = torch.optim.Adam(model.parameters(), lr=0.01)
for _ in range(4000):
opt.zero_grad()
((model(Xt) - Xt) ** 2).mean().backward()
opt.step()
with torch.no_grad():
codes = enc(Xt).numpy().ravel()
lo, hi = codes.min(), codes.max()
grid = torch.linspace(lo, hi, 400).unsqueeze(1) # invent codes across the range
made = dec(grid).numpy()
# how far is each invented point from the nearest real one?
d = np.sqrt(((made[:, None, :] - X[None, :, :]) ** 2).sum(-1)).min(1)
print(f"code values seen in the data run from {lo:.2f} to {hi:.2f}")
print("distance from an invented point to the nearest real point:")
print(f" average {d.mean():.3f} worst {d.max():.3f}")
print(f" fraction of invented points further than 0.2 away: {(d > 0.2).mean():.2f}")code values seen in the data run from -4.79 to 6.86 distance from an invented point to the nearest real point: average 0.069 worst 0.550 fraction of invented points further than 0.2 away: 0.12
These figures move in the last decimal place between machines, for the threading reason given above. The shape of the answer — a worst miss over half a radius, and roughly one invented code in eight landing nowhere — is what holds.
The spiral has radius 1, so a miss of 0.550 is more than half a radius — the decoded point is nowhere near the shape it was meant to lie on.
Twelve percent of invented codes land somewhere the data never goes. Every one of those codes sits inside the observed range, between two values the encoder genuinely produced. There is nothing at the surface to warn you.
This is the reason a plain autoencoder is a poor generator. The bottleneck was optimised so that the codes it produces decode well. Nothing in the training objective asked the space between them to mean anything.
The variational autoencoder attacks exactly this. It makes the encoder output a small blob of possible codes rather than a single point, and adds a penalty pulling those blobs towards a standard bell-shaped cloud. Overlapping blobs leave no gaps to fall into. The cost is more blur, because a code now has to decode acceptably for a whole neighbourhood, not one point.
Common mistakes
No activation function. Measured above: it is PCA, and slower.
A bottleneck as wide as the input. With enough width the network learns the identity function, the loss goes to nearly zero, and the code carries no structure. If your reconstruction is suspiciously perfect, this is why. Narrow the bottleneck, or use a denoising or sparse variant instead.
Not scaling the input. Reconstruction error is a sum over columns. A column measured in thousands dominates the loss and the model spends all its capacity there.
Using nn.MSELoss with a Sigmoid output on data outside [0, 1]. The output cannot reach the target and the loss plateaus at a value you cannot explain.
Judging an anomaly detector on training data. The threshold must come from held-out normal data. Fitted on the same rows the model trained on, it is far too tight and you will alarm constantly in production.
Expecting it to compress arbitrary files. It reproduces the kind of data it was trained on, and produces confident nonsense on anything else.
Reaching for an autoencoder to learn image features. For that job, masked or contrastive self-supervised methods now do better. Reach for an autoencoder when you want compression into a working latent space, or reconstruction error as a signal.
Try it yourself
In beats_pca.py, widen the bottleneck from 1 to 3, so the code is as wide as the input.
Predict the error before running it. Both the linear and the curved model should now drop to nearly zero — and the code has become a re-encoding of the input rather than a description of it. Print the codes and see whether they carry any structure at all.
Then take fault_alarm.py and change the bottleneck from 2 to 4, so the code is as wide as the sensor count. Print each fault's error as a multiple of the threshold rather than reading the raw numbers.
Every fault still alarms, so at a glance nothing has changed. The margins tell a different story. sensor 4 disconnected goes from about 50 times the threshold down to under 3 times — a detector that was shouting is now whispering. With four code slots for four sensors, the network can pass sensor 4 through almost untouched instead of rebuilding it from the other three, so breaking it costs almost nothing.
The alarm was never about the network's power. It was about the bottleneck being narrow enough that only genuine patterns fit through it.
What to learn next
- Transformers — the architecture that now sits on top of these code spaces.
- Diffusion models — the generator that runs inside an autoencoder's latent space.
- Clustering — what to do with the codes once you have them.
Researcher — Mathematics and papers.
Definition and objective
An autoencoder is a pair of parameterised maps trained to approximate the identity through a constrained intermediate space:
f_θ : R^n → R^k encoder
g_φ : R^k → R^n decoder
θ*, φ* = argmin E_{x ~ p_data} [ L( x, g_φ(f_θ(x)) ) ]n— input dimension;k— code dimensionL— reconstruction loss: squared error for continuous data, Bernoulli cross-entropy for data in[0,1]h = f_θ(x) ∈ R^k— the code, or latent representation
An autoencoder is undercomplete when k < n, which is the constraint that forces structure to be learned. Overcomplete autoencoders with k ≥ n need a different regulariser, since the identity map is available and optimal.
The linear case is exactly PCA
Baldi & Hornik (1989) analysed the linear autoencoder g(f(x)) = W_2 W_1 x under squared loss. Every critical point corresponds to projection onto a subspace spanned by some subset of eigenvectors of the data covariance. All are saddle points except the global minimum, which projects onto the subspace of the top k eigenvectors — the PCA subspace.
The solution is unique only up to an invertible k × k transform: W_1 → A W_1, W_2 → W_2 A^{-1} leaves the reconstruction unchanged. So a linear autoencoder recovers the principal subspace, not the ordered principal axes. The Developer section measures the consequence: identical reconstruction error to four decimals, from different weights.
Kunin et al. (2019) show that adding L2 regularisation to a linear autoencoder breaks this symmetry and recovers the ordered principal directions, which explains why the two methods are not always as interchangeable as the theorem suggests.
Non-linear encoders and decoders generalise this to non-linear manifolds, and the factor-of-ten error reduction on the spiral is that generalisation observed directly.
Regularised variants
Denoising autoencoder (Vincent et al., 2008). Corrupt the input and reconstruct the clean original: minimise E_{x, x̃ ~ q(x̃|x)} [ L(x, g(f(x̃))) ]. The identity map is no longer a solution even when overcomplete. Vincent et al. (2010) show that for small additive Gaussian corruption the objective is equivalent to penalising the Frobenius norm of the encoder Jacobian, and Alain & Bengio (2014) show the trained model estimates the score ∂ log p(x) / ∂x — the connection that leads directly to score-based diffusion models.
Contractive autoencoder (Rifai et al., 2011). Add λ ‖ ∂f(x)/∂x ‖_F^2 explicitly, making the representation locally insensitive to input perturbation while reconstruction keeps it from collapsing.
Sparse autoencoder. Penalise the average activation of code units, via an L1 term or a KL penalty towards a low target rate. Recently prominent again for a different purpose: sparse autoencoders trained on the residual stream of language models to extract interpretable, monosemantic features (Bricken et al., 2023; Templeton et al., 2024).
The variational autoencoder
Kingma & Welling (2013) replace the deterministic code with a latent-variable model p_θ(x) = ∫ p_θ(x|z) p(z) dz, with p(z) = N(0, I). The marginal is intractable, so introduce a variational posterior q_φ(z|x) and optimise the evidence lower bound:
log p_θ(x) ≥ E_{q_φ(z|x)} [ log p_θ(x|z) ] − KL( q_φ(z|x) ‖ p(z) ) = ELBO(x)q_φ(z|x)— the encoder, typicallyN(μ_φ(x), diag(σ²_φ(x)))p_θ(x|z)— the decoder likelihood- The first term is reconstruction; the second is the regulariser that fills the latent space in
- The gap between
log p_θ(x)and the ELBO is exactlyKL(q_φ(z|x) ‖ p_θ(z|x)), so maximising the bound both fits the data and sharpens the posterior approximation
The reparameterisation trick makes this trainable: sample z = μ_φ(x) + σ_φ(x) ⊙ ε with ε ~ N(0, I). The randomness is moved to an input, so the gradient flows into μ_φ and σ_φ with variance far lower than a score-function estimator gives.
The KL term is precisely what a plain autoencoder lacks, and precisely what the Developer section's latent-hole measurement diagnoses. Overlapping posteriors leave no unclaimed regions in latent space.
Two well-documented pathologies:
- Posterior collapse. With a sufficiently expressive decoder — an autoregressive one especially — the model can maximise the ELBO by setting
q_φ(z|x) ≈ p(z), driving the KL term to zero and ignoringzentirely. Remedies: KL annealing, the free-bits constraint (Kingma et al., 2016), or weakening the decoder. - Blur. A Gaussian decoder likelihood makes the reconstruction term a squared error, whose optimum under uncertainty is the conditional mean. Averaging over plausible completions is a smudge. This is the mirror image of the GAN's mode collapse: mode-covering blurs, mode-seeking sharpens.
β-VAE (Higgins et al., 2017) scales the KL term by β > 1, trading reconstruction for disentangled factors. Locatello et al. (2019) subsequently proved that unsupervised disentanglement is impossible without inductive biases or supervision, and that reported results were highly seed-dependent — a useful corrective to that literature.
Discrete latents and the modern role
VQ-VAE (van den Oord, Vinyals & Kavukcuoglu, 2017) replaces the continuous code with a lookup into a learned codebook e_1..e_K, z_q = e_{argmin_j ‖ z_e − e_j ‖}. The argmin is non-differentiable, so the straight-through estimator copies the decoder gradient to the encoder, and two auxiliary terms keep the codebook and encoder aligned. Discrete codes avoid posterior collapse and let a separate autoregressive prior be fitted over the code grid afterwards.
This is where autoencoders sit in current systems. VQGAN (Esser, Rombach & Ommer, 2021) adds a patch discriminator and a perceptual loss to the VQ-VAE, giving sharp reconstructions at high compression. Latent diffusion (Rombach et al., 2022) then runs the entire diffusion process inside that latent space rather than on pixels, cutting compute by roughly an order of magnitude at 512×512. Stable Diffusion is a diffusion model over an autoencoder's code space.
So the autoencoder is not the generator. It is the coordinate system the generator works in, and the adversarial term from the GAN lesson is what keeps its decoder from blurring.
Masked autoencoders (He et al., 2022) mask 75 percent of image patches and reconstruct them with an asymmetric encoder-decoder, where the encoder sees only visible patches. High masking ratios make the task non-trivial and the asymmetry makes pretraining cheap. BERT (Devlin et al., 2019) is the same idea for text, and both are denoising autoencoders under a specific corruption process.
Autoencoders for anomaly detection, with caveats
Score by reconstruction error, s(x) = ‖ x − g(f(x)) ‖², thresholded on held-out normal data. The Developer section demonstrates both the strength and the failure mode.
Two documented caveats matter in production.
Generalisation to anomalies. A sufficiently expressive autoencoder often reconstructs unseen anomalous inputs well, because the learned mapping is not restricted to the training manifold. Memory-augmented variants (Gong et al., 2019) restrict the decoder input to a convex combination of stored prototypes, forcing anomalies to be reconstructed from normal patterns and inflating their error.
Out-of-distribution likelihood. Nalisnick et al. (2019) showed deep generative models can assign higher likelihood to out-of-distribution inputs than to their own training data — a VAE trained on CIFAR-10 scoring SVHN higher. Reconstruction error is not likelihood, but the result is a standing warning against treating any single generative score as an anomaly detector without validation.
The practical consequence is the point made with data above: an autoencoder flags inputs that violate learned relationships, not inputs that are only rare. Pair it with a marginal density check or a per-feature range check.
Cost
Forward O( |θ| + |φ| ) per sample; roughly 2× a classifier of the same width
Latent size k << n is the constraint doing the work
Latent diffusion operating at 64×64×4 instead of 512×512×3 is a ~48× reduction
in the elements the diffusion model must processPapers
- Baldi, P. & Hornik, K. (1989). Neural Networks and Principal Component Analysis. Neural Networks 2(1).
- Hinton, G. & Salakhutdinov, R. (2006). Reducing the Dimensionality of Data with Neural Networks. Science 313(5786).
- Vincent, P. et al. (2008). Extracting and Composing Robust Features with Denoising Autoencoders. ICML.
- Rifai, S. et al. (2011). Contractive Auto-Encoders. ICML.
- Kingma, D. P. & Welling, M. (2013). Auto-Encoding Variational Bayes. — arxiv.org/abs/1312.6114
- Alain, G. & Bengio, Y. (2014). What Regularized Auto-Encoders Learn from the Data-Generating Distribution. JMLR 15.
- Higgins, I. et al. (2017). beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. ICLR.
- van den Oord, A., Vinyals, O. & Kavukcuoglu, K. (2017). Neural Discrete Representation Learning (VQ-VAE). NeurIPS — arxiv.org/abs/1711.00937
- Nalisnick, E. et al. (2019). Do Deep Generative Models Know What They Don't Know? ICLR — arxiv.org/abs/1810.09136
- Locatello, F. et al. (2019). Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. ICML.
- Esser, P., Rombach, R. & Ommer, B. (2021). Taming Transformers for High-Resolution Image Synthesis (VQGAN). CVPR — arxiv.org/abs/2012.09841
- He, K. et al. (2022). Masked Autoencoders Are Scalable Vision Learners. CVPR — arxiv.org/abs/2111.06377
- Rombach, R. et al. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. CVPR — arxiv.org/abs/2112.10752
What to learn next
- Transformers — the architecture that now sits on top of these code spaces.
- Diffusion models — the generator that runs inside an autoencoder's latent space.
- Clustering — what to do with the codes once you have them.