Clustering in Depth

Using clusters as features

A fitted clustering can feed a supervised model new columns — cluster ids and distances to centres — turning unlabelled structure into predictive signal, if you avoid the leakage trap.

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A clustering does not have to be the final answer — its output can become extra input columns that make a prediction model smarter.

Think about PIN codes. The postal department drew zones over the map long ago, for its own sorting convenience. Yet today a delivery app predicts arrival time better because "PIN 500081" is a column in its data. Someone else's grouping became your predictive feature.

Clustering lets you invent such zones for any data, not only maps. Group your customers, then hand each customer's group membership to the model that predicts churn.

Why this matters

Some models can only draw straight dividing lines through the data. A pattern shaped like two interlocking crescents defeats them — no straight line separates crescents. This limitation is called being a linear model, and logistic regression is the classic example.

Clusters offer a cheap escape. Describe each point by its relationship to a handful of cluster centres. Which centre is nearest, and how far away each one sits. Curved structure gets translated into columns a straight-line model can use. You upgrade the features instead of the model.

How it works

original columns          new columns from a fitted clustering

x1, x2             ->     x1, x2,  dist to centre 1,
                                   dist to centre 2,
                                   ...
                                   dist to centre 12

The distances are like asking: how far is this house from the railway station, the market, the hospital? Twelve landmarks give twelve numbers that pin down where the point sits in the data's shape — information the raw coordinates express only crookedly.

One rule is sacred. The clustering must be learned from training data only, then applied unchanged to test data. Learn the landmarks from data the exam will be graded on, and your score becomes a lie. That mistake is called leakage: letting information from the test set sneak into training.

A real example you have seen

Fraud systems cluster normal transaction behaviour into patterns. Each new transaction then carries features like "distance to the nearest normal pattern". A far-from-every-landmark transaction is exactly the kind that deserves a closer look.

Remember this

  • Cluster ids and distances to centres make honest, useful new columns.
  • They let straight-line models handle curved patterns.
  • Fit the clustering on training data only — anything else is leakage.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; last-digit drift across platforms is normal.

Rescuing logistic regression with 12 landmarks

The two-moons shape is unlearnable for a straight-line classifier. Cluster distances translate it:

cluster_features.py
import numpy as np
from sklearn.cluster import KMeans
from sklearn.datasets import make_moons
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = make_moons(n_samples=600, noise=0.15, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)

plain = LogisticRegression().fit(Xtr, ytr)
print("raw features only:      ", round(plain.score(Xte, yte), 3))

km = KMeans(n_clusters=12, n_init=10, random_state=0).fit(Xtr)   # fit on train ONLY

def with_cluster_features(X):
    return np.hstack([X, km.transform(X)])   # distance to each of the 12 centres

rich = LogisticRegression(max_iter=1000).fit(with_cluster_features(Xtr), ytr)
print("plus cluster distances: ", round(rich.score(with_cluster_features(Xte), yte), 3))
Output
raw features only:       0.86
plus cluster distances:  0.913

The walkthrough

km.transform(X) is the underused gem. Everyone knows predict; transform returns the distance from each row to every centre — a 300×12 matrix here. Those 12 distances are the landmark description.

The labels y never touched the clustering. K-means saw only the shape of X. The clusters carry predictive power because the two moons occupy different regions, so "near centre 3" correlates with class 0. Unlabelled structure became supervised signal.

Note what was fitted where. km was fitted on Xtr alone, then applied to both splits. Refitting K-means on the test rows — or on all rows — would let test-set geometry leak into the features. With 600 clean points the damage would be small; in production pipelines this exact mistake produces beautiful offline scores that vanish on live traffic.

12 is a tuning knob, not magic. More centres means a finer-grained description and more columns to overfit with. Treat n_clusters here like any hyperparameter: tune it with cross-validation on the training side. The choosing-k tools do not apply — you want predictive lift, not pretty clusters.

Common mistakes

Feeding the raw cluster id as a number. Cluster 7 is not "more" than cluster 2 — the ids are names. If you use hard membership, one-hot encode it. Distances avoid the issue entirely and carry more information.

Fitting the clusterer inside no pipeline, then cross-validating. Each CV fold must refit the clusterer on that fold's training part. Wrap KMeans and the classifier in a scikit-learn Pipeline and cross-validate the whole thing; hand-rolled loops almost always leak.

Skipping feature scaling. Distances again. If spend is in rupees and age in years, every centre distance is a spend measurement in disguise. Scale before clustering.

Expecting lift on every dataset. When the classes are already linearly separable, or the model is a gradient-boosted tree that carves regions natively, cluster features often add nothing. They shine for linear models and neural nets on lumpy, multi-modal data.

Try it yourself

Replace the distance features with a one-hot of km.predict(X) — hard membership instead of distances — and compare test accuracy. Then try n_clusters of 4 and 30. Distances versus membership at each size tells you how much of the lift comes from geometry versus region ids.

What to learn next

Researcher — Mathematics and papers.

Why distance features linearise lumpy problems

Appending distances $\lVert x - \mu_j \rVert$ for centres ${\mu_j}_{j=1}^{k}$ is an explicit feature map $\phi: \mathbb{R}^d \to \mathbb{R}^{d+k}$. A linear function over $\phi(x)$ can express thresholded combinations of distances to landmarks — piecewise-spherical structure the raw coordinates cannot express linearly. Replace distances with kernels $\exp(-\gamma \lVert x - \mu_j \rVert^2)$ and you have an RBF network with unsupervisedly-chosen centres (Broomhead and Lowe, 1988); the K-means variant is the classic Moody and Darken (1989) architecture.

The same construction underlies Nyström approximation: landmarks subsample the kernel matrix to give a low-rank feature map with error bounds governed by how well centres cover the data (Zhang, Tsang, Kwok, 2008 use K-means centres precisely because quantisation error controls approximation error).

Vector quantisation as representation learning

Coates and Ng (2012), Learning feature representations with K-means, showed single-layer K-means features with whitening rivalling early deep networks on CIFAR-10 — the strongest empirical statement of this lesson's thesis. Key ingredients transfer to tabular practice: whiten before clustering, use soft assignments (their "triangle" encoding: $\max(0, \bar{d} - d_j)$, mean distance minus each distance), and use many more centres than you would for descriptive clustering.

The modern descendant is VQ in deep architectures — VQ-VAE (van den Oord et al., 2017) learns the codebook jointly, but the interface is identical: represent a point by its relationship to learned prototypes.

Leakage, formalised

With clustering as a transform $T_{\hat{\theta}}$ whose parameters $\hat{\theta}$ are estimated from data, the estimand of interest is generalisation error with $T$ estimated inside the training procedure. Estimating $\hat{\theta}$ on $D_{\text{train}} \cup D_{\text{test}}$ biases error estimates optimistically, even though no labels are involved, because the feature map adapts to test-set geometry. Kaufman et al. (2012) taxonomise this as feature leakage; the empirical bias is small for stable transforms (PCA, K-means on large $n$) and grows with transform flexibility and small $n$. The clean rule — all estimation inside the CV fold — costs little and removes the question.

Practical notes from applied literature

  • Credit-scoring and churn pipelines have long used "segment id" features from K-means or RFM-style rules; lift concentrates where the outcome-feature relationship differs by segment — cluster features approximate interaction terms.
  • Distance-to-normal-clusters features connect directly to anomaly scoring; the feature and the detector are the same quantity.
  • For high-cardinality landmarks, product quantisation (Jégou et al., 2011) compresses distance computations — relevant when $k$ reaches thousands, as in retrieval.

What to learn next