Clustering in Depth

Gaussian mixture models

A Gaussian mixture model describes data as overlapping bumps and gives every point a membership probability for each group, instead of a single hard label.

Read these first

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A Gaussian mixture model treats your data as a blend of overlapping crowds.

It tells you how strongly each point belongs to each crowd, as a share rather than a verdict.

Picture the exit gate after an India–Australia match. Two crowds stream out and mingle: home fans and away fans. Someone in a plain white shirt walks past. You cannot be certain which crowd they came from. But you can be 60–40 about it.

K-means refuses to say 60–40. It draws a hard wall and files the white shirt on one side, with total false confidence.

Why this matters

Real groups overlap. Customers sit between "bargain hunter" and "brand loyal". A song sits between two moods. Forcing every borderline case into one box throws away the most useful information about it — the fact that it is borderline.

A Gaussian mixture model, or GMM, keeps that information. It describes each group as a smooth bump of likelihood — dense at the centre, fading at the edges. Every point then gets a share split across the bumps.

How it works

   crowd A          overlap          crowd B
   ......         ..::||::..         ......
 ..::||::..     here a point is    ..::||::..
.::||||||::.    maybe 60% A,      .::||||||::.
             ->   40% B         <-

Fitting a GMM means adjusting three things per group until the bumps explain the data well. Where each bump sits. How wide it spreads, and in which directions. And how big a share of everyone it holds.

The spread part is a quiet superpower. K-means implicitly assumes every group is a circle of the same size. A GMM lets each group be a stretched, tilted oval. Groups shaped like "income rises with age" stop getting shredded into arbitrary slices.

Because each point gets probabilities, you also get a warning system for free. A point scoring around half-and-half is the model saying: I am not sure about this one.

A real example you have seen

Voice assistants must decide whether a sound is speech or background noise. Sound sits on a smooth range, and the two categories overlap heavily. Mixtures of bumps over sound features were the backbone of speech systems for years. The reason is exactly this: "80% sure this is speech" beats a hard yes.

Remember this

  • A GMM models data as overlapping bumps, one per group.
  • Every point gets a probability per group, not a single stamp.
  • Groups can be stretched ovals of different sizes — a shape K-means cannot express.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; expect last-digit drift elsewhere.

Hard wall versus soft split

Two student batches with overlapping marks, and one student caught between them:

soft_vs_hard.py
import numpy as np
from sklearn.cluster import KMeans
from sklearn.mixture import GaussianMixture

rng = np.random.default_rng(0)
batch_a = rng.normal([55, 60], [8, 6], size=(100, 2))   # marks in two subjects
batch_b = rng.normal([70, 72], [7, 8], size=(100, 2))
X = np.vstack([batch_a, batch_b])

km = KMeans(n_clusters=2, n_init=10, random_state=0).fit(X)
gmm = GaussianMixture(n_components=2, random_state=0).fit(X)

student = np.array([[62.0, 66.0]])          # sits between the two groups
print("k-means centres:\n", np.round(km.cluster_centers_, 1))
print("k-means says: cluster", km.predict(student)[0], "- full stop, no doubt expressed")
print("GMM says:", np.round(gmm.predict_proba(student)[0], 2), "- probability per group")
Output
k-means centres:
 [[69.7 72. ]
 [53.7 60.1]]
k-means says: cluster 0 - full stop, no doubt expressed
GMM says: [0.41 0.59] - probability per group

The walkthrough

The student at (62, 66) is genuinely ambiguous — roughly midway between the centres. K-means files them into cluster 0, the higher-marks batch, with no hint of doubt. The GMM reports 41% versus 59%: a lean, honestly labelled as a lean.

predict_proba is the point of the model. gmm.predict(student) also exists and returns the hard argmax label, but if you only ever take the argmax, most of the machinery was wasted.

Component numbering is arbitrary, for GMM and K-means alike. Component 0 of one is not component 0 of the other, and refitting can swap them. Match groups by their learned centres, never by index. Comparing two groupings properly is its own lesson.

covariance_type controls the bump shapes. The default "full" gives every component its own tilted oval. "diag" forbids tilt, "spherical" forces circles, "tied" shares one shape across components. Fewer shape parameters resist overfitting on small data.

Common mistakes

Reading component probabilities as truth about the world. The 59% is a statement within the model's assumptions: bumps are Gaussian and there are exactly two. Change either assumption and the number moves.

Using covariance_type="full" with many features and little data. Each full oval needs many parameters; with 50 features that is over a thousand per component. Fits become singular or absurd. Drop to "diag", or reduce dimensions first.

Letting a component collapse onto a few points. A bump shrunk around near-duplicate points gives the model infinite-looking likelihood — the classic mixture failure. scikit-learn's reg_covar guards against it; if you see needle-thin components, raise it.

Choosing n_components by eyeballing. GMMs are likelihood models, so information criteria apply: fit several counts and compare gmm.bic(X). Lower is better. This is the likelihood-based answer to how many clusters.

Try it yourself

Move the student to (55, 60), then to (75, 75), printing predict_proba each time. Then refit with covariance_type="spherical" and check how the probabilities at (62, 66) change. Shape assumptions move real answers.

What to learn next

Researcher — Mathematics and papers.

The model

A GMM with $K$ components defines the density:

$$ p(x) = \sum_{k=1}^{K} \pi_k \, \mathcal{N}(x \mid \mu_k, \Sigma_k) $$

Where:

  • $\pi_k$ — mixing weights, $\pi_k \ge 0$, $\sum_k \pi_k = 1$: each component's share of the data.
  • $\mu_k \in \mathbb{R}^d$ — component means.
  • $\Sigma_k \in \mathbb{R}^{d \times d}$ — component covariance matrices: the size and tilt of each bump.
  • $\mathcal{N}(\cdot \mid \mu, \Sigma)$ — the multivariate normal density.

The membership probability (the responsibility) of component $k$ for point $x$ is the posterior:

$$ r_k(x) = \frac{\pi_k \, \mathcal{N}(x \mid \mu_k, \Sigma_k)}{\sum_{j} \pi_j \, \mathcal{N}(x \mid \mu_j, \Sigma_j)} $$

This is what predict_proba returns. Fitting maximises the log-likelihood $\sum_i \log p(x_i)$, which has no closed form — the standard optimiser is expectation-maximisation, covered next.

K-means as a degenerate GMM

Fix $\Sigma_k = \sigma^2 I$ for all components with equal weights, and let $\sigma^2 \to 0$. The responsibilities harden to indicators of the nearest mean, and the EM updates become exactly Lloyd's algorithm. K-means is the zero-variance, spherical, equal-weight corner of the GMM family — which explains both its speed and its blindness to shape.

Parameter count and covariance structure

Per component: $d$ for the mean, plus covariance parameters — $\frac{d(d+1)}{2}$ (full), $d$ (diag), or $1$ (spherical); tied shares one full matrix across components. Model selection via BIC $= p \ln n - 2 \ln \hat{L}$ (with $p$ total parameters, $n$ points, $\hat{L}$ the maximised likelihood) tends to choose both $K$ and the covariance structure sensibly; AIC penalises less and picks larger models.

The unbounded-likelihood pathology: letting $\det \Sigma_k \to 0$ around coincident points sends the likelihood to infinity, so the global MLE does not exist in the unconstrained problem (Day, 1969). Practical fixes are covariance floors (reg_covar), Bayesian priors, or constraints on eigenvalue ratios.

Identifiability and inference caveats

Mixtures are identifiable only up to label permutation, and finite mixtures of Gaussians are generically identifiable (Yakowitz and Spragins, 1968). Likelihood surfaces are multimodal; multiple EM restarts (n_init) are as necessary here as in K-means.

Beyond vanilla GMMs

  • Bayesian / variational GMMs (BayesianGaussianMixture): a Dirichlet process prior lets surplus components fade to near-zero weight — an alternative route to choosing $K$ (Blei and Jordan, 2006).
  • Mixtures of t-distributions (Peel and McLachlan, 2000): heavier tails resist outliers that drag Gaussian components.
  • Mixture-of-experts and deep variants: responsibilities produced by networks; and modern density estimation increasingly uses normalising flows where GMMs once served.
  • The standard reference: McLachlan and Peel, Finite Mixture Models (2000).

What to learn next