Clustering in Depth

Clustering mixed numeric and categorical data

Distance stops having one right answer when one column is rupees and another is a payment method — Gower distance and column weighting make mixed-type clustering honest.

Read these first

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Clustering needs a notion of "how far apart" two rows are, and that notion breaks the moment one column holds numbers and another holds categories.

Compare two people by height: 160 cm and 172 cm are 12 cm apart. Easy. Now compare them by mother tongue: how far is Telugu from Bengali? The question has no answer. Categories have no ruler — two values are either the same or different, and that is all.

Most clustering algorithms assume every column has a ruler. Real tables — spend in rupees, city, payment method, age — do not cooperate.

Why this matters

The common workaround is to turn categories into 0/1 columns and cluster as usual. It runs without error, which is the dangerous part. Every category difference now counts as the same fixed hop, while numeric columns count in their own units. Whichever side of the table happens to produce bigger numbers silently wins the whole clustering.

Nobody decided that. The units decided it.

How it works

The honest fix is to score each column in its own fair way, then average. This recipe is called Gower distance:

column type      how far apart are two rows on this column?

numeric      ->  gap between values, rescaled so the biggest
                 possible gap counts as 1
category     ->  same value: 0        different value: 1

overall      ->  average the per-column scores (weights allowed)

Every column now speaks in the same range, from 0 to 1. A rupee gap and a payment-method difference become comparable — by construction, not by accident.

One decision remains, and no formula makes it for you: the weights. Should paying by card versus UPI matter as much as a ₹4,000 spend gap? That is a question about your business, and the weights are where you answer it.

A real example you have seen

Delivery apps group customers for offers using exactly such tables: order value (number), veg or non-veg preference (category), city (category), orders per month (number). Any segmentation over that table had to solve this problem before clustering could start.

Remember this

  • Categories have no ruler — only same or different.
  • Gower distance scores every column between 0 and 1, then averages, so no column wins by units.
  • Weights are a business decision. No algorithm can choose them for you.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 and numpy 1.26.4 on CPU.

Gower distance by hand, then clustering on it

Six customers, one numeric column and one categorical column. Watch the weights change the answer:

gower_mixed.py
import numpy as np
from sklearn.cluster import AgglomerativeClustering

# Six customers: monthly spend in rupees, and how they usually pay.
spend = np.array([500.0, 550.0, 620.0, 4800.0, 5100.0, 4900.0])
pay = np.array(["upi", "upi", "card", "card", "card", "upi"])

# Gower distance, by hand. Numeric column: range-scaled gap, between 0 and 1.
num = np.abs(spend[:, None] - spend[None, :]) / (spend.max() - spend.min())
# Categorical column: 0 if same, 1 if different.
cat = (pay[:, None] != pay[None, :]).astype(float)

equal = (num + cat) / 2                    # each column gets an equal vote
agg = AgglomerativeClustering(n_clusters=2, metric="precomputed", linkage="average")
print("equal weights: ", agg.fit_predict(equal))

spendy = 0.8 * num + 0.2 * cat             # spend gets four times the say
print("spend-weighted:", agg.fit_predict(spendy))
Output
equal weights:  [0 0 1 1 1 0]
spend-weighted: [1 1 1 0 0 0]

The walkthrough

The two lines disagree, and both are correct. With equal weights, the clusters are UPI users versus card users — read the labels against the pay array. With spend weighted up, the clusters become small spenders versus big spenders. Same six customers, two defensible segmentations. The weights are the analysis.

Why did the category win under equal weights? After range-scaling, the biggest spend gap counts as 1.0, so the ₹50 gap between neighbours counts as about 0.01. A category mismatch counts as a full 1.0 every time. One loud binary column outvotes a numeric column whose values huddle together — a thing to check, not to fear.

metric="precomputed" lets AgglomerativeClustering accept our distance matrix directly. Most scikit-learn clusterers compute Euclidean distance internally; hierarchical clustering is the one that happily takes yours. K-means cannot — its centre-averaging step has no meaning for "the average of UPI and card".

Building the matrix with broadcasting — spend[:, None] - spend[None, :] — gives the full 6×6 table in one line. For $n$ rows the matrix holds $n^2$ entries — about 0.8 GB at 10,000 rows in float64, 3.2 GB at 20,000, and beyond most machines past roughly 30,000.

Common mistakes

One-hot encoding into K-means without thinking. It runs, and the arbitrary scale ratio between dummies and numerics decides your segments. If you must go this route, scale deliberately and read the result with suspicion.

Range-scaling with outliers present. One customer with a ₹9,00,000 spend stretches the range so every normal gap rounds to zero. Cap or clean outliers before computing Gower.

Ordinal categories treated as plain categories. Sizes S / M / L have an order — same-or-different throws that away. Map ordered categories to ranks and treat them as numeric instead.

Forgetting k-prototypes exists. For large mixed tables, the kmodes package's k-prototypes algorithm clusters without a full distance matrix — numeric means plus category modes as centres. Worth knowing when $n^2$ memory is the blocker.

Try it yourself

Add a seventh customer: spend 5,000, pays by UPI. Recompute both matrices and predictions. Under which weighting does this customer switch groups? Explain the switch from the two distance tables before running the code.

What to learn next

Researcher — Mathematics and papers.

Gower's coefficient, formally

Gower (1971), A general coefficient of similarity and some of its properties, defines similarity between records $i$ and $j$ as:

$$ S_{ij} = \frac{\sum_{f} w_f \, \delta_{ijf} \, s_{ijf}}{\sum_{f} w_f \, \delta_{ijf}} $$

Where:

  • $s_{ijf}$ — the per-feature similarity: $1 - |x_{if} - x_{jf}| / R_f$ for numeric features with range $R_f$; the indicator $\mathbb{1}[x_{if} = x_{jf}]$ for categorical.
  • $\delta_{ijf}$ — 1 when feature $f$ is comparable for the pair (both values present), 0 otherwise — Gower handles missing data natively, a routinely forgotten strength.
  • $w_f$ — feature weights.

Distance is taken as $d_{ij} = 1 - S_{ij}$, matching the developer block's average. $\sqrt{1 - S_{ij}}$ is Euclidean-embeddable when the similarity matrix is positive semi-definite, which Gower proves for his coefficient — relevant when a method requires metric embedding.

k-prototypes

Huang (1998) extends K-means to mixed data. The cost for cluster $l$ with centre $(\mu_l, m_l)$:

$$ \sum_{i \in C_l} \Big( \sum_{f \in \text{num}} (x_{if} - \mu_{lf})^2 + \gamma \sum_{f \in \text{cat}} \mathbb{1}[x_{if} \ne m_{lf}] \Big) $$

Numeric parts update as means; categorical parts as per-cluster modes; $\gamma$ trades the two — the same weighting decision, wearing an algorithmic costume. Complexity per iteration stays $O(nkd)$ and needs no $n^2$ matrix. Its pure-categorical special case is k-modes (Huang, 1997).

Alternatives and their trade-offs

  • Embed, then cluster: learn dense representations of categories (target encoding, entity embeddings — Guo and Berkhahn, 2016) and cluster in the embedded space. Powerful with many high-cardinality categories; the metric becomes opaque.
  • Model-based: latent class analysis and mixtures with per-type likelihoods (Gaussian × multinomial), fitted by EM. Principled uncertainty; strong independence assumptions.
  • Tree-derived distances: random-forest proximities handle types natively and capture interactions, at the price of a supervised or artificial target.
  • Hierarchical clustering over a Gower matrix remains the small-data workhorse precisely because every modelling choice stays visible.

The unavoidable subjectivity

Hennig and Liao (2013), How to find an appropriate clustering for mixed-type variables with application to socio-economic stratification, is the essential read. Its argument: no automatic weighting is neutral, because the analyst's concept of similarity is a modelling decision that precedes the mathematics. Equal weights feel objective and are still a choice.

What to learn next