One-class SVM
A one-class SVM learns a single flexible fence around your normal data, with a dial that states upfront how many training points you are willing to leave outside.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A one-class SVM draws one fence around your normal data, as tight as it can manage.
A few good points are allowed to fall outside. Everything beyond the fence is flagged.
A shepherd pens his flock for the night. Build the fence too generous and wolves wander inside the pen with room to spare. Build it too tight and half the sheep are locked out bleating. The craft is a snug fence: hugging the flock's true shape, tolerating a few stragglers outside.
One-class SVM is that fence, drawn in data instead of grass.
Why this matters
This is a tool for the novelty mode of anomaly detection. You hold a trusted set of normal examples. Future arrivals are judged against them. You have no examples of the bad thing — wolves are rare, varied, and unhelpful enough to never pose for training photos.
Ordinary classifiers need both classes. A one-class method learns the shape of a single class and answers only: inside or outside?
The fence is not a circle. Real normal data forms bananas, ridges and blobs. The method uses a kernel. That is a trick for stretching the space, so a straight fence in the stretched version becomes a curved, flock-hugging fence in the real one.
How it works
___----___
/ x x \
| x x x x | x = normal training data
\ x x _ / o = new arrival, inside -> ok
\_ x _/ x <- a few stragglers left out, on purpose
\__/
O <- new arrival, outside -> alarm!The deliberate stragglers are the honest part. You set one dial before training, called nu: the share of training sheep you accept leaving outside. Set it near zero and the fence stretches to swallow every last point — including mistakes hiding in your "clean" data. Set it at five percent and the fence stays snug, sacrificing the strangest five percent of training points to hold its shape.
That dial is a promise you make about your own data's dirtiness.
A real example you have seen
Machine-health monitoring in factories. Months of vibration and temperature readings from a healthy motor exist; recordings of every possible failure do not. A fence around "healthy" flags any departure — including failure types nobody predicted. The same one-sided logic guards server metrics and medical readings.
Remember this
- Learns from normal data only; answers inside-or-outside for new points.
- nu is your stated tolerance: the share of training points allowed outside the fence.
- The kernel lets the fence hug curved shapes, not only circles.
What to learn next
- Anomaly detection by reconstruction error — the deep learning route to one-class thinking.
- Choosing a threshold from costs — turning fence distances into operational decisions.
- Outliers vs novelties — the mode distinction this tool depends on.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
Fencing a healthy machine
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.svm import OneClassSVM
rng = np.random.default_rng(0)
healthy = rng.normal([60, 40], [8, 5], size=(300, 2)) # temperature, vibration
scaler = StandardScaler().fit(healthy)
fence = OneClassSVM(nu=0.05).fit(scaler.transform(healthy))
new = np.array([[62.0, 41.0], [95.0, 70.0], [60.0, 15.0]])
print("verdicts on new readings:", fence.predict(scaler.transform(new)))
inside = fence.predict(scaler.transform(healthy))
print("training points left outside:", int((inside == -1).sum()), "of 300")verdicts on new readings: [ 1 -1 -1] training points left outside: 21 of 300
The walkthrough
The three verdicts, in order. (62, 41) sits amid the training cloud: normal. (95, 70) is far outside on both axes: flagged. (60, 15) has a perfectly healthy temperature with impossibly low vibration — anomalous in combination: also flagged. Per-column rules from the statistical lesson would clear that third reading; a fence in the full 2-D space does not.
The nu promise held: 21 of 300 is 7%, near the requested 5%. The mathematics guarantees the training-set violation fraction lands close to nu. This is the parameter's real meaning — a stated false-alarm budget on data you trust — and small-sample wobble around it is expected.
Scaling is not optional. The kernel measures distances; unscaled axes distort the fence exactly as they distort K-means and LOF. Note the scaler is fitted on the healthy data alone, then applied to newcomers.
The default kernel is RBF, with a width setting gamma. gamma="scale" (the default) works from the data's variance and is the sensible start. Big gamma wraps the fence tightly around individual points; small gamma inflates it toward an ellipse.
Common mistakes
Tuning gamma by eye on training data alone. With a big enough gamma, the fence hugs every training point perfectly — and rejects everything new. This is overfitting in its purest costume. Validate on held-out normal data: a good fence should admit roughly 1 - nu of it.
Feeding contaminated "normal" data. Every hidden anomaly in training either stretches the fence to cover it or spends part of the nu budget. A few percent contamination degrades it gracefully; more, and the fence learns the wrong flock.
Expecting calibrated scores. decision_function gives a signed distance from the fence — useful for ranking, meaningless as a probability. If someone downstream needs "94% sure", that is a calibration problem.
Using it beyond ~10,000 training points without thought. Training cost grows steeply with sample count. SGDOneClassSVM in scikit-learn gives a linear-time approximation; isolation forest is the usual pragmatic swap.
Try it yourself
Sweep nu over 0.01, 0.05, 0.2 and record both numbers: training points outside, and the verdict on (60, 15). Then set gamma=5.0 and watch the fence turn paranoid about all three newcomers. Two dials, and you have felt both.
What to learn next
- Anomaly detection by reconstruction error — the deep learning route to one-class thinking.
- Choosing a threshold from costs — turning fence distances into operational decisions.
- Outliers vs novelties — the mode distinction this tool depends on.
Researcher — Mathematics and papers.
The optimisation problem
Schölkopf, Platt, Shawe-Taylor, Smola and Williamson (2001), Estimating the support of a high-dimensional distribution. Map data by $\phi$ into the kernel's feature space and separate it from the origin with maximum margin:
$$ \min_{w, \rho, \xi} \;\; \frac{1}{2}\lVert w \rVert^2 + \frac{1}{\nu n} \sum_{i=1}^{n} \xi_i - \rho \quad \text{s.t.} \quad \langle w, \phi(x_i) \rangle \ge \rho - \xi_i, \;\; \xi_i \ge 0 $$
Where:
- $w, \rho$ — the separating hyperplane in feature space; the decision function is $\mathrm{sign}(\langle w, \phi(x)\rangle - \rho)$.
- $\xi_i$ — slack variables: how far each training point may fall on the wrong side.
- $\nu \in (0, 1]$ — the regularisation dial.
- $n$ — training count.
The $\nu$-property (Proposition 3 of the paper): $\nu$ is an upper bound on the fraction of training outliers and a lower bound on the fraction of support vectors, with equality asymptotically under mild conditions. This is the "21 of 300 ≈ 5%" guarantee, proved rather than hoped.
SVDD and the equivalence
Tax and Duin (2004), Support vector data description, enclose the data in a minimal feature-space ball instead of separating from the origin. For translation-invariant kernels (RBF included), SVDD and the one-class SVM give identical solutions — the ball and the hyperplane are the same fence in different coordinates. For non-stationary kernels they differ, and SVDD's geometry is usually the better mental model.
Kernel choice as density-level-set estimation
With the RBF kernel, the one-class SVM estimates a level set of the underlying density — the connection to minimum-volume sets (Scott and Nowak, 2006) made in the first lesson. Vert and Vert (2006) prove consistency of the RBF one-class SVM for density-level-set estimation with suitable bandwidth decay. In practice bandwidth ($\gamma$) selection remains the sore point; heuristics based on median pairwise distances predate gamma="scale" and remain competitive.
Scaling and the deep successor
Training solves a QP: $O(n^2)$–$O(n^3)$ depending on solver and cache; prediction is $O(n_{sv} \cdot d)$ per point with $n_{sv}$ support vectors. Linear-time routes: SGDOneClassSVM (an online approximation over kernel approximation features such as Nyström or random Fourier features) and random-feature one-class SVMs.
Deep SVDD (Ruff et al., 2018) replaces $\phi$ with a learned network that contracts normal data around a centre — the SVDD objective with a trainable kernel. Its collapse pathology (network maps everything to the centre) and fixes are the entry point into the deep one-class literature, surveyed in Ruff et al. (2021). For tabular work, the empirical message from ADBench (Han et al., 2022) is humbler: the classical one-class SVM stays competitive when features are informative and data fits in memory.
What to learn next
- Anomaly detection by reconstruction error — the deep learning route to one-class thinking.
- Choosing a threshold from costs — turning fence distances into operational decisions.
- Outliers vs novelties — the mode distinction this tool depends on.