Manufacturing and Predictive Maintenance

Building a model from four failures

When a machine has failed only a handful of times in its whole recorded history, no model's accuracy number deserves to be trusted, and the honest response is to change the question being asked.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When a machine has only failed four times in its entire recorded history, no model's accuracy score deserves to be trusted. The honest response is to ask a different question.

Think about a doctor who has only ever seen four patients with a specific rare disease in their whole career. No matter how careful or experienced that doctor is, they cannot claim to reliably recognise every version of that disease. Four cases is not enough to have seen the full range of how it can show up. A responsible doctor says so plainly, rather than projecting false confidence.

The same honesty is needed for a machine that has genuinely, rarely failed.

Why it exists

The earlier lessons in this section used made-up examples with plenty of failures to learn from. Real critical equipment — a large industrial compressor, a specialised furnace, a well-maintained turbine — often fails only a handful of times across years of operation. That is a good thing for the business. It is a very hard starting point for machine learning.

With only four known failures, a model trained the ordinary way is really being evaluated on four individual events, not a statistically meaningful sample. Two of those four might look sharply abnormal and get caught reliably. The other two might look subtle, barely different from normal operation, and get missed every time. Averaging that into one tidy "50% recall" number hides how little that average actually tells you about the next failure.

How it works

4 known failures, checked one at a time:

  Failure 1 (subtle):    MISSED
  Failure 2 (subtle):    MISSED
  Failure 3 (obvious):   CAUGHT
  Failure 4 (obvious):   CAUGHT

Reported recall: 2 out of 4 = 50%

But with only 4 examples, the TRUE recall could honestly be
anywhere from about 7% to 93% -- the data cannot say more than that.

Reporting "50% recall" without that context implies a precision the data does not support. The honest version of this result says: two failures stood out sharply and were caught, and two looked subtle and were missed. There is not enough history yet to know how the next failure will look.

Where you have already seen it

  • A rare medical condition with few documented cases, where doctors explicitly rely on published case studies rather than claiming statistical confidence a small sample cannot provide.
  • An airline investigating a rare type of mechanical failure. Regulators demand root-cause engineering analysis of each individual incident, rather than trusting a statistical average built from a handful of events.
  • A cricket team analysing a bowler's rare, unusual delivery. It has only happened a few times on camera, so commentators are careful to say "small sample," rather than treating four instances as a proven pattern.

An honest warning

This is exactly the situation where a confident-looking model is more dangerous than an honestly uncertain one. A model reporting "94% accuracy" on this kind of data is almost always measuring something else entirely. It shows how well it predicts the overwhelming majority of normal days, not how well it catches the rare failures that actually matter. For any machine where a missed failure has real safety consequences, this kind of model should support a qualified engineer's judgement, never replace it.

Remember this

  • With only a handful of known failures, no accuracy or recall number is statistically reliable — say so plainly rather than hiding it.
  • A model trained on very few positive examples has effectively memorised those specific cases, not learned the general pattern of failure.
  • The honest response is often to change the question, from "predict failure" to "rank how unusual this looks," which needs no failure labels at all.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas scikit-learn statsmodels

Minimal runnable code

400 days of normal readings and only 4 known failures — two sharply abnormal, two subtle. We evaluate a classifier the honest way: one held-out failure at a time, with a confidence interval on the result. Then we try an anomaly-detection framing that needs no failure labels at all.

rare_failure_data.py
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier, IsolationForest
from statsmodels.stats.proportion import proportion_confint

rng = np.random.default_rng(14)

# 400 days of normal operation, and only 4 recorded failures -- two sharply
# abnormal, two subtle -- a realistic count for one well-maintained machine.
n_normal = 400
normal = pd.DataFrame({
    "temperature": rng.normal(70, 3, n_normal),
    "current_draw": rng.normal(12, 1, n_normal),
})
failures = pd.DataFrame({
    "temperature":  [75.0, 76.0, 90.0, 92.0],
    "current_draw": [13.2, 13.5, 17.5, 19.0],
})
data = pd.concat([normal.assign(label=0), failures.assign(label=1)], ignore_index=True)
X = data[["temperature", "current_draw"]].to_numpy()
y = data["label"].to_numpy()
failure_idx = np.where(y == 1)[0]

print(f"total days: {len(data)}, failures: {y.sum()}")
print()

# Leave-one-failure-out: hold out ONE known failure at a time, train on
# everything else (all normals plus the other 3 failures), and check if
# the model catches the one held out. With only 4 failures, this is the
# entire evaluation -- there is no "average over many test failures."
results = []
for i in failure_idx:
    test_mask = np.zeros(len(y), dtype=bool)
    test_mask[i] = True
    model = RandomForestClassifier(n_estimators=200, class_weight="balanced", random_state=0)
    model.fit(X[~test_mask], y[~test_mask])
    caught = model.predict(X[test_mask])[0] == 1
    results.append(caught)
    print(f"held-out failure at temp={X[i,0]:.0f}: {'CAUGHT' if caught else 'MISSED'}")

recall = np.mean(results)
print()
print(f"observed recall: {sum(results)}/4 = {recall:.0%}")

low, high = proportion_confint(sum(results), 4, method="beta")
print(f"95% confidence interval for the TRUE recall: [{low:.0%}, {high:.0%}]")
print("(this interval is almost the entire range -- four examples cannot pin it down)")
print()

# An anomaly-detection framing instead: score every day by how unusual it
# is, using no failure labels at all, and see where the known failures rank.
iso = IsolationForest(contamination=0.01, random_state=0).fit(normal[["temperature", "current_draw"]])
anomaly_score = -iso.score_samples(data[["temperature", "current_draw"]])
data["anomaly_score"] = anomaly_score
data["anomaly_rank"] = data["anomaly_score"].rank(ascending=False)

print("where each known failure ranks by unusualness, out of 404 days (1 = most unusual):")
print(data[data["label"] == 1][["temperature", "current_draw", "anomaly_rank"]].round(1))
Output
total days: 404, failures: 4

held-out failure at temp=75: MISSED
held-out failure at temp=76: MISSED
held-out failure at temp=90: CAUGHT
held-out failure at temp=92: CAUGHT

observed recall: 2/4 = 50%
95% confidence interval for the TRUE recall: [7%, 93%]
(this interval is almost the entire range -- four examples cannot pin it down)

where each known failure ranks by unusualness, out of 404 days (1 = most unusual):
     temperature  current_draw  anomaly_rank
400         75.0          13.2          52.0
401         76.0          13.5          21.0
402         90.0          17.5           1.5
403         92.0          19.0           1.5

What actually happened

Leave-one-failure-out trains on everything except one specific failure, then checks whether that one failure gets caught. Doing this for all four gives the entire, honest picture: the two sharply abnormal failures (temperature 90 and 92) were caught every time; the two subtle ones (75 and 76, barely above the normal range) were missed every time.

proportion_confint computes a formal confidence interval for the true recall, given only 2 successes out of 4 trials. The result, roughly 7% to 93%, spans almost the entire possible range. That is not a coding mistake — it is an accurate reflection of how little four data points can tell you. Reporting "50% recall" without this interval would be technically true and practically misleading.

The IsolationForest step asks a different question entirely: not "is this a failure," but "how unusual is this compared to normal operation," using zero failure labels to train on. The two obvious failures rank as the two most unusual days out of 404. The two subtle failures rank 21st and 52nd most unusual — not flagged by a strict top-few alert, but ranked meaningfully higher than an ordinary day, which is genuinely useful information for an engineer prioritising which readings to review.

Common mistakes

Reporting a recall or accuracy score from a handful of positive examples without a confidence interval. The number alone invites false confidence. Always show how wide the honest uncertainty is when the positive class count is this small.

Treating the anomaly detector's ranking as a replacement for engineering judgement. A high anomaly rank means "worth a look," not "confirmed failure." It is a prioritisation tool, not a verdict.

Waiting for more failures to accumulate before doing anything. For genuinely rare, expensive failures, that wait can be years. The anomaly-detection framing exists precisely so useful monitoring can start immediately, without needing many failure examples first.

Ignoring fleet data when it exists. If the plant has thirty near-identical pumps, and one specific pump has only failed four times, the other twenty-nine pumps' complete operating history — including their own failures — is directly relevant data being wasted if only the one pump's history is used.

Try it yourself

Change contamination=0.01 to contamination=0.1 in the IsolationForest. Re-run, and check how the anomaly ranks and the implied alert threshold shift — this parameter is effectively a knob for how many "worth reviewing" alerts the system generates per day.

What to learn next

Researcher — Mathematics and papers.

Why standard supervised learning theory does not apply here

The generalisation bounds covered in What is machine learning? assume a sample large enough for the law of large numbers to apply to the minority class specifically, not only the dataset as a whole. With n_pos = 4, no consistent estimator of recall, precision, or any other minority-class statistic has meaningfully shrinking variance — the sampling distribution of the estimate stays wide regardless of how the majority class is sampled. The confidence interval computed above, via the exact Clopper-Pearson method, makes this explicit rather than letting a point estimate imply false precision.

One-class and semi-supervised framings

Framing this as anomaly detection, as in the developer example, is formally one-class classification: learn a model of the normal class's support and flag deviations, without requiring examples of every possible failure mode. IsolationForest (Liu, Ting and Zhou, 2008, Isolation Forest) isolates points via random recursive partitioning, scoring anomalies by how few splits are needed to isolate them — anomalies, being sparse and different, isolate faster than dense normal clusters. One-Class SVM (Schölkopf et al., 2001) and autoencoder reconstruction error are common alternatives, covered in One-class SVM and Anomaly detection by reconstruction error.

A middle ground, when a handful of labelled failures do exist, is semi-supervised anomaly detection: train primarily on the normal class, but use the few known failures purely for threshold selection and validation — as done above by checking where they rank — rather than as training examples a classifier tries to generalise from directly.

Extreme value theory as a principled alternative

Where failure is defined by a sensor value crossing into its extreme tail (a temperature or pressure spike, rather than a qualitatively different pattern), extreme value theory (EVT) offers an approach grounded in different statistical machinery than either classification or generic anomaly scoring. The peaks-over-threshold method fits a generalised Pareto distribution to exceedances above a high threshold u:

P(X - u > x | X > u) approx (1 + xi*x/sigma)^(-1/xi)
  • u — a chosen high threshold
  • xi — the shape parameter, governing tail heaviness
  • sigma — the scale parameter

This lets a model estimate the probability of an even more extreme event than any yet observed, from the shape of the tail already seen — directly relevant when the "failure" event of interest is rarer than anything in the historical record. Coles (2001), An Introduction to Statistical Modeling of Extreme Values, is the standard reference.

Fleet pooling and hierarchical models

Where genuinely identical or near-identical units exist across a plant, pooling their histories using a hierarchical (multilevel) model — a unit-specific failure rate drawn from a shared population-level distribution, updated per unit as more of that specific unit's own data arrives — directly formalises the fleet-pooling suggestion in the developer section, using the same shrinkage logic discussed for new-product forecasting in Forecasting a product with no history. Gelman and Hill (2006), Data Analysis Using Regression and Multilevel/Hierarchical Models, is the standard reference for this class of model.

Key references

  • Liu, F. T., Ting, K. M. & Zhou, Z.-H. (2008). Isolation Forest. IEEE International Conference on Data Mining.
  • Coles, S. (2001). An Introduction to Statistical Modeling of Extreme Values. Springer.
  • Clopper, C. & Pearson, E. (1934). The Use of Confidence or Fiducial Limits Illustrated in the Case of the Binomial. Biometrika 26(4).
  • Gelman, A. & Hill, J. (2006). Data Analysis Using Regression and Multilevel/Hierarchical Models. Cambridge University Press.

What to learn next