Manufacturing and Predictive Maintenance

Control charts versus machine learning

Statistical control charts have monitored factories for a century and remain the right tool for many jobs, but they are slow to notice a small, sustained drift that a purpose-built chart or model can catch far sooner.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Statistical control charts have monitored factories for a century, and they remain the right tool for most everyday monitoring. But they are slow to notice a small, sustained drift.

Think about checking your own body temperature. You know roughly what "normal" feels like — around 37°C, give or take a bit through the day. If it suddenly spikes to 39°C, you notice immediately; that is a big, obvious jump. But imagine your temperature crept up by only half a degree, and quietly stayed there for a week. You might not notice at all. A small, sustained change is far harder to catch than a big sudden one — even though it can matter every bit as much.

Factories have used a formal version of "knowing what normal feels like" for a hundred years, and it has exactly this same blind spot.

Why it exists

A control chart plots a process measurement over time against fixed control limits. Those limits usually sit three standard deviations above and below the normal average. They are set using data collected while the process was known to be running well. A point outside those limits signals something has genuinely changed, worth investigating. This method, invented by Walter Shewhart in the 1920s, is simple, well understood, and still runs on production lines everywhere today.

It has one well-known weakness. It is built to catch a single point standing out sharply, and it is genuinely slow to notice a small, steady drift. That kind of drift never crosses the line on any single reading, even as the process quietly moves away from where it should be. A tool wearing down gradually, a sensor calibration slipping bit by bit — these can hide from a standard control chart for a long time.

This is not a reason to abandon control charts. It is a reason to know when a different tool is called for.

How it works

Normal process readings, centred around 50, control limits at 44.5 and 56.2

Then a small, sustained shift begins: the true average quietly moves to 51
  -- well within the control limits on every single reading

A basic control chart:   never notices, because no single point crosses 44.5 or 56.2

A chart built to notice ACCUMULATING small changes ("CUSUM"):
  adds up the small deviations over time, and catches the drift once
  enough small deviations have accumulated to be real evidence

The fix is not a more complicated model. It is a chart specifically designed to accumulate small, persistent evidence, instead of only reacting to one extreme point at a time.

Where you have already seen it

  • A bathroom scale you check every morning. A single unusual reading might only be noise. A whole month trending in one direction is real, even if no single day looked alarming.
  • A cricket team's batting average over a season. It reveals a real slump far more reliably than any single low score, which could be nothing more than a bad day.
  • A savings account balance drifting down slowly, unnoticed month to month, versus a single large, obvious unauthorised transaction that would be spotted immediately.

Remember this

  • Classic control charts catch large, sudden changes well, and remain a simple, well-proven, widely used tool.
  • They are genuinely slow to catch a small change that persists over time, since no single reading crosses the limit.
  • A chart built to accumulate small deviations, like CUSUM, catches exactly this kind of slow drift that a basic control chart misses.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

A process runs normally for 100 readings, then its true average quietly shifts by only half a standard deviation and stays there. We compare a classic Shewhart control chart against a CUSUM chart on the same data.

control_charts.py
import numpy as np

rng = np.random.default_rng(25)

# A process running in control for 100 readings, mean 50, std 2.
n_in_control = 100
in_control = rng.normal(50, 2, n_in_control)

# Then a SMALL, sustained shift begins: the true mean quietly creeps to
# 51 (only half a standard deviation) and stays there -- a slow drift,
# like a tool wearing down gradually, not a sudden breakage.
n_shifted = 60
shifted = rng.normal(51, 2, n_shifted)

process = np.concatenate([in_control, shifted])
shift_starts_at = n_in_control

# Classic Shewhart control limits: centre line +/- 3 standard deviations,
# estimated from the known in-control period.
center = in_control.mean()
sigma = in_control.std()
ucl, lcl = center + 3 * sigma, center - 3 * sigma
print(f"control limits: {lcl:.2f} to {ucl:.2f} (centre {center:.2f})")

shewhart_flags = np.where((process > ucl) | (process < lcl))[0]
shewhart_after_shift = shewhart_flags[shewhart_flags >= shift_starts_at]
print(f"Shewhart chart: points outside limits after the shift began: {len(shewhart_after_shift)} / {n_shifted}")
if len(shewhart_after_shift):
    print(f"  first flagged {shewhart_after_shift[0] - shift_starts_at} readings after the shift started")
else:
    print("  never flagged the shift at all, in this window")
print()


def cusum(data, target, k, h):
    """Cumulative sum control chart: accumulates small, persistent deviations."""
    s_high = 0.0
    s_low = 0.0
    alarms = []
    for i, x in enumerate(data):
        s_high = max(0, s_high + (x - target) - k)
        s_low = min(0, s_low + (x - target) + k)
        if s_high > h or s_low < -h:
            alarms.append(i)
            s_high, s_low = 0.0, 0.0   # reset after an alarm, as standard practice
    return alarms


# k is set to half the shift we want to be sensitive to; h controls how
# much accumulated evidence is required before raising an alarm.
alarms = cusum(process, target=center, k=sigma / 2, h=4 * sigma)
alarms_after_shift = [a for a in alarms if a >= shift_starts_at]
print(f"CUSUM chart: alarms raised after the shift began: {len(alarms_after_shift)}")
if alarms_after_shift:
    print(f"  first alarm {alarms_after_shift[0] - shift_starts_at} readings after the shift started")
Output
control limits: 44.50 to 56.19 (centre 50.35)
Shewhart chart: points outside limits after the shift began: 0 / 60
  never flagged the shift at all, in this window

CUSUM chart: alarms raised after the shift began: 1
  first alarm 37 readings after the shift started

What actually happened

The Shewhart chart's control limits, 44.50 to 56.19, comfortably contain the shifted process too — an average of 51 with a standard deviation of 2 rarely produces a single reading anywhere near those limits. It genuinely never flagged the drift across all 60 shifted readings.

cusum tracks two running totals: s_high, accumulating evidence of a persistent upward drift, and s_low, doing the same downward. Each reading only nudges the total a little (k acts as a tolerance, ignoring noise smaller than that), but a genuine sustained shift keeps pushing the same direction, and the total eventually crosses the alarm threshold h. Here, it caught the drift 37 readings after it began — far from instant, but a real detection where the Shewhart chart found nothing at all.

Common mistakes

Assuming a control chart with no flagged points means the process is healthy. As shown here, it can mean the process has a real, sustained problem that is too small for that specific chart to catch.

Setting CUSUM's k and h without a principled choice. k should reflect the smallest shift genuinely worth catching — set too small, CUSUM fires constantly on ordinary noise; too large, it becomes as blind as a Shewhart chart.

Treating this as "old statistics versus new machine learning." CUSUM is itself decades-old statistics, not machine learning, and it solves this specific problem more directly and more cheaply than most ML anomaly detectors would. Reach for the simplest tool that solves the actual problem.

Forgetting that both charts assume the "in control" baseline is genuinely representative. If in_control itself was collected during an unusual period, both the control limits and the CUSUM target inherit that mistake.

Try it yourself

Change the shift size from 51 to 54 — a much larger, more sudden change. Re-run and watch the Shewhart chart start flagging points almost immediately, while CUSUM also reacts, likely faster than before. Big shifts are exactly what Shewhart charts were designed for.

What to learn next

Researcher — Mathematics and papers.

Shewhart charts and the average run length

A Shewhart control chart's key performance measure is the average run length (ARL): the expected number of samples before a false alarm occurs (ARL when the process is in control, ideally very large) versus the expected number of samples to detect a real shift of a given size (ARL when the process is out of control, ideally very small). For a chart with 3-sigma limits monitoring a shift of size delta (in standard deviation units), the out-of-control ARL for an independent, normally distributed process is:

ARL = 1 / (1 - Phi(3 - delta) + Phi(-3 - delta))
  • Phi — the standard normal cumulative distribution function
  • delta — the shift size, in units of the process standard deviation

At delta = 0.5 (the shift size used in the developer example), the theoretical ARL is roughly 155 samples — meaning a Shewhart chart is expected to need well over a hundred readings, on average, to detect a shift this small, consistent with the empirical result above (though any single run, as demonstrated, can differ substantially from the average).

CUSUM's formal derivation and optimality

CUSUM (Page, 1954, Continuous Inspection Schemes, Biometrika) is not an ad hoc accumulator — it is derived from the sequential probability ratio test (Wald, 1945), and is provably optimal in a specific sense: among all detection procedures achieving a given in-control ARL, CUSUM minimises the out-of-control ARL for the specific shift size it is tuned to detect (Moustakides, 1986, Optimal Stopping Times for Detecting Changes in Distributions, Annals of Statistics). This optimality is conditional on k being set correctly for the target shift — a CUSUM tuned for a large shift performs no better than Shewhart against a small one, and vice versa.

The exponentially weighted moving average (EWMA) chart offers similar small-shift sensitivity to CUSUM through a different mechanism — a smoothed running average with exponentially decaying weight on older observations:

z_t = lambda * x_t + (1 - lambda) * z_{t-1}
  • lambda — the smoothing weight, 0 < lambda <= 1; smaller values weight history more heavily and detect smaller shifts more sensitively, at the cost of slower response and more complex control-limit calculation

EWMA and CUSUM detect small sustained shifts with broadly comparable statistical efficiency (Lucas and Saccucci, 1990, Exponentially Weighted Moving Average Control Schemes, Technometrics); the choice between them in practice is often driven by implementation convention and ease of tuning rather than a meaningful performance gap.

Where multivariate SPC and machine learning genuinely add value

Both charts above monitor one variable at a time. Real processes have dozens of correlated sensors, and a shift that is invisible in any single variable can be glaringly obvious in their joint relationship — this is exactly Hotelling's T-squared statistic and multivariate SPC's domain, and it is also where nonlinear machine-learning anomaly detectors (autoencoders, isolation forests) earn their added complexity over classical charts: not by replacing the single-variable case covered here, but by handling correlation structure across many variables that univariate control charts cannot see by construction. The honest comparison is rarely "control charts versus ML" as competitors — it is choosing the simplest tool that matches the actual structure of the problem, univariate drift versus multivariate pattern change.

Key references

  • Shewhart, W. (1931). Economic Control of Quality of Manufactured Product. Van Nostrand.
  • Page, E. S. (1954). Continuous Inspection Schemes. Biometrika 41(1/2).
  • Lucas, J. & Saccucci, M. (1990). Exponentially Weighted Moving Average Control Schemes: Properties and Enhancements. Technometrics 32(1).
  • Montgomery, D. C. (2019). Introduction to Statistical Quality Control. Wiley — the standard applied textbook across this whole field.

What to learn next