Manufacturing and Predictive Maintenance
Predictive maintenance
Predictive maintenance uses sensor data to warn that a machine is about to fail, instead of fixing it on a fixed calendar or waiting for it to break.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Predictive maintenance uses a machine's own sensor readings to warn that it is about to fail — before it actually does.
Think about the three ways you might treat a nagging cough. You could ignore it until it gets so bad you cannot get out of bed — fixing the problem only once it has become a crisis. You could visit a doctor every three months on a fixed schedule, whether you feel fine or not. That treats everyone the same, regardless of how they actually feel. Or a smartwatch could notice your resting heart rate creeping up over several days. It could nudge you to get checked, based on what is actually happening in your body right now.
Factories face the exact same three choices for every machine they run.
Why it exists
Reactive maintenance means fixing a machine after it breaks. It is simple, and it is often the most expensive option — an unplanned stop can halt an entire production line, not only one machine.
Preventive maintenance means servicing a machine on a fixed schedule — every 90 days, say — whether it needs it or not. This is safer than reactive maintenance, but wasteful. Healthy machines get serviced anyway. And a machine can still fail early, between scheduled visits, if something unusual is wearing it down faster than normal.
Predictive maintenance — often shortened to PdM — replaces the fixed calendar with the machine's actual sensor data. It does not ask "has enough time passed?" It asks a sharper question: "does this machine's vibration, temperature, or current draw look like the pattern that has preceded a failure before?"
How it works
Vibration reading, days before a known past failure:
Day -60 to -11: steady around 2.0 -> nothing unusual
Day -10: 2.3 -> a little higher
Day -8: 2.9 -> climbing
Day -5: 3.6 -> climbing faster
Day -2: 4.1 -> well outside normal range
A model trained on many past failures like this one learns to recognise
that climbing pattern -- and can raise an alarm days before the failure,
on a machine it has never seen fail before.The model is not predicting the future in some mysterious sense. It is recognising that the current sensor pattern resembles the pattern that, in the past, was followed by a breakdown.
Where you have already seen it
- A car's "check engine" light, triggered by sensor readings that resemble patterns linked to real problems, not by a fixed mileage counter.
- A laptop battery health warning, based on how the battery is actually degrading, not only its age.
- An elevator maintenance company prioritising which lifts to inspect first, using motor and cable sensor data instead of visiting every building on the same fixed rotation.
An honest warning
A predictive-maintenance model is a statistical pattern-matcher, not a physical inspection. It can be wrong in both directions. It might miss a real failure whose pattern it has never seen, or raise a false alarm over a harmless fluctuation.
For any machine where a missed failure risks safety, not only cost, a model's warning should trigger one thing only. That thing is a proper physical inspection by a qualified engineer. It should never trigger a decision to skip one. Building a genuinely reliable predictive-maintenance system for critical, safety-relevant equipment takes careful engineering validation, not only a good-looking model on historical data.
Remember this
- Predictive maintenance uses real sensor patterns to warn of failure, instead of a fixed calendar or waiting for a breakdown.
- It works by recognising sensor patterns that resembled past failures, not by directly foreseeing the future.
- A model's warning is a signal to inspect, not a replacement for the inspection itself, especially where safety is involved.
What to learn next
- Remaining useful life — predicting not only "will it fail soon" but "how many days are left."
- Building a model from four failures — what happens when real failures are rare enough that this lesson's easy example stops applying.
- Anomaly detection — the general idea of spotting unusual patterns, applied here to a machine's health.
Developer — Code and libraries.
Setup
pip install numpy pandas scikit-learnMinimal runnable code
Five run-to-failure cycles of a machine. Vibration climbs in the last 10 days before each failure. We train a classifier on four cycles and test it on a held-out fifth one.
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import precision_score, recall_score
rng = np.random.default_rng(8)
# Five run-to-failure cycles. In the last 10 days before each failure,
# vibration climbs -- a realistic early-warning pattern.
cycle_length = 60
n_cycles = 5
rows = []
for cycle in range(n_cycles):
for day in range(cycle_length):
days_to_failure = cycle_length - day
baseline_vibration = 2.0 + rng.normal(0, 0.15)
if days_to_failure <= 10:
baseline_vibration += (10 - days_to_failure) * 0.25
rows.append({"cycle": cycle, "day": day, "vibration": baseline_vibration,
"days_to_failure": days_to_failure})
df = pd.DataFrame(rows)
df["failing_within_7_days"] = (df["days_to_failure"] <= 7).astype(int)
df["vibration_trend"] = df.groupby("cycle")["vibration"].diff(5)
df = df.dropna()
print(df[["cycle", "day", "vibration", "vibration_trend", "failing_within_7_days"]].iloc[[0, 45, 52, 58]])
print()
print("share of days labelled 'failing within 7 days':", df["failing_within_7_days"].mean().round(3))
print()
# Split by CYCLE, not randomly -- the last cycle is held out entirely,
# so no day from the test failure event ever leaks into training.
train = df[df["cycle"] < n_cycles - 1]
test = df[df["cycle"] == n_cycles - 1]
features = ["vibration", "vibration_trend"]
model = RandomForestClassifier(n_estimators=200, class_weight="balanced", random_state=0)
model.fit(train[features], train["failing_within_7_days"])
pred = model.predict(test[features])
truth = test["failing_within_7_days"]
print(f"test cycle: {len(test)} days, {truth.sum()} of them within 7 days of failure")
print(f"precision: {precision_score(truth, pred):.2f}")
print(f"recall: {recall_score(truth, pred):.2f}")
print()
alarm_days = test.loc[pred == 1, "days_to_failure"]
if len(alarm_days):
print(f"model first predicts failure {alarm_days.max()} days before it happens")cycle day vibration vibration_trend failing_within_7_days 5 0 5 1.971665 0.232405 0 50 0 50 1.829910 -0.166174 0 57 0 57 4.130715 1.803306 1 68 1 8 2.093854 0.143396 0 share of days labelled 'failing within 7 days': 0.127 test cycle: 55 days, 7 of them within 7 days of failure precision: 1.00 recall: 1.00 model first predicts failure 7 days before it happens
What actually happened
failing_within_7_days is the label: 1 for any day within the final week before a breakdown, 0 otherwise. Only 12.7% of days carry a positive label — most of a machine's life is spent running normally, which is exactly why this is a harder problem than an ordinary balanced classification task.
vibration_trend captures the rate of change over the last 5 days, not only the current reading — a machine at vibration 3.6 that got there suddenly is a very different situation from one that has read 3.6 steadily for months. That single engineered feature carries most of the useful signal here.
The train/test split is by cycle, not by randomly shuffled rows. This matters more than it might look. A random row-level split would put some days from the test failure event into training — the model would effectively be tested on data it had already partly seen, producing a falsely optimistic score. Splitting by whole cycles keeps the entire held-out failure genuinely unseen.
On that clean held-out cycle, the model catches the failure exactly at the edge of its 7-day target window. Real factory data is far noisier than this synthetic example — perfect precision and recall like this is not something to expect outside a teaching example.
Common mistakes
Splitting data randomly instead of by machine or by run-to-failure cycle. This is the single most common way a predictive-maintenance model looks great in testing and fails in production — covered in depth in Your model scores 99%: what to suspect first.
Reporting accuracy instead of precision and recall. With only 12.7% of days labelled positive, a model that predicts "never failing" scores 87% accuracy while being completely useless. Always check precision and recall, or better, the full cost of false alarms against missed failures.
Picking the failure window arbitrarily. "7 days" here is a business decision — how much lead time does maintenance actually need to schedule a repair? A model tuned for a 7-day window will not automatically work well for a 2-day or 30-day one.
Treating a false alarm and a missed failure as equally bad. They almost never are. A missed failure can mean an unplanned production stop or a safety incident; a false alarm means an unnecessary inspection. Setting the decision threshold correctly needs both costs, using the same asymmetric-cost thinking as Choosing a threshold from costs.
Try it yourself
Change class_weight="balanced" to None and re-run. Watch recall drop — an unweighted model, faced with mostly-negative training data, becomes more cautious about predicting the rare positive class, which is exactly the wrong instinct for catching failures early.
What to learn next
Researcher — Mathematics and papers.
Framing predictive maintenance as survival analysis
The developer example frames PdM as a fixed-horizon binary classification problem — "will this machine fail within the next 7 days?" This is simple and effective, but it discards information: it does not use the fact that a machine 2 days from failure and one 6 days from failure receive the identical label. Survival analysis models the full time until failure directly:
S(t) = P(T > t) the survival function: probability of running past time t
h(t) = f(t) / S(t) the hazard function: instantaneous failure rate given survival to tT— the random variable for time until failuref(t)— the probability density of failure at exactly timeth(t)— the hazard rate, the standard quantity condition-monitoring systems actually want to track over time
The Cox proportional hazards model (Cox, 1972, Regression Models and Life-Tables, JRSS-B) extends this to covariates — sensor readings — by assuming h(t | x) = h_0(t) * exp(beta^T x), letting sensor features scale a baseline hazard multiplicatively. Random survival forests (Ishwaran et al., 2008) and deep survival models (DeepSurv, Katzman et al., 2018) relax the proportional-hazards assumption and handle nonlinear sensor interactions, at the cost of interpretability.
Censoring is the central statistical complication
Most machines in a historical maintenance dataset have not failed yet by the time the dataset is assembled — their true time-to-failure is right-censored, exactly as in You measure sales, not demand, but here the censoring mechanism is "still running" rather than "sold out." Treating a censored, still-healthy machine's last observed day as if it were a confirmed non-failure (as the simple binary-label approach implicitly does for machines beyond the labelling window) systematically biases a naive model toward underestimating failure risk. Proper survival-analysis likelihoods handle censored observations correctly by construction; naive classification labelling schemes generally do not, unless deliberately corrected for it.
Evaluation metrics specific to this framing
Ordinary classification metrics (precision, recall, ROC-AUC at a fixed horizon) remain useful for the binary framing. For survival-based models, the standard metric is Harrell's concordance index (C-index): the fraction of comparable machine pairs where the model correctly ranks which one failed first. This generalises ROC-AUC to a continuous time-to-event target and is the metric of choice in most published PdM benchmarks, including the widely used NASA CMAPSS turbofan degradation dataset.
Why the fixed-horizon classification framing remains dominant in industry
Despite the statistical elegance of survival models, most production PdM systems in industry still use fixed-horizon classification, for a concrete operational reason: maintenance scheduling itself usually operates on a fixed planning horizon (a weekly or monthly maintenance window), so "will this fail within the next scheduling cycle" is often literally the decision-relevant question, not an approximation of it. Survival models earn their added complexity most when maintenance lead time itself varies by failure mode, or when the business genuinely needs a continuous remaining-life estimate — the subject of the next lesson.
Key references
- Cox, D. R. (1972). Regression Models and Life-Tables. Journal of the Royal Statistical Society: Series B 34(2).
- Ishwaran, H., Kogalur, U., Blackstone, E. & Lauer, M. (2008). Random Survival Forests. Annals of Applied Statistics 2(3).
- Saxena, A. & Goebel, K. (2008). Turbofan Engine Degradation Simulation Data Set. NASA Ames Prognostics Data Repository — origin of the widely used CMAPSS benchmark.