Manufacturing and Predictive Maintenance
Remaining useful life
Remaining useful life predicts how many days a machine has left, not only whether it will fail soon, and the honest version of that number gets more confident as the machine gets closer to failing.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Remaining useful life predicts how many days a machine has left, not only a yes-or-no warning that trouble is coming.
Think about a phone battery's "time remaining" estimate. It does not only say "low battery, replace soon." It gives an actual number. That number gets more accurate as the battery gets closer to zero. An estimate at 90% charge is a rough guess. One at 5% charge is quite precise.
Remaining useful life, usually shortened to RUL, is the same idea applied to machines. It does not ask "is this about to fail?" Instead it asks a more precise question: "roughly how many days, hours, or production cycles does this machine have left?"
Why it exists
The previous lesson's yes-or-no warning is useful, but it does not tell a maintenance planner everything they need. "This machine will fail within 7 days" gives no way to prioritise machines against each other. It also gives no way to schedule a repair around the rest of the factory's calendar.
A number changes that. If Machine A has an estimated 40 days left and Machine B has 4, a planner knows exactly which one needs attention first. Parts and downtime can now be scheduled around a real number, instead of a fixed calendar or a plain alarm.
How it works
A machine's "degradation index" -- a combined health score built from its
sensors -- climbs steadily from 0 (brand new) toward 100 (failure).
Day 0: degradation 1 -> a lot of life left, hard to be precise
Day 60: degradation 54 -> roughly half worn, moderate confidence
Day 115: degradation 98 -> nearly done, easy to be precise
A model trained on many machines that have already run to failure learns
the relationship between "how worn does this look" and "how many days
were actually left" once similar machines got to that same point.Notice the pattern: RUL estimates are naturally less certain early in a machine's life and more certain near the end. A trustworthy RUL system should show that shrinking uncertainty honestly, not only a single confident-looking number throughout.
Where you have already seen it
- A car dashboard's "estimated tyre life remaining."
- A printer's "ink running low, approximately 40 pages left" warning.
- A gym trainer estimating how many more sets you have in you. They judge this from how your form has been degrading over the workout — a very human version of the same idea.
Remember this
- Remaining useful life predicts a number of days or cycles left, not only a yes-or-no alarm.
- A number lets maintenance teams prioritise and plan, instead of reacting to a fixed calendar or a plain warning.
- RUL predictions are naturally rough early in a machine's life and sharpen as failure approaches — trust the early numbers less.
What to learn next
- Building a model from four failures — what happens when there is barely enough failure history to learn from at all.
- Predictive maintenance — the yes-or-no version of this same problem.
- Prediction intervals for regression — showing a range instead of one confident number.
Developer — Code and libraries.
Setup
pip install numpy pandas scikit-learnMinimal runnable code
Six machines, each degrading at its own pace until failure, tracked by a combined "degradation index." We train a linear model on five of them and test it on a machine the model has never seen fail.
import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error
rng = np.random.default_rng(12)
# Several machines, each degrading at its own pace until failure.
# A "degradation index" -- a stand-in for a combined sensor health score --
# climbs from 0 (brand new) toward 100 (failure) at a machine-specific rate.
machines = []
for machine_id in range(6):
life = rng.integers(80, 140)
rate = 100 / life
noise = rng.normal(0, 1.5, life)
degradation = np.clip(np.cumsum(np.full(life, rate)) + np.cumsum(noise) * 0.1, 0, 100)
for day, value in enumerate(degradation):
machines.append({"machine_id": machine_id, "day": day, "degradation": value,
"rul_true": life - day})
df = pd.DataFrame(machines)
print(df[df["machine_id"] == 0].iloc[[0, 30, 60, -1]])
print()
train = df[df["machine_id"] < 5]
test = df[df["machine_id"] == 5]
model = LinearRegression().fit(train[["degradation"]], train["rul_true"])
test = test.copy()
test["rul_predicted"] = model.predict(test[["degradation"]]).clip(min=0)
mae = mean_absolute_error(test["rul_true"], test["rul_predicted"])
print(f"mean absolute error on held-out machine: {mae:.1f} days")
print()
sample = test.iloc[[0, 30, 60, -5, -1]]
print(sample[["day", "degradation", "rul_true", "rul_predicted"]].round(1)) machine_id day degradation rul_true
0 0 0 1.018990 116
30 0 30 26.828564 86
60 0 60 53.914987 56
115 0 115 100.000000 1
mean absolute error on held-out machine: 12.2 days
day degradation rul_true rul_predicted
492 0 0.8 121 99.6
522 30 27.1 91 73.5
552 60 53.0 61 47.7
608 116 98.1 5 2.9
612 120 100.0 1 1.0What actually happened
The model learns a single relationship — roughly, "degradation score X has historically meant Y days left" — pooled across every training machine's full run to failure.
Look at the errors across the held-out machine's life. At day 0, the true remaining life was 121 days but the model predicted only 99.6 — a 21-day miss. By day 120, with the true remaining life at 1 day, the prediction was 1.0 — essentially exact. This is the pattern the beginner section described: RUL prediction sharpens as failure approaches, because a machine near the end of its life looks distinctly different from a healthy one, while a brand-new machine looks similar regardless of how long it will ultimately last.
Common mistakes
Reporting one MAE number and calling the model "accurate to within 12 days." As shown above, that average hides a real pattern: early-life predictions are far less reliable than late-life ones. Report accuracy broken down by how degraded the machine currently is, not as one blended number.
Training on too few machines. Six machines here is barely enough to illustrate the idea, and is not enough for a trustworthy real model. Each machine's individual quirks — a slightly different wear pattern, a different operating environment — need many examples to average out.
Treating a linear relationship between degradation and RUL as guaranteed. Real degradation is often closer to exponential near the end of life — see the researcher section below. A plain linear model, as used here for simplicity, can systematically overestimate remaining life right before failure.
Ignoring that RUL is fundamentally a censored problem, much like demand. A machine still running when the dataset was built has an unknown true RUL — treating it as "definitely healthy" rather than "not yet known to be near failure" introduces the same bias covered in You measure sales, not demand.
Try it yourself
Compute the mean absolute error separately for test[test["day"] < 30] versus test[test["day"] > 90]. The gap between those two numbers is the honest way to describe this model's accuracy — a single blended MAE hides it.
What to learn next
Researcher — Mathematics and papers.
RUL as a regression-on-degradation-trajectory problem
Formally, given a degradation signal x(t) (raw sensor data or an engineered health index) observed up to the current time t_c for a unit whose failure time is T, remaining useful life is:
RUL(t_c) = T - t_cThe developer example regresses RUL directly on the current degradation level x(t_c), discarding the trajectory's shape and rate of change. Production RUL models typically also condition on trend features — rate of change over a recent window, curvature, time since the trend became statistically significant — since two machines at the same degradation level but different recent rates of change genuinely have different expected remaining life.
Why degradation is often better modelled as exponential, not linear
Many physical degradation processes — bearing wear, corrosion, fatigue crack growth — accelerate as damage accumulates, following approximately exponential or power-law dynamics rather than the linear ramp used in the developer example for simplicity. The Paris–Erdogan law for fatigue crack growth,
da/dN = C * (delta_K)^ma— crack lengthN— number of load cyclesdelta_K— the stress intensity factor rangeC,m— material-specific constants fitted empirically
is the classical physical model underlying this: crack growth rate scales with a power of the current crack size, producing accelerating, not linear, degradation. A purely data-driven RUL model that ignores this tends to underestimate risk right before failure, exactly the systematic error flagged in the developer section.
Deep learning approaches on the CMAPSS benchmark
The NASA CMAPSS turbofan degradation dataset (Saxena and Goebel, 2008) is the standard public benchmark for RUL prediction. LSTM and 1D-CNN architectures (Zheng et al., 2017, Long Short-Term Memory Network for Remaining Useful Life Estimation) dominate published results, learning the trajectory shape directly from raw multivariate sensor sequences rather than a single hand-engineered health index. A near-universal practical trick in this literature: clip the RUL target at a maximum value (commonly 125–130 cycles for CMAPSS), because early-life RUL is both harder to predict accurately and less operationally relevant — nobody schedules maintenance around a machine reported as "115 vs 130 days left," so training the model to distinguish those cases precisely wastes model capacity that would help more near end-of-life.
Uncertainty quantification is not optional here
Because RUL feeds a scheduling and cost decision directly, a point estimate without an honest uncertainty band is close to unusable in practice: a maintenance planner needs to know whether "40 days remaining" means "confidently between 35 and 45" or "could be anywhere from 10 to 90." Bayesian deep learning (Monte Carlo dropout, deep ensembles) and conformal prediction, covered generally in Conformal prediction, are both used in production RUL systems specifically to produce this interval, not only a mean.
Key references
- Saxena, A. & Goebel, K. (2008). Turbofan Engine Degradation Simulation Data Set. NASA Ames Prognostics Data Repository.
- Zheng, S., Ristovski, K., Farahat, A. & Gupta, C. (2017). Long Short-Term Memory Network for Remaining Useful Life Estimation. IEEE International Conference on Prognostics and Health Management.
- Paris, P. & Erdogan, F. (1963). A Critical Analysis of Crack Propagation Laws. Journal of Basic Engineering 85(4).