Manufacturing and Predictive Maintenance

Remaining useful life

Remaining useful life predicts how many days a machine has left, not only whether it will fail soon, and the honest version of that number gets more confident as the machine gets closer to failing.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Remaining useful life predicts how many days a machine has left, not only a yes-or-no warning that trouble is coming.

Think about a phone battery's "time remaining" estimate. It does not only say "low battery, replace soon." It gives an actual number. That number gets more accurate as the battery gets closer to zero. An estimate at 90% charge is a rough guess. One at 5% charge is quite precise.

Remaining useful life, usually shortened to RUL, is the same idea applied to machines. It does not ask "is this about to fail?" Instead it asks a more precise question: "roughly how many days, hours, or production cycles does this machine have left?"

Why it exists

The previous lesson's yes-or-no warning is useful, but it does not tell a maintenance planner everything they need. "This machine will fail within 7 days" gives no way to prioritise machines against each other. It also gives no way to schedule a repair around the rest of the factory's calendar.

A number changes that. If Machine A has an estimated 40 days left and Machine B has 4, a planner knows exactly which one needs attention first. Parts and downtime can now be scheduled around a real number, instead of a fixed calendar or a plain alarm.

How it works

A machine's "degradation index" -- a combined health score built from its
sensors -- climbs steadily from 0 (brand new) toward 100 (failure).

  Day 0:    degradation 1    -> a lot of life left, hard to be precise
  Day 60:   degradation 54   -> roughly half worn, moderate confidence
  Day 115:  degradation 98   -> nearly done, easy to be precise

A model trained on many machines that have already run to failure learns
the relationship between "how worn does this look" and "how many days
were actually left" once similar machines got to that same point.

Notice the pattern: RUL estimates are naturally less certain early in a machine's life and more certain near the end. A trustworthy RUL system should show that shrinking uncertainty honestly, not only a single confident-looking number throughout.

Where you have already seen it

  • A car dashboard's "estimated tyre life remaining."
  • A printer's "ink running low, approximately 40 pages left" warning.
  • A gym trainer estimating how many more sets you have in you. They judge this from how your form has been degrading over the workout — a very human version of the same idea.

Remember this

  • Remaining useful life predicts a number of days or cycles left, not only a yes-or-no alarm.
  • A number lets maintenance teams prioritise and plan, instead of reacting to a fixed calendar or a plain warning.
  • RUL predictions are naturally rough early in a machine's life and sharpen as failure approaches — trust the early numbers less.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas scikit-learn

Minimal runnable code

Six machines, each degrading at its own pace until failure, tracked by a combined "degradation index." We train a linear model on five of them and test it on a machine the model has never seen fail.

remaining_useful_life.py
import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error

rng = np.random.default_rng(12)

# Several machines, each degrading at its own pace until failure.
# A "degradation index" -- a stand-in for a combined sensor health score --
# climbs from 0 (brand new) toward 100 (failure) at a machine-specific rate.
machines = []
for machine_id in range(6):
    life = rng.integers(80, 140)
    rate = 100 / life
    noise = rng.normal(0, 1.5, life)
    degradation = np.clip(np.cumsum(np.full(life, rate)) + np.cumsum(noise) * 0.1, 0, 100)
    for day, value in enumerate(degradation):
        machines.append({"machine_id": machine_id, "day": day, "degradation": value,
                          "rul_true": life - day})

df = pd.DataFrame(machines)
print(df[df["machine_id"] == 0].iloc[[0, 30, 60, -1]])
print()

train = df[df["machine_id"] < 5]
test = df[df["machine_id"] == 5]

model = LinearRegression().fit(train[["degradation"]], train["rul_true"])
test = test.copy()
test["rul_predicted"] = model.predict(test[["degradation"]]).clip(min=0)

mae = mean_absolute_error(test["rul_true"], test["rul_predicted"])
print(f"mean absolute error on held-out machine: {mae:.1f} days")
print()

sample = test.iloc[[0, 30, 60, -5, -1]]
print(sample[["day", "degradation", "rul_true", "rul_predicted"]].round(1))
Output
     machine_id  day  degradation  rul_true
0             0    0     1.018990       116
30            0   30    26.828564        86
60            0   60    53.914987        56
115           0  115   100.000000         1

mean absolute error on held-out machine: 12.2 days

     day  degradation  rul_true  rul_predicted
492    0          0.8       121           99.6
522   30         27.1        91           73.5
552   60         53.0        61           47.7
608  116         98.1         5            2.9
612  120        100.0         1            1.0

What actually happened

The model learns a single relationship — roughly, "degradation score X has historically meant Y days left" — pooled across every training machine's full run to failure.

Look at the errors across the held-out machine's life. At day 0, the true remaining life was 121 days but the model predicted only 99.6 — a 21-day miss. By day 120, with the true remaining life at 1 day, the prediction was 1.0 — essentially exact. This is the pattern the beginner section described: RUL prediction sharpens as failure approaches, because a machine near the end of its life looks distinctly different from a healthy one, while a brand-new machine looks similar regardless of how long it will ultimately last.

Common mistakes

Reporting one MAE number and calling the model "accurate to within 12 days." As shown above, that average hides a real pattern: early-life predictions are far less reliable than late-life ones. Report accuracy broken down by how degraded the machine currently is, not as one blended number.

Training on too few machines. Six machines here is barely enough to illustrate the idea, and is not enough for a trustworthy real model. Each machine's individual quirks — a slightly different wear pattern, a different operating environment — need many examples to average out.

Treating a linear relationship between degradation and RUL as guaranteed. Real degradation is often closer to exponential near the end of life — see the researcher section below. A plain linear model, as used here for simplicity, can systematically overestimate remaining life right before failure.

Ignoring that RUL is fundamentally a censored problem, much like demand. A machine still running when the dataset was built has an unknown true RUL — treating it as "definitely healthy" rather than "not yet known to be near failure" introduces the same bias covered in You measure sales, not demand.

Try it yourself

Compute the mean absolute error separately for test[test["day"] < 30] versus test[test["day"] > 90]. The gap between those two numbers is the honest way to describe this model's accuracy — a single blended MAE hides it.

What to learn next

Researcher — Mathematics and papers.

RUL as a regression-on-degradation-trajectory problem

Formally, given a degradation signal x(t) (raw sensor data or an engineered health index) observed up to the current time t_c for a unit whose failure time is T, remaining useful life is:

RUL(t_c) = T - t_c

The developer example regresses RUL directly on the current degradation level x(t_c), discarding the trajectory's shape and rate of change. Production RUL models typically also condition on trend features — rate of change over a recent window, curvature, time since the trend became statistically significant — since two machines at the same degradation level but different recent rates of change genuinely have different expected remaining life.

Why degradation is often better modelled as exponential, not linear

Many physical degradation processes — bearing wear, corrosion, fatigue crack growth — accelerate as damage accumulates, following approximately exponential or power-law dynamics rather than the linear ramp used in the developer example for simplicity. The Paris–Erdogan law for fatigue crack growth,

da/dN = C * (delta_K)^m
  • a — crack length
  • N — number of load cycles
  • delta_K — the stress intensity factor range
  • C, m — material-specific constants fitted empirically

is the classical physical model underlying this: crack growth rate scales with a power of the current crack size, producing accelerating, not linear, degradation. A purely data-driven RUL model that ignores this tends to underestimate risk right before failure, exactly the systematic error flagged in the developer section.

Deep learning approaches on the CMAPSS benchmark

The NASA CMAPSS turbofan degradation dataset (Saxena and Goebel, 2008) is the standard public benchmark for RUL prediction. LSTM and 1D-CNN architectures (Zheng et al., 2017, Long Short-Term Memory Network for Remaining Useful Life Estimation) dominate published results, learning the trajectory shape directly from raw multivariate sensor sequences rather than a single hand-engineered health index. A near-universal practical trick in this literature: clip the RUL target at a maximum value (commonly 125–130 cycles for CMAPSS), because early-life RUL is both harder to predict accurately and less operationally relevant — nobody schedules maintenance around a machine reported as "115 vs 130 days left," so training the model to distinguish those cases precisely wastes model capacity that would help more near end-of-life.

Uncertainty quantification is not optional here

Because RUL feeds a scheduling and cost decision directly, a point estimate without an honest uncertainty band is close to unusable in practice: a maintenance planner needs to know whether "40 days remaining" means "confidently between 35 and 45" or "could be anywhere from 10 to 90." Bayesian deep learning (Monte Carlo dropout, deep ensembles) and conformal prediction, covered generally in Conformal prediction, are both used in production RUL systems specifically to produce this interval, not only a mean.

Key references

  • Saxena, A. & Goebel, K. (2008). Turbofan Engine Degradation Simulation Data Set. NASA Ames Prognostics Data Repository.
  • Zheng, S., Ristovski, K., Farahat, A. & Gupta, C. (2017). Long Short-Term Memory Network for Remaining Useful Life Estimation. IEEE International Conference on Prognostics and Health Management.
  • Paris, P. & Erdogan, F. (1963). A Critical Analysis of Crack Propagation Laws. Journal of Basic Engineering 85(4).

What to learn next