Manufacturing and Predictive Maintenance

Soft sensors

A soft sensor predicts a slow, expensive, or destructive lab measurement continuously from cheap sensors that already run every second, filling the long gaps between real lab tests.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A soft sensor predicts a slow, expensive lab measurement in real time, using the cheap sensors that are already running continuously.

Think about a fruit seller checking whether a mango is ripe. They do not cut every mango open to check — that would destroy it, and then it cannot be sold. Instead, they press it gently, smell it, and look at its colour, judging ripeness from things that are quick and harmless to check.

A factory faces the same trade-off constantly. The one measurement that truly matters — product quality — often needs a slow, expensive lab test. Sometimes that test even destroys the sample being checked.

Why it exists

Temperature and pressure sensors report a new number every second, cheaply and continuously. Many important quality measurements need a physical sample taken to a lab. The concentration of a chemical, the viscosity of a liquid, the strength of a material — all work this way. The result can take an hour or more to come back. Some tests destroy the sample entirely, meaning not every batch can even be tested.

This creates an awkward gap. A factory can watch its temperature and pressure every second, but can only truly know its product quality once every hour, or less often. For fifty-nine minutes out of every sixty, nobody actually knows how the product is doing.

A soft sensor — also called a virtual or inferential sensor — closes that gap. It is a model trained to predict the slow lab measurement directly from fast, always-available sensor readings. It learns that relationship from the handful of moments when both are known at once.

How it works

Fast sensors (every minute):     temperature, flow rate
Slow lab test (every 60 minutes): true product quality

              [ learn the relationship, using only the
                minutes where BOTH are known ]
                          |
                          v
       Now predict quality for EVERY minute, filling
       the huge gaps between real lab tests

The soft sensor does not replace the lab test entirely. The real lab measurement is still the ground truth. It is used both to build the model, and to periodically check the model is still accurate. What it replaces is the long, blind wait between lab results.

Where you have already seen it

  • A doctor estimating your body fat percentage from your height and weight (BMI), instead of a slower, more precise scan every time.
  • A weather forecaster estimating humidity from temperature and dew point, when a direct humidity sensor is unavailable or unreliable.
  • A cook judging a curry is ready by its smell, colour and how the oil separates, instead of formally testing every ingredient's exact concentration.

Remember this

  • A soft sensor predicts a slow, expensive, or destructive measurement using fast, cheap sensors that are already running.
  • It is trained on the rare moments when both the fast and slow readings exist together, then used to fill every gap in between.
  • The real lab measurement remains the source of truth. The soft sensor is a continuous estimate built on top of it, not a replacement for it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas scikit-learn

Minimal runnable code

Temperature and flow rate are reported every minute. True product quality only gets a lab reading every 60 minutes. We train a soft sensor on the rare moments both exist, then use it to estimate quality continuously.

soft_sensor.py
import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error

rng = np.random.default_rng(17)

n_minutes = 300
t = np.arange(n_minutes)

# Fast, cheap sensors: reported every single minute.
temperature = 80 + 5 * np.sin(t / 40) + rng.normal(0, 0.5, n_minutes)
flow_rate = 12 + 2 * np.cos(t / 55) + rng.normal(0, 0.3, n_minutes)

# The TRUE product quality (e.g. viscosity) is a real physical quantity
# driven by temperature and flow -- but nobody can read it directly.
true_quality = 0.6 * temperature - 1.2 * flow_rate + rng.normal(0, 0.4, n_minutes)

# The lab can only test a physical sample every 60 minutes, and it takes
# time to get the result back -- most minutes have NO quality reading at all.
lab_minutes = np.arange(0, n_minutes, 60)
lab_quality = true_quality[lab_minutes] + rng.normal(0, 0.3, len(lab_minutes))

print(f"sensor readings available: every minute, {n_minutes} total")
print(f"lab quality readings available: every 60 minutes, {len(lab_minutes)} total")
print()

# Train the soft sensor ONLY on the minutes where a lab reading exists.
X_train = np.column_stack([temperature[lab_minutes], flow_rate[lab_minutes]])
y_train = lab_quality
soft_sensor = LinearRegression().fit(X_train, y_train)

# Now predict quality for EVERY minute, filling the huge gaps between lab tests.
X_all = np.column_stack([temperature, flow_rate])
predicted_quality = soft_sensor.predict(X_all)

mae_at_lab_times = mean_absolute_error(lab_quality, predicted_quality[lab_minutes])
mae_overall = mean_absolute_error(true_quality, predicted_quality)
print(f"soft sensor error at the moments a lab reading exists: {mae_at_lab_times:.2f}")
print(f"soft sensor error across EVERY minute (vs the true, unmeasured quality): {mae_overall:.2f}")
print()

sample_minutes = [10, 25, 40, 55]
print("minute | true quality | soft-sensor estimate | a lab reading available?")
for m in sample_minutes:
    has_lab = "yes" if m in lab_minutes else "no"
    print(f"  {m:>3}  |   {true_quality[m]:6.2f}     |      {predicted_quality[m]:6.2f}         |  {has_lab}")
Output
sensor readings available: every minute, 300 total
lab quality readings available: every 60 minutes, 5 total

soft sensor error at the moments a lab reading exists: 0.16
soft sensor error across EVERY minute (vs the true, unmeasured quality): 0.37

minute | true quality | soft-sensor estimate | a lab reading available?
   10  |    31.12     |       31.16         |  no
   25  |    33.77     |       33.22         |  no
   40  |    34.43     |       34.16         |  no
   55  |    34.77     |       34.69         |  no

What actually happened

The soft sensor was trained on only 5 data points — the 5 minutes where a lab reading happened to exist. That is an extremely small training set by ordinary machine learning standards, and it works here because the underlying relationship between temperature, flow and quality is simple and consistent, not because 5 examples is generally enough.

The genuinely interesting numbers are the four sample minutes, none of which ever had a real lab reading. In this synthetic example we can peek at true_quality, since we generated it — in a real factory, that column would not exist at all outside the 5 lab-tested minutes. The soft sensor's estimates track it closely throughout, turning five isolated snapshots into a continuous, usable read of product quality.

The error at real lab moments (0.16) is smaller than the error across every minute (0.37) — expected, since the model was fitted to match the lab points exactly, and its errors on the in-between minutes depend on how well the learned relationship generalises across the whole range of operating conditions, not only the five it was trained on.

Common mistakes

Trusting the soft sensor forever without rechecking it against new lab results. Process conditions drift over months — a soft sensor validated once needs periodic recalibration against fresh lab data, not a one-time fit.

Using a soft sensor's output as if it were a real lab measurement, without tracking its own uncertainty. A soft sensor's confidence should be lowest far from any conditions seen during training, and a good implementation flags exactly that, rather than reporting one confident number regardless of how unusual current conditions are.

Training on too few, too similar lab samples. Five lab readings that all happened during similar operating conditions teach the model very little about how it behaves outside that narrow range — a soft sensor is only as good as the diversity of conditions it was trained across.

Forgetting that correlation in the training window does not guarantee causation. If a soft sensor is trained during a period when two sensors happened to move together for an unrelated reason, it can learn a relationship that quietly breaks once conditions change — see Why a great model answers the wrong question.

Try it yourself

Reduce lab_minutes to only two readings, np.array([0, 299]), at the very start and end. Re-run, and check how much mae_overall grows — this shows how much a soft sensor's real-world reliability depends on the number and spread of the lab readings it was trained against.

What to learn next

Researcher — Mathematics and papers.

Soft sensors as an inverse and semi-supervised regression problem

Formally, a soft sensor learns f: X -> y, where X is a vector of fast, continuously available process variables and y is a slow, sparsely labelled quality variable, with |{labelled samples}| << |{X observations}|. This is a specific instance of semi-supervised regression: the fast sensor stream itself, even where y is unobserved, carries information about the shape and structure of X's distribution that a well-designed soft sensor should exploit, not only the handful of labelled pairs.

Just-in-time (JIT) and adaptive soft sensors

Chemical and process industries have used soft sensors in production for decades, and the applied literature (Kadlec, Gabrys and Strandt, 2009, Data-Driven Soft Sensors in the Process Industry, Computers & Chemical Engineering) distinguishes two broad families:

  • Global models, fitted once on historical data, as in the developer example, and periodically retrained.
  • Just-in-time (JIT) models, which fit a fresh local model at prediction time, using only the historical samples most similar to the current operating point — a form of local regression closely related to k-nearest-neighbours weighting. JIT models adapt naturally to slow process drift without an explicit retraining schedule, at the cost of higher per-prediction computation.

Handling the label-sparsity problem more rigorously

Beyond ordinary regression fitted to the few labelled points, two techniques are common in the process-industry literature specifically because labelled samples are this scarce:

  • Partial least squares (PLS) regression, which handles the typically high collinearity among process sensor variables (temperature, pressure and flow are rarely independent) better than ordinary least squares, and remains a de facto standard baseline in this field for exactly that reason.
  • Co-training and self-training semi-supervised schemes, where a preliminary model's own confident predictions on unlabelled fast-sensor data are added back into the training set as pseudo-labels, expanding the effective labelled set beyond the handful of true lab measurements — used cautiously, since pseudo-label errors can compound if the preliminary model is systematically biased.

Dynamic and time-delay effects

The developer example assumes quality at time t depends only on sensor readings at time t. Real processes usually have transport delay and residence time — a temperature change now affects product quality some minutes or hours later, once material has physically moved through the process. Production soft sensors typically include lagged features (temperature[t - k] for several values of k) or an explicit dynamic model (a state-space or ARX model) rather than assuming an instantaneous relationship, since ignoring transport delay systematically misattributes cause and effect between sensor and lab measurements taken at slightly different points in the process timeline.

Validation against genuine process drift

Because the labelled set is small, cross-validation on it alone is a weak check of soft-sensor reliability going forward. The standard practice is continuous prediction monitoring: track the soft sensor's residual against every new lab result as it arrives over time, and trigger retraining when residuals trend away from zero — a direct application of the production-monitoring ideas in Working with delayed labels, adapted to a setting where "ground truth" itself arrives on a slow, sparse schedule.

Key references

  • Kadlec, P., Gabrys, B. & Strandt, S. (2009). Data-Driven Soft Sensors in the Process Industry. Computers & Chemical Engineering 33(4).
  • Fortuna, L., Graziani, S., Rizzo, A. & Xibilia, M. G. (2007). Soft Sensors for Monitoring and Control of Industrial Processes. Springer.

What to learn next