Feature and Data Pipelines in Production

Missing features at serve time

A model must still answer when a feature is missing at request time, so a pipeline needs a deliberate, agreed strategy for gaps instead of a crash.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A serving system needs a deliberate plan for a missing feature, because "crash the request" is rarely the right answer.

The analogy you have already lived

A doctor examining a patient without one lab result still has to make a decision. She does not refuse to treat the patient — she uses what she has, notes what is missing, and proceeds carefully.

A model at serving time is in the same position. One feature failed to arrive in time. The request still needs an answer.

Why it exists

A feature can be missing for good reasons. A brand-new user with no history yet. An upstream service that timed out. A low-latency lookup that came back empty for a key that was never written.

None of these are bugs in the request itself. They are normal, expected situations a live system has to survive. That means someone has to decide, ahead of time, what happens next.

How it works

feature arrives?  --- yes --->  use it
      |
      no
      |
      v
 pick ONE, decided in advance:
   - use an agreed default value
   - use the last known value, even if a little old
   - reject the request and say why

The wrong choice is making this decision silently, differently, in whichever code path happens to touch it first.

A real example you have seen

A ride-hailing app's price estimate for a brand-new rider who has no ride history yet. The app does not crash. It falls back to a citywide average instead, an openly rough starting point rather than a personalised number.

Remember this

  • A missing feature is a normal situation, not a bug — plan for it.
  • The main choices are a default value, a last-known value, or a clear rejection — pick one on purpose.
  • Whatever the model does with a missing feature should be visible in the response, not hidden.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]" httpx

Three requests, three situations

missing_features.py
from typing import Optional

from fastapi import FastAPI
from fastapi.testclient import TestClient
from pydantic import BaseModel

app = FastAPI()

# A per-feature fallback used only when the feature store could not answer
# in time. Not a guess pulled from nowhere -- typically the training-set
# median, agreed with whoever owns the feature.
DEFAULTS = {"recent_txn_count": 0, "avg_order_value": 220.0}


class ScoreRequest(BaseModel):
    user_id: int
    recent_txn_count: Optional[int] = None
    avg_order_value: Optional[float] = None


@app.post("/score")
def score(req: ScoreRequest):
    used_default = []
    txn_count = req.recent_txn_count
    if txn_count is None:
        txn_count = DEFAULTS["recent_txn_count"]
        used_default.append("recent_txn_count")

    order_value = req.avg_order_value
    if order_value is None:
        order_value = DEFAULTS["avg_order_value"]
        used_default.append("avg_order_value")

    # toy scoring formula standing in for a real model
    risk = round(min(1.0, txn_count * 0.05 + order_value / 2000), 3)
    return {"user_id": req.user_id, "risk": risk, "used_default_for": used_default}


client = TestClient(app)

full = client.post("/score", json={"user_id": 1, "recent_txn_count": 4, "avg_order_value": 300.0})
print("all features present   :", full.json())

partial = client.post("/score", json={"user_id": 2, "recent_txn_count": 4})
print("one feature missing     :", partial.json())

cold_start = client.post("/score", json={"user_id": 3})
print("brand-new user, no data :", cold_start.json())
Output
all features present   : {'user_id': 1, 'risk': 0.35, 'used_default_for': []}
one feature missing     : {'user_id': 2, 'risk': 0.31, 'used_default_for': ['avg_order_value']}
brand-new user, no data : {'user_id': 3, 'risk': 0.11, 'used_default_for': ['recent_txn_count', 'avg_order_value']}

Line-by-line walkthrough

Optional[int] = None is what lets a feature legitimately be absent from the request, instead of the request failing validation outright — compare this to model serving's Applicant model, where every field is required, because that lesson's fields always come directly from the caller and have no upstream feature pipeline that could fail.

used_default_for is the important part of the response. A caller — or a later debugging session — can see exactly which numbers were real and which were filled in, instead of guessing after the fact.

Common mistakes

Silently substituting zero for every missing number. Zero is rarely a neutral value — for avg_order_value, zero looks like a customer who never buys anything, which is a very different signal from "we do not know yet".

Not recording that a default was used. Six months later, nobody can tell which historical predictions were based on real data and which were guesses. Always log it alongside the prediction.

One fallback strategy for every feature. A slightly stale "account balance" is often fine to reuse. A stale "is this card blocked" is not — reusing an old value there could approve a transaction that should have been refused. Decide per feature, as in feature freshness.

Treating a timeout the same as "no such feature exists". A timeout might resolve if retried; a genuinely absent feature will not. Handling them identically throws away information that could have improved the fallback.

Try it yourself

Add a third strategy: reject the request outright with a 422 when recent_txn_count is missing, but keep the default for avg_order_value. Decide which features in your own project deserve which strategy, and write the reason down.

What to learn next

Researcher — Mathematics and papers.

Missingness is not one mechanism

Statistics distinguishes three mechanisms, and the right handling strategy depends on which one applies:

  • MCAR (missing completely at random) — the absence has no relationship to anything, observed or not. Rare in production; a purely random sensor dropout is close to this.
  • MAR (missing at random) — the probability of being missing depends on observed data. A new user's feature is missing precisely because they are new — "new" is itself observable and predictive.
  • MNAR (missing not at random) — the probability of being missing depends on the unobserved value itself. A fraud feature that fails to compute specifically for accounts under active attack is the dangerous case: the missingness is correlated with the very outcome being predicted.

Most serving-time missingness in practice is MAR or MNAR, not MCAR — which is why filling gaps with a single global default, uninformed by why the value is missing, routinely biases the resulting prediction rather than only adding noise to it.

Making missingness a feature, not a hidden default

A stronger pattern than silent imputation: pass an explicit is-missing indicator alongside the imputed value, so the model can learn a different response when a feature was absent versus when it was genuinely low. This requires the indicator to exist at training time too — meaning the training pipeline must simulate the same missingness the serving pipeline will see, not train exclusively on complete rows.

$$\hat{y} = f(x_1, \dots, x_k, m_1, \dots, m_k), \quad m_i = \mathbb{1}[x_i \text{ was missing}]$$

Tree ensembles (XGBoost, LightGBM) handle missing values natively by learning a default split direction per node, which is a more principled version of the same idea, learned rather than hand-set.

Fallback strategy as a per-feature contract

A production feature registry (see feature versioning) should declare, per feature, a machine-readable fallback policy: reject, default(value), last_known(max_age), or model_native (pass a sentinel and let the model's own missing-value handling take over). Making this declarative, rather than implicit in serving code, is what lets it be audited and changed without a code deploy.

Papers and systems

  • Rubin, Inference and Missing Data, Biometrika 1976 — the original MCAR/MAR/MNAR taxonomy.
  • Chen and Guestrin, XGBoost: A Scalable Tree Boosting System, KDD 2016 — the learned default-direction handling of missing values.
  • Breck et al., The ML Test Score, IEEE Big Data 2017 — Data 4 in its rubric specifically requires testing feature-generation code against known missing-value cases.

What to learn next