Feature and Data Pipelines in Production

Feature freshness

Feature freshness is how old a stored number is right now, and whether that age is still safe to trust for the decision it is about to make.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Feature freshness is how old a stored number is, measured from right now.

The analogy you have already lived

You check the date printed on a milk packet before drinking it. Milk from three days ago might already be fine, or might already be spoiled. The date is what tells you, not how it tastes at first sip.

A stored feature is the same. A number sitting in a database is only useful if you know how long ago it was computed. And whether that age is still acceptable.

Why it exists

Point-in-time correctness makes sure a training example never sees the future. Freshness is the live, running version of the same worry. Does a feature used for a prediction right now still describe reality right now?

A "current balance" feature computed by last night's batch job is up to 24 hours old by evening. For a loan approval that is often fine. For a UPI fraud check happening mid-transaction, 24 hours is far too old — the account could be empty by then.

How it works

feature: user_42:last_login_days
computed at: 11:30 AM
now:         12:00 PM
age:         30 minutes

allowed age for this feature: 15 minutes
30 > 15  ->  STALE. Do not trust this value for a live decision.

Every feature needs an agreed maximum age. That number is a business decision, not a technical one — ask whoever owns the use case, not only the pipeline.

A real example you have seen

A food delivery app's "estimated preparation time" feature. Say it was computed five minutes ago, and the kitchen has suddenly got busy since. The app is now showing you an estimate the kitchen has already outgrown.

The honest part

Freshness limits are rarely obvious. Setting them too strict wastes computation recomputing things nobody needed refreshed yet. Setting them too loose ships stale numbers as if they were current. There is no universal right answer — only "picked deliberately" versus "never decided at all".

Remember this

  • Freshness is a feature's age, not its correctness — a fresh wrong number is still wrong.
  • Every feature needs an agreed maximum age, decided with the business, not guessed by an engineer alone.
  • The right limit depends entirely on how fast the underlying thing changes, and how costly a stale answer is.

What to learn next

Developer — Code and libraries.

Setup

No installation needed — this uses only Python's standard library.

A small freshness monitor

freshness_check.py
from datetime import datetime

# A tiny in-memory feature cache: feature_name -> (value, computed_at)
feature_cache = {
    "user_42:avg_order_value":  (540.0, datetime(2026, 9, 2, 8, 0)),
    "user_42:last_login_days":  (2,     datetime(2026, 9, 2, 11, 30)),
    "user_42:churn_risk_score": (0.71,  datetime(2026, 8, 30, 9, 0)),  # stale
}

# Each feature is allowed to be this many minutes old before it is "stale".
max_age_minutes = {
    "user_42:avg_order_value":  24 * 60,  # daily batch job, one day is fine
    "user_42:last_login_days":  15,       # should update near-live
    "user_42:churn_risk_score": 60,       # hourly job
}

now = datetime(2026, 9, 2, 12, 0)

def check_freshness(computed_at, limit_minutes):
    age_minutes = (now - computed_at).total_seconds() / 60
    return age_minutes, age_minutes > limit_minutes

print(f"{'feature':<28} {'age (min)':>10} {'limit':>7}  status")
for name, (value, computed_at) in feature_cache.items():
    age, stale = check_freshness(computed_at, max_age_minutes[name])
    status = "STALE - do not trust" if stale else "fresh"
    print(f"{name:<28} {age:>10.0f} {max_age_minutes[name]:>7}  {status}")
Output
feature                       age (min)   limit  status
user_42:avg_order_value             240    1440  fresh
user_42:last_login_days              30      15  STALE - do not trust
user_42:churn_risk_score           4500      60  STALE - do not trust

Line-by-line walkthrough

feature_cache stands in for what a real feature store keeps: a value, and the timestamp it was computed. Nothing here is special about the storage — a dictionary, Redis and a full feature store all keep the same two facts.

max_age_minutes is the part teams skip. Deciding it up front, per feature, is what turns "seems fine" into a number you can actually alert on.

Common mistakes

No computed_at stored at all. Without it, freshness cannot be checked — only guessed. Always store the timestamp next to the value, not only the value.

One global freshness rule for every feature. A slow-moving feature like "account age in years" and a fast-moving one like "items in cart" do not belong to the same rule.

Treating a stale value as an error instead of a decision. Sometimes serving a slightly stale number is the right trade-off against failing the request entirely — see missing features at serve time for that decision in full.

Try it yourself

Add a feature user_42:cart_total with a 2-minute limit, and a computed_at from 90 seconds ago. Confirm it reports fresh, then move it to 3 minutes ago and confirm it flips to stale.

What to learn next

Researcher — Mathematics and papers.

Freshness as a service-level objective

Treat freshness the way reliability engineering treats latency: as a distribution with a target, not a single number. Define it as

$$P(\text{age}(f, t) \le L) \ge \tau$$

where $\text{age}(f, t)$ is a feature's age at query time $t$, $L$ is the agreed maximum age, and $\tau$ is the fraction of requests that must meet it — for example, 99% of requests must see a feature no older than 15 minutes.

This framing matters because pipelines fail intermittently, not uniformly. A single freshness check at deploy time misses the tail: a job that is usually fast but occasionally stalls for two hours will violate freshness for a small, real fraction of production traffic, invisible to a spot check.

Where staleness comes from

  • Batch cadence — a nightly or hourly job has an inherent floor on freshness equal to its own schedule.
  • Pipeline backlog — a streaming job that falls behind its input rate accumulates a growing lag, distinct from its steady-state latency.
  • Upstream outage — a source system going down does not stop requests; it silently increases every downstream feature's age.
  • Partial failure — one shard, partition or entity segment lagging while the rest of the pipeline looks healthy in aggregate metrics.

Detecting it in production

The direct measurement is event-time lag: the gap between the current wall-clock time and the watermark of the most recently fully-processed event, per feature or per pipeline stage. Streaming systems (Flink, Spark Structured Streaming, Kafka Streams) expose this as a first-class metric, since it is also what they use internally to decide when a windowed aggregate is final.

The naive alternative — alerting only when a job fails — misses staleness caused by a job that succeeds but runs slowly, which in practice is the more common failure mode at scale.

Papers and systems

  • Akidau et al., The Dataflow Model, VLDB 2015 — watermarks, the formal tool behind "how do we know this window is complete".
  • Kreps, I heart logs, O'Reilly 2014 — the reasoning behind treating a feature pipeline's backlog as a first-class operational metric.
  • Uber's Michelangelo and Airbnb's Zipline engineering write-ups both describe freshness SLOs as a per-feature, not per-pipeline, configuration in production feature platforms.

What to learn next