Feature and Data Pipelines in Production

Point-in-time correctness

Point-in-time correctness means every number shown to a model was actually known at that exact moment in the past, never filled in using something from later.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Point-in-time correctness means using only the information that was actually known at that exact moment in the past.

The analogy you have already lived

Think of your grandmother's old bank passbook. Each page shows the balance on that exact date, stamped by a bank clerk.

If you want to know her balance on 3 March, you check the 3 March entry. You never peek at June's entry and pretend it was already true in March.

Why it exists

A model learns from past examples. Each example needs feature values as they stood at that moment, plus the true outcome.

If a feature value from AFTER that moment sneaks in, the model is peeking at the future. In training it looks brilliant. In production it fails, because the future has not happened yet when a real prediction is needed.

This mistake is called leakage — information from outside the allowed window reaching the model by accident. It is one of the most common ways a machine learning project quietly lies to its own team.

How it works

training example:  "was this user a big spender AS OF 15 Jan?"

purchases table:
  05 Jan -- Rs 200
  12 Jan -- Rs 150
  25 Jan -- Rs 900    <-- happened AFTER 15 Jan

WRONG (leaky):         sum everything        -> Rs 1250
RIGHT (as-of 15 Jan):  sum rows on/before it  -> Rs 350

The "as of" date travels with every training example. A correct pipeline always asks "what did we know at that time", never "what do we know now".

A real example you have seen

A UPI fraud model trains on "was this transaction later reported as fraud". It must use only the account's history up to that transaction's moment. Never later transactions on the same account, and never the fraud report itself, which by definition arrives afterwards.

The honest part

This bug rarely announces itself with an error. It shows up months later as a model that scored 95% offline and then quietly disappoints in production. Read that sentence twice — it is the single costliest mistake in this whole area.

Remember this

  • Every feature value must reflect what was known at that exact past moment.
  • Using a later value by accident is called leakage — it makes results look better than they really are.
  • A feature pipeline needs an "as of" timestamp on every training row, not only "today".

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Only Python's built-in sqlite3 module is used for the core example — no extra service needed to see the bug.

The bug, made concrete

pit_bug.py
import sqlite3

conn = sqlite3.connect(":memory:")
conn.execute("""
CREATE TABLE purchases (
    user_id INTEGER,
    amount REAL,
    ts TEXT   -- 'YYYY-MM-DD'
)
""")

rows = [
    (1, 200.0, "2026-01-05"),
    (1, 150.0, "2026-01-12"),
    (1, 900.0, "2026-01-25"),   # happens AFTER the training example's date
    (2,  50.0, "2026-01-03"),
    (2,  80.0, "2026-01-09"),
]
conn.executemany("INSERT INTO purchases VALUES (?,?,?)", rows)

# A training example: "was user 1 a big spender as of 2026-01-15?"
as_of = "2026-01-15"
user = 1

# WRONG: sums every row in the table, including purchases that had not
# happened yet on 2026-01-15. The model would train on information from
# the future.
leaky = conn.execute(
    "SELECT SUM(amount) FROM purchases WHERE user_id = ?", (user,)
).fetchone()[0]

# RIGHT: only rows whose timestamp is on or before the as-of date.
correct = conn.execute(
    "SELECT SUM(amount) FROM purchases WHERE user_id = ? AND ts <= ?",
    (user, as_of),
).fetchone()[0]

print(f"as_of date          : {as_of}")
print(f"leaky total (wrong)  : {leaky}")
print(f"point-in-time (right): {correct}")
Output
as_of date          : 2026-01-15
leaky total (wrong)  : 1250.0
point-in-time (right): 350.0

Line-by-line walkthrough

The leaky query has no time filter — it treats "now" as if it had always been true. Every past training row silently gets full knowledge of the future.

The correct query adds AND ts <= ?. That single condition is the whole fix. A feature store's real job is making this condition automatic, so nobody has to remember it by hand on every query.

Common mistakes

Joining a feature table on user ID alone, with no date filter. This is the exact bug above, hidden inside a bigger join. It is the single most common leakage bug in tabular machine learning.

Using a row's updated_at instead of the true event time. A late-arriving correction can push updated_at after your as-of date, even though the underlying fact was true earlier.

Trusting a pipeline that "worked last time". Point-in-time bugs rarely throw an error. They show up as a model that mysteriously stops matching its offline score once it meets real traffic.

Try it yourself

Add a fourth purchase dated exactly "2026-01-15" — the as-of date itself. Decide, then test, whether <= should include that day. Most teams do include it, but write the decision down as an explicit rule.

What to learn next

Researcher — Mathematics and papers.

The formal statement

For a feature value $f(u, t)$ of entity $u$ at time $t$, point-in-time correctness requires:

$$f(u, t) = g\big({e \in E_u : \text{time}(e) \le t}\big)$$

where $E_u$ is the full event history for entity $u$, $\text{time}(e)$ is each event's true occurrence time, and $g$ is the aggregation function — sum, count, average, or similar. No part of $f(u,t)$'s computation may depend on an event with $\text{time}(e) > t$.

Two timestamps, not one

Production feature stores distinguish event time (when a fact became true in the real world) from processing time (when the pipeline recorded it). A late-arriving event — a delayed webhook, a batch job running a day behind — has an event time in the past but a processing time close to now.

A correct as-of join filters on event time. Filtering on processing time instead is a subtler version of the same bug: it looks time-ordered while still leaking information a production system would not yet have had at that moment.

Point-in-time joins at scale

Feast, Tecton and Databricks Feature Store all implement this as a point-in-time join: for each (entity_id, as_of_timestamp) pair in a training set, find the latest feature value with event_time <= as_of_timestamp. A naive implementation costs $O(n \times m)$ over $n$ training rows and $m$ feature rows; production systems use sorted merge-joins or window functions — ASOF JOIN in DuckDB and ClickHouse, merge_asof in pandas — to bring this closer to $O((n + m) \log m)$.

Papers and systems

  • Akidau et al., The Dataflow Model, VLDB 2015 — the watermark and trigger concepts it introduces are the same machinery streaming feature pipelines use to decide when a value is final.
  • Feast's point-in-time join documentation is the most concrete open-source reference implementation — docs.feast.dev
  • pandas merge_asof — the smallest correct implementation of the join, worth reading once.

What to learn next

What to learn next

These follow on from what you just read.

  • Feature and Data Pipelines in Production

    Feature freshness

    Feature freshness is how old a stored number is right now, and whether that age is still safe to trust for the decision it is about to make.

  • Feature and Data Pipelines in Production

    Backfilling a new feature

    Backfilling means computing a brand-new feature for every past date it needs to exist, not only going forward from today.

  • Feature and Data Pipelines in Production

    Streaming feature aggregation

    Streaming aggregation keeps a running total updated as each new event arrives, instead of recomputing it from scratch by rescanning everything.