Feature and Data Pipelines in Production

Backfilling a new feature

Backfilling means computing a brand-new feature for every past date it needs to exist, not only going forward from today.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Backfilling a feature means computing it for every date in the past it needs to exist, not only from today onward.

The analogy you have already lived

Imagine a school adds a new column to its attendance register: "arrived by bicycle, yes or no". Starting today is easy. But the term's attendance report needs that column filled in for every past day too. Someone has to go back through old records and fill it in retroactively.

A new model feature works the same way. It is easy to compute going forward. The hard part is filling it in for every historical date a model will train on.

Why it exists

Say you invent a useful new feature: "amount spent in the last 7 days". You can compute it for today onward with no trouble.

But your training data has thousands of past examples. Each one needs that same feature as it stood on ITS OWN date, respecting point-in-time correctness. Without a backfill, the new feature only exists for brand-new rows. Every past training row has a gap.

How it works

new feature invented today: "spend_last_7d"

for EVERY past date a training row might need:
    recompute spend_last_7d as of that date
    store it, dated correctly

result: the feature now has a full history,
        not only values starting from today

This is usually a large, one-time batch job — sometimes the biggest single compute cost in adding one new feature.

A real example you have seen

A shopping app adds "days since last return" as a new fraud signal. Training a model on the last year of orders needs that signal for every one of those past orders. Not only for orders placed after the feature was invented.

The honest part

Backfills are slow, and it is tempting to skip them and start collecting the feature from today instead. That choice quietly throws away a year of training data for that one feature. It is often a worse trade than the backfill's own cost.

Remember this

  • A new feature needs values for every past date it will be trained on, not only going forward.
  • Backfilling means running the same logic used for live traffic, against historical data.
  • Skipping a backfill is a real choice with a real cost — usually a smaller, weaker training set.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Only sqlite3 from the standard library is used in the example itself.

Backfilling a rolling-window feature

backfill.py
import sqlite3
import time
import random
from datetime import datetime, timedelta

random.seed(0)
conn = sqlite3.connect(":memory:")
conn.execute("CREATE TABLE events (user_id INTEGER, amount REAL, ts TEXT)")

# 500 synthetic purchase events for 20 users over 90 days.
start = datetime(2026, 1, 1)
rows = []
for _ in range(500):
    user = random.randint(1, 20)
    day = start + timedelta(days=random.randint(0, 89))
    amount = round(random.uniform(50, 500), 2)
    rows.append((user, amount, day.strftime("%Y-%m-%d")))
conn.executemany("INSERT INTO events VALUES (?,?,?)", rows)
conn.commit()

conn.execute("""
CREATE TABLE feature_7d_spend (
    user_id INTEGER, as_of_date TEXT, value REAL
)
""")

# The NEW feature: "amount spent in the 7 days up to and including as_of_date".
# It has to be computed for every past date a training example might use,
# not only from today onward. That is the backfill.
dates = [(start + timedelta(days=d)).strftime("%Y-%m-%d") for d in range(90)]
users = list(range(1, 21))

t0 = time.perf_counter()
inserted = 0
for as_of in dates:
    window_start = (datetime.strptime(as_of, "%Y-%m-%d") - timedelta(days=6)).strftime("%Y-%m-%d")
    for user in users:
        total = conn.execute(
            "SELECT COALESCE(SUM(amount), 0) FROM events "
            "WHERE user_id = ? AND ts BETWEEN ? AND ?",
            (user, window_start, as_of),
        ).fetchone()[0]
        conn.execute(
            "INSERT INTO feature_7d_spend VALUES (?, ?, ?)", (user, as_of, total)
        )
        inserted += 1
conn.commit()
elapsed = time.perf_counter() - t0

print(f"backfilled rows      : {inserted}  ({len(dates)} dates x {len(users)} users)")
print(f"time on this machine : {elapsed:.2f}s  (timing only, not a benchmark claim)")

sample = conn.execute(
    "SELECT * FROM feature_7d_spend WHERE user_id = 3 ORDER BY as_of_date LIMIT 5"
).fetchall()
print("sample rows for user 3:", sample)
Output
backfilled rows      : 1800  (90 dates x 20 users)
time on this machine : 0.04s  (timing only, not a benchmark claim)
sample rows for user 3: [(3, '2026-01-01', 317.56), (3, '2026-01-02', 317.56), (3, '2026-01-03', 317.56), (3, '2026-01-04', 317.56), (3, '2026-01-05', 317.56)]

That 0.04 seconds is one small in-memory run on one machine, not a real-world benchmark. A real backfill runs against millions of rows and is measured in hours, which is exactly why the next lesson exists.

Line-by-line walkthrough

The double loop over dates and users is doing, by brute force, exactly what a real backfill job does at scale: recompute the feature's definition for every entity, at every point in time it is needed.

Notice the query inside the loop is the same shape as a live lookup would use — SUM over a bounded window ending at as_of. A backfill is not different logic from serving; it is the same logic, run once over history instead of once per request.

Common mistakes

Backfilling with today's logic against data that used to mean something else. If a column's meaning changed last year, backfilling blindly recomputes history using rules that were not true back then. Check feature versioning before trusting an old backfill.

Running one giant query instead of chunking by date or user. At real scale this either times out or uses more memory than the machine has. Chunk it, and make each chunk resumable.

No idempotency. If the job is interrupted halfway and re-run, does it produce duplicate rows, or pick up cleanly? Design for "re-run safely" from the start — see how the RAG ingestion lesson handles this same problem for documents.

Forgetting to backfill the SAME feature that will run live. Two slightly different implementations — one for backfill, one for serving — is called training-serving skew, and it is one of the hardest bugs to find later.

Try it yourself

Change the window from 7 days to 30 days, and re-run. Time it again. The backfill cost scales with the number of (user, date) pairs, not with the window length — confirm that from your own timing.

What to learn next

Researcher — Mathematics and papers.

Backfill as a batch materialisation problem

A backfill computes $f(u, t)$ for a Cartesian product of entities $U$ and timestamps $T$: $|U| \times |T|$ evaluations of the feature function. Naive implementations recompute each evaluation independently, at cost $O(|U| \times |T| \times c)$ where $c$ is the per-window aggregation cost.

For a fixed-window aggregate (sum, count, mean over a trailing $w$-day window), this is wasteful: consecutive as-of dates for the same entity share almost all of their underlying events. A sliding accumulator, processed once per entity in event-time order, reduces this to $O(|U| \times |E_u|)$ total work — the same complexity argument behind streaming feature aggregation, applied to a historical batch instead of a live stream.

Backfill consistency with the live path

The correctness requirement is that the backfilled value and the value the live pipeline would have produced, had it existed on that date, are identical:

$$f_{\text{backfill}}(u, t) = f_{\text{live}}(u, t) \quad \forall (u, t)$$

Feature platforms enforce this by sharing one feature-definition artifact between the batch and streaming code paths (a single SQL or Python transform compiled to both engines), rather than maintaining two hand-written implementations. Tecton and Databricks Feature Store both make this the central design constraint of their transform APIs, precisely because divergence here is training-serving skew by another name.

Scheduling and cost

Backfills are typically the largest single compute cost in a feature platform's lifecycle for a given feature, and are usually run once, on dedicated batch infrastructure (Spark, BigQuery, DuckDB over Parquet), separate from the low-latency online path. Chunking by date range with checkpointed progress, and by entity partition, is standard practice for jobs that would otherwise run for many hours and risk a partial failure losing all progress.

Papers and systems

  • Hermann and Del Balso, Meta's feature store: the missing data layer in ML pipelines, describes the shared batch/streaming transform design.
  • Feast's materialization engine documentation covers backfill scheduling and incremental re-materialization directly — docs.feast.dev
  • Karau and Warren, High Performance Spark — the chunked, checkpointed job pattern this lesson's backfill loop is a miniature of.

What to learn next