Data Engineering for AI

Data versioning

Versioning data means being able to say exactly which rows trained a model, so that when a score changes you can find out what moved.

On this page 8
  1. Why this became a problem
  2. The question versioning answers
  3. How it works, in outline
  4. What you store alongside it
  5. Somewhere you have seen this
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Versioning data means recording exactly which rows you used. Later, you can prove what a model was trained on.

Think about a family recipe passed around on paper. Someone adds a pinch more salt and does not mention it. Someone else swaps the oil. Six months on, the dish tastes different and nobody can say which change did it.

If every version had been written down and dated, you could compare and find out.

Data changes the same way, quietly, by many hands.

Why this became a problem

Code has had version control for decades. You can see every line that changed, who changed it, and when.

Data got no such treatment. It sits in a database or a folder, and it changes underneath you. Somebody fixes a typo. A daily import brings 4,000 new rows. A column gets renamed.

Then the model's accuracy drops, and the only honest answer to "what changed?" is a shrug.

The question versioning answers

Six months from now, someone will ask: which exact rows produced this model?

Without versioning you cannot answer. With it, you have a short code that identifies the dataset precisely, saved next to the model.

How it works, in outline

   dataset  ->  read every row  ->  one short code
                                    "a0d1573c38cf"

   change one cell in one row
                                 ->  a completely different code
                                    "9ad1ba519e37"

   shuffle the rows, change nothing
                                 ->  the SAME code
                                    "a0d1573c38cf"

That short code is called a fingerprint. It is produced by a rule that reads all the data and returns a fixed-length string.

Two things make it useful. Any change to the contents, however small, changes the fingerprint. And harmless changes, like row order, do not.

What you store alongside it

The fingerprint alone tells you that something changed, not what. So teams save a small summary next to it:

  • How many rows and columns.
  • How many blanks in each column.
  • What fraction of rows have each answer.
  • A one-line note about what was done and why.

Compare two of those summaries and the difference usually jumps out.

Somewhere you have seen this

Your phone's photo backup does not upload a picture twice, even if you copy it into another folder. It compares fingerprints and recognises the file as one it already has.

Software downloads publish a checksum for the same reason. You can confirm the file you received is byte-for-byte the file that was published.

What is honestly hard here

You cannot copy the whole dataset every day. At any real size that is unaffordable in both storage and patience.

So there is a genuine trade-off between how much history you keep and what it costs. Different teams land in different places, and none of the answers are elegant.

Databases make it harder still. A table that is updated in place has no natural version. You have to design one in, usually by never deleting and always appending.

Remember this

  • Versioning answers: which exact rows trained this model?
  • A fingerprint changes when the contents change, and ignores harmless reordering.
  • Save a small summary next to the fingerprint. It is what makes a difference readable.

What to learn next

Developer — Code and libraries.

A fingerprint plus a manifest gets you most of the way, in about fifteen lines. Understanding this makes the tools that do it for you far less mysterious.

Setup

bash
pip install pandas

hashlib and json are in the standard library.

Fingerprint and manifest

fingerprint.py
import hashlib, json
import pandas as pd

def fingerprint(df):
    """A short, stable id for the exact contents of a frame."""
    ordered = df.sort_index(axis=1).sort_values(list(sorted(df.columns))).reset_index(drop=True)
    payload = ordered.to_csv(index=False).encode("utf-8")
    return hashlib.sha256(payload).hexdigest()[:12]

v1 = pd.DataFrame({
    "id":     [1, 2, 3, 4],
    "city":   ["pune", "mumbai", "nagpur", "pune"],
    "income": [42.0, 88.0, 31.0, 55.0],
    "repaid": [1, 1, 0, 1],
})
v2 = v1.copy()
v2.loc[2, "income"] = 310.0            # someone fixed a typo. Nobody wrote it down.
v3 = v1.iloc[[3, 0, 2, 1]]             # same rows, different order

for name, frame in [("v1", v1), ("v2 (one cell changed)", v2), ("v3 (rows reordered)", v3)]:
    print(f"{name:24s} {fingerprint(frame)}  rows={len(frame)}")

def manifest(df, note):
    return {
        "fingerprint": fingerprint(df),
        "rows": len(df),
        "columns": list(df.columns),
        "null_counts": {c: int(df[c].isna().sum()) for c in df.columns},
        "label_rate": round(float(df["repaid"].mean()), 4),
        "note": note,
    }

print()
print(json.dumps(manifest(v2, "income typo fixed for id=3"), indent=2))
Output
v1                       a0d1573c38cf  rows=4
v2 (one cell changed)    9ad1ba519e37  rows=4
v3 (rows reordered)      a0d1573c38cf  rows=4

{
  "fingerprint": "9ad1ba519e37",
  "rows": 4,
  "columns": [
    "id",
    "city",
    "income",
    "repaid"
  ],
  "null_counts": {
    "id": 0,
    "city": 0,
    "income": 0,
    "repaid": 0
  },
  "label_rate": 0.75,
  "note": "income typo fixed for id=3"
}

The two properties that matter

One changed cell produced a completely different fingerprint. a0d1573c38cf to 9ad1ba519e37. Not a similar string — a different one. That is what a cryptographic hash does, and it is why you cannot judge similarity from fingerprints, only equality.

Reordering the rows changed nothing. v3 holds the same four rows in a different order and hashes identically. That is the sort_values line doing its work.

Without that sort, every run that read rows in a different order would produce a new fingerprint, and you would learn to ignore the alerts. A version identifier that changes when nothing meaningful changed is worse than none, because it trains people to stop looking.

Line by line

df.sort_index(axis=1) sorts columns by name, so reordering columns does not change the fingerprint. Whether you want that is a decision — if column order is meaningful in your pipeline, drop this line.

sort_values(list(sorted(df.columns))) sorts rows by every column in turn, making row order irrelevant.

.to_csv(index=False) gives a deterministic text serialisation. Beware: floating-point formatting can vary across pandas versions, so a fingerprint is stable within an environment, not necessarily across upgrades. Pin your pandas version if fingerprints must survive across years, or serialise with explicit rounding.

hexdigest()[:12] truncates to 12 hex characters, 48 bits. Fine for identifying a few thousand dataset versions. If you are content-addressing millions of blobs, keep the full 64 characters.

Recording it with the experiment

The fingerprint earns its keep when it sits beside the score:

log_run.py
# continues the file above; `manifest` and the frames are already defined

def log_run(df, model_name, auc, note, path="runs.jsonl"):
    row = manifest(df, note) | {"model": model_name, "auc": auc}   # dict union needs Python 3.9+
    with open(path, "a") as f:
        f.write(json.dumps(row) + "\n")

log_run(v1, "logreg", 0.81, "baseline")
log_run(v2, "logreg", 0.77, "after the income typo fix")
print(open("runs.jsonl").read().strip())
Output
{"fingerprint": "a0d1573c38cf", "rows": 4, "columns": ["id", "city", "income", "repaid"], "null_counts": {"id": 0, "city": 0, "income": 0, "repaid": 0}, "label_rate": 0.75, "note": "baseline", "model": "logreg", "auc": 0.81}
{"fingerprint": "9ad1ba519e37", "rows": 4, "columns": ["id", "city", "income", "repaid"], "null_counts": {"id": 0, "city": 0, "income": 0, "repaid": 0}, "label_rate": 0.75, "note": "after the income typo fix", "model": "logreg", "auc": 0.77}

Two lines, and the investigation is over before it starts. The score fell from 0.81 to 0.77, the fingerprint changed, and the note says which cell moved.

Now a dropped score is a diff between two lines of runs.jsonl. Same fingerprint and a different score means the change is in your code or your seed. Different fingerprint and a different score means the data moved, and the manifest tells you which column.

That single distinction saves days. Without it, every regression is investigated from scratch.

The tools that do this properly

DVC stores a hash of each file in Git while the bytes live in object storage. git checkout of an old commit plus dvc checkout restores the exact data that commit referenced. It fits naturally when data is files.

Delta Lake, Apache Iceberg and Apache Hudi add transactional versioning to table storage. Each write creates a new snapshot, and you can query as of a version or a timestamp — SELECT * FROM t VERSION AS OF 42. This is the right shape when data is tables.

LakeFS applies Git semantics — branch, commit, merge — to an object store, so you can develop against a branch of production data.

Append-only design with a valid-from column is the version you can build in any database today. Never update a row; insert a new one with a timestamp, and read the latest row per key as of a chosen moment. This also gives you point-in-time correctness, which feature stores depend on.

Common mistakes

Versioning the model and not the data. Then a reproducible model is one you cannot reproduce, because its input is gone.

Naming files final_v2_new_FIXED.csv. Everyone has done it. It carries no information about contents and cannot be checked.

Hashing the file rather than the contents. Re-export the same table with a different column order and the file hash changes while the data has not. Hash a canonical form, as above.

Storing a fingerprint with no manifest. You learn that something changed, and nothing about what. The manifest is the useful half.

Mutating tables in place with no history. The single most common cause of "we cannot reproduce last quarter's model". Append, never overwrite.

Try it yourself

Add a fifth row to v1 and print the fingerprint. Then remove it again and confirm the fingerprint returns to a0d1573c38cf. Content addressing is reversible in a way that filenames never are.

Then write a diff_manifests(a, b) function that prints only the keys whose values differ. That tiny function is the one people actually use every day.

What to learn next

Researcher — Mathematics and papers.

Content addressing and Merkle structures

A content-addressed store keys objects by $h(\text{contents})$ for a collision-resistant hash $h$. Properties that follow directly: deduplication is automatic, integrity is verifiable without a trusted index, and equality is decidable in $O(1)$ after hashing.

For large datasets, hashing the whole object on every check is $O(n)$ per verification. Merkle trees (Merkle, 1987) reduce this. Chunk the data, hash each chunk as a leaf, and hash concatenated pairs upward to a root:

$$ h_{\text{node}} = h\big(h_{\text{left}} \,|\, h_{\text{right}}\big) $$

Verifying one chunk against a known root costs $O(\log n)$ hashes. Diffing two versions costs $O(k \log n)$ for $k$ changed chunks, rather than $O(n)$. This is the structure underneath Git, IPFS, ZFS and Delta Lake's file-level manifests.

Chunk boundaries matter. Fixed-size chunking suffers the boundary shift problem: an insertion of one byte at offset 0 changes every subsequent chunk. Content-defined chunking (Rabin fingerprinting) places boundaries where a rolling hash meets a condition, so an insertion affects only the local chunk. Expected chunk size is tuned by the number of low bits required to be zero.

Collision probability

With a $b$-bit digest and $k$ stored objects, the birthday bound gives:

$$ P(\text{collision}) \;\approx\; 1 - e^{-k^2 / 2^{b+1}} \;\approx\; \frac{k^2}{2^{b+1}} $$

For the truncated 48-bit fingerprint in the developer section and $k = 10^4$ dataset versions, this is around $1.8 \times 10^{-7}$. Acceptable for a human-facing label. For a content-addressed store where a collision means silent data loss, use the full digest.

Note that SHA-1 is no longer collision-resistant (Stevens et al., 2017, SHAttered), which is why Git has been migrating to SHA-256. Use SHA-256 for anything new.

Table formats and snapshot isolation

Delta Lake (Armbrust et al., 2020, VLDB) implements ACID transactions over object storage by writing an ordered log of JSON actions plus periodic Parquet checkpoints. A table version is a log offset. Readers resolve a snapshot by replaying the log to a version, giving serialisable snapshot isolation without a coordinating server, using atomic put-if-absent on the log file as the concurrency primitive.

Iceberg takes a different route: a metadata tree of manifest lists and manifest files, with an atomic pointer swap on the current metadata location. It handles very high partition counts better, since a scan plan reads manifests rather than listing directories.

Both give time travel — reading a table as of a version or timestamp — which is precisely the reproducibility primitive that ad-hoc CSV workflows lack.

Point-in-time correctness

Reproducibility of the inputs is necessary but not sufficient. For any model consuming time-varying features, the correct training row for an event at time $t$ uses feature values as known at $t$, not as known now.

Formally, given an event log $E$ and a feature computation $\phi$, the training feature must be $\phi\big({e \in E : e.\text{ts} \le t}\big)$. Computing $\phi$ over the whole table at training time is a temporal leak, and it is invisible to every fingerprint scheme, because the bytes are correct and the join is wrong.

This requires bitemporal modelling: distinguish valid time (when the fact was true in the world) from transaction time (when the system learned it). Snodgrass's work on temporal databases is the reference treatment; Kleppmann's Designing Data-Intensive Applications, chapter 11, gives the practical version. See feature stores for the machinery.

What a run record should contain

Minimum viable provenance, per training run:

  1. Dataset fingerprint, plus fingerprints of each upstream source.
  2. Manifest statistics: row and column counts, null rates per column, label distribution, and per-column min/max/mean for numerics.
  3. Code commit SHA, and the resolved dependency set — a lockfile hash, not a requirements range.
  4. All random seeds, including any framework-level global seed.
  5. Hardware and library versions where numerics matter, since floating-point reduction order varies with thread count and device.
  6. The transformation lineage: which function produced this table from which inputs.

Items 1 to 3 are cheap and almost always omitted. Item 6 is what W3C PROV formalises, and what lineage tools such as OpenLineage and Marquez collect automatically.

Reproducibility is a spectrum

Worth naming precisely, because teams claim the strongest form and deliver the weakest:

  • Repeatable — same team, same setup, same result.
  • Reproducible — different team, same artefacts, same result.
  • Replicable — different team, independently collected data, same conclusion.

Data versioning is necessary for the second and irrelevant to the third. Confusing them produces a great deal of misplaced confidence.

Reading

  • Armbrust et al., Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores, VLDB 2020.
  • Merkle, A Digital Signature Based on a Conventional Encryption Function, CRYPTO 1987.
  • Kleppmann, Designing Data-Intensive Applications, O'Reilly 2017 — chapters 3 and 11.
  • Stevens et al., The First Collision for Full SHA-1, CRYPTO 2017.
  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — on data dependencies as the costliest debt.

What to learn next

What to learn next

These follow on from what you just read.

  • Data Engineering for AI

    Feature stores

    A feature store keeps one definition of each feature for both training and live serving, which is the only reliable cure for a model that scores well offline and badly in production.

  • Data Engineering for AI

    ETL for machine learning

    ETL moves data from where it lands to where a model can use it, and the job that matters is making that move safe to run twice.

  • Data Engineering for AI

    Streaming data

    Streaming means handling data that never stops arriving, where events show up late and out of order, so any number you report is only true so far.