Laying out an ML project
A predictable folder layout — raw data kept sacred, code as a package, configs and runs separated — is what lets another person, or future you, run your project at all.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A good project layout means anyone can find anything without asking you — including you, six months from now.
Walk into a well-run kitchen. Knives hang in one place, spices line one shelf, raw vegetables never touch the cooked-food counter. A new cook finds the salt without asking, because everything lives where a reasonable person would look. The organisation is information.
A messy ML project is the opposite kitchen: final_v2_REAL.py beside test_new.py, data files scattered among scripts, results pasted into a chat. Every question needs the original cook — and the original cook has forgotten too.
Why it exists
ML projects have more kinds of things than ordinary software: code, raw data, cleaned data, configs, trained models, run results. Without assigned places, these mix — and the mixtures are poisonous. Code edits sneak into data folders. Cleaned files overwrite raw ones. Nobody knows which model came from which run.
One rule matters above all: raw data is sacred. The files you received are never edited, never overwritten. Everything derived from them goes elsewhere, and can therefore always be rebuilt.
How it works
project/
├── README.md ← how to run, in one screen
├── configs/ ← every experiment's knobs
├── data/
│ ├── raw/ ← sacred: never edited, never overwritten
│ └── processed/ ← rebuildable: derived from raw by code
├── src/<package>/ ← the actual code: data, features, train, evaluate
├── tests/ ← checks that code AND data behave
└── runs/ ← results, one folder per run: config + metrics + modelThe flow runs one direction: raw data → processing code → processed data → training → runs. Nothing flows backwards, so any later folder can be deleted and rebuilt.
A real example you have seen
Hospitals archive the original scan and write reports as separate documents — nobody annotates the only X-ray with a marker pen. Banks never edit a passbook entry; corrections are new entries. Any system that must survive mistakes keeps originals untouchable and derivations separate. Your data/raw/ is the passbook.
Remember this
- One assigned place per kind of thing: code, configs, raw data, processed data, runs.
- Raw data is sacred — derived files are disposable because code can rebuild them.
- If running the project needs knowledge that is only in your head, the layout is not done.
What to learn next
- Getting out of the notebook — moving exploration code into src/ without losing momentum.
- Configs instead of constants scattered everywhere — what fills configs/ and runs/.
- Asserting on your data before you train — what fills tests/test_data.py.
Developer — Code and libraries.
Setup
pip install scikit-learnVerified on Python 3.10, Windows and Linux. The snippet below creates a real skeleton on disk, so run it in a scratch directory.
A skeleton you can generate and inspect
import pathlib
skeleton = {
"README.md": "# churn\nHow to train: python -m churn.train\n",
"configs/baseline.json": '{"model": "logistic", "seed": 0}\n',
"data/raw/.gitkeep": "",
"data/processed/.gitkeep": "",
"src/churn/__init__.py": "",
"src/churn/data.py": "# loading and cleaning lives here\n",
"src/churn/features.py": "# every feature, one function each\n",
"src/churn/train.py": "# the only file that trains\n",
"src/churn/evaluate.py": "# metrics, never mixed into training\n",
"tests/test_data.py": "# assertions about the data itself\n",
"runs/.gitkeep": "",
}
root = pathlib.Path("churn-project")
for path, text in skeleton.items():
full = root / path
full.parent.mkdir(parents=True, exist_ok=True)
full.write_text(text)
for p in sorted(root.rglob("*"), key=lambda q: str(q).lower()):
depth = len(p.relative_to(root).parts) - 1
print(" " * depth + p.name)configs
baseline.json
data
processed
.gitkeep
raw
.gitkeep
README.md
runs
.gitkeep
src
churn
__init__.py
data.py
evaluate.py
features.py
train.py
tests
test_data.pyThe walkthrough
src/churn/ is a package, not a script pile. Code lives in an importable package with a name, run as python -m churn.train. This one decision kills the sys.path hacks and from utils import * confusion that plague script piles, and it makes the code testable — tests/ imports churn like any consumer.
Four modules, four verbs. data.py loads and cleans. features.py transforms. train.py fits. evaluate.py measures. When a file grows a second verb, split it. The point is not purity — it is that a reader with a question knows which file answers it.
.gitkeep files are empty markers so git tracks the empty folders — git ignores folders with nothing in them. Meanwhile the contents of data/ and runs/ should be git-ignored (large, regenerable, or private), even though the folders themselves ship with the repo.
runs/ completes the story from configs and run records: one folder per run, holding the resolved config, the metrics, and the model file. The repo answers "what produced this number?" without archaeology.
The README's job is one screen: what this is, how to install, the one command that trains. Every question you answer twice in chat becomes a README line.
What did not make the skeleton, and when it does
notebooks/— add it when you have notebooks; keep them for exploration, promote working code intosrc/(the next lesson is exactly this).scripts/— one-off utilities (download, export) that are not part of the package.Dockerfile,requirements.txt/lockfile — the environment layer; see Docker for ML.- Model registry, orchestration — when runs outgrow folders, graduate to MLflow and friends. The layout above is what makes that graduation painless.
Common mistakes
Editing raw data "only this once". The original is gone forever and every downstream result becomes unverifiable. If raw data has errors, code fixes them on the way to processed/ — visibly, repeatably.
Paths hardcoded to one machine. C:/Users/you/Desktop/data.csv breaks for everyone else. Paths belong in configs, relative to the project root.
A utils.py that eats the project. Grab-bag modules grow until every import touches them. Name modules for what they do; when you cannot name it, you do not understand it yet.
Committing the 2 GB processed dataset. Git holds code and configs; data lives in storage with data versioning pointing at it.
Try it yourself
Take a current messy project and only move files into this shape — no code edits beyond fixing imports and paths. Keep a list of every path you had to fix. That list is a measurement of how much hidden machine-specific knowledge your project was carrying.
What to learn next
- Getting out of the notebook — moving exploration code into src/ without losing momentum.
- Configs instead of constants scattered everywhere — what fills configs/ and runs/.
- Asserting on your data before you train — what fills tests/test_data.py.
Researcher — Mathematics and papers.
The debt this layout services
Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems, NeurIPS — the canonical taxonomy of how ML systems rot: boundary erosion, entangled data dependencies, "pipeline jungles", configuration debt, and undeclared consumers. The layout above is a direct counter-position: data flowing one direction through named stages attacks pipeline jungles; configs as first-class artefacts attack configuration debt (which the paper ranks among the most underestimated); the sacred-raw rule makes data dependencies explicit and rebuildable.
Amershi et al. (2019), Software Engineering for Machine Learning: A Case Study, ICSE-SEIP — Microsoft's field study: the three dominant ML-specific pain points were data discovery/management/versioning, reuse, and the fact that ML components entangle in ways software modules do not. Layout conventions are the cheapest mitigation available to a small team: they impose module boundaries (Sculley's "strong abstraction boundaries") by directory fiat.
Convention as coordination
The specific shape matters less than its predictability — the value is a Schelling point. Cookiecutter Data Science (Driven Data, 2016 onward) popularised near-exactly this layout and articulated its two load-bearing rules: data is immutable and analysis is a DAG (raw → processed → results, no back-edges). The DAG property is what makes processed/ and runs/ disposable, which in turn is what makes reproduction possible: the repo's state is fully determined by (raw data, code, configs) — the provenance coordinates from the config lesson.
The src/-layout (package under src/, not at repo root) is Python-packaging best practice for a subtle reason: it forces tests to run against the installed package rather than the working directory, catching packaging errors and import-order accidents — pytest and packaging documentation both recommend it.
Testing surface
The tests/ folder in an ML repo carries three distinct test kinds, per Breck et al. (2017), The ML Test Score: A Rubric for ML Production Readiness, IEEE Big Data:
- Code tests — ordinary unit tests of
features.pyand friends. - Data tests — schema, ranges, distributions of what enters training (asserting on your data).
- Model tests — behavioural checks on trained artefacts (invariances, minimum quality bars).
Their rubric's empirical finding: teams systematically over-invest in (1) and under-invest in (2) and (3), while production incidents concentrate in the latter. A layout with tests/test_data.py present from day one is a nudge against that gradient; the CI side is covered in CI/CD for ML.
Monorepo vs. per-project, briefly
At organisation scale the per-project layout meets the monorepo question: shared feature code, shared data access layers, many models. The published positions (Google's monorepo experience; feature stores as the anti-entanglement answer to shared features) sit beyond this lesson; the invariant that survives every scale is the one-directional flow raw → code → derived → results, with each stage addressable and rebuildable.
What to learn next
- Getting out of the notebook — moving exploration code into src/ without losing momentum.
- Configs instead of constants scattered everywhere — what fills configs/ and runs/.
- Asserting on your data before you train — what fills tests/test_data.py.