The one-page ML project brief
One page holding the decision, the prediction, the metric, the baseline, and the constraints — agreed before building — is the cheapest insurance an ML project can buy.
- 7 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A project brief is one page that states what will be predicted, for which decision, measured how, against what baseline — signed before any model is built.
Before a wedding card goes to print, someone checks one proof sheet: names, date, venue, time. Printing five hundred cards with the wrong date is expensive; reading one sheet is free. The proof sheet exists because fixing things before commitment is cheap.
An ML project brief is that proof sheet. Every scoping lesson so far produced a decision — this page is where those decisions get written down and agreed, while changing them still costs nothing.
Why it exists
Without a brief, each person carries a private version of the project. The engineer optimises accuracy; the manager expects fewer complaints; the analyst assumed weekly predictions when the ops team needed real-time ones. Everyone works hard, and the versions collide in month three.
The brief forces the collision into week one, on paper, where it is a conversation instead of a crisis.
How it works
Ten lines, each answering one question a lesson in this section asked:
DECISION what action will change?
PREDICTION what exactly is predicted, per what unit?
METRIC how is success measured?
TARGET what number makes this a success?
BASELINE what does the current approach score?
DATA which tables, how far back, audited when?
CONSTRAINT the hardest of the five doors
OWNER who decides, who builds?
REVIEW DATE when do we honestly check?The test of a good brief: a stranger reads it and can say what the project does. More valuable still, they can say what it does not do.
If the brief will not fit on one page, the project is not understood yet. That sentence sounds harsh and has saved a lot of money.
A real example you have seen
Every cricket scorecard is a brief of sorts: target, overs, required rate, on one screen. Nobody plays a chase without knowing the target. Teams build models without knowing theirs, surprisingly often — the brief is the scorecard that prevents it.
Remember this
- The brief is one page: decision, prediction, metric, target, baseline, data, constraint, owner, review date.
- Its job is making disagreements happen in week one, when they are cheap.
- No baseline line, no project — a target without a baseline is a wish.
What to learn next
- Getting knowledge out of a domain expert — several brief lines require an expert's input.
- Deciding a model is good enough — the review date's question, answered honestly.
- Model cards — the brief's post-deployment descendant.
Developer — Code and libraries.
Setup
No installation needed — the standard library is enough.
A brief that validates itself
Keeping the brief as data means the required fields are checkable, printable, and versionable in git next to the code it governs.
REQUIRED = ["decision", "prediction", "unit", "metric", "target_value",
"baseline", "data_source", "constraint", "owner", "review_date"]
brief = {
"decision": "Which support tickets get answered first each morning",
"prediction": "Probability this ticket concerns a payment failure",
"unit": "one support ticket, scored at the moment it arrives",
"metric": "recall on payment tickets at 20% review budget",
"target_value": "recall >= 0.85 (rule-based triage today: 0.61)",
"baseline": "keyword rule: contains 'refund', 'charged' or 'payment'",
"data_source": "tickets table, 14 months, ~48k rows, labels from agent tags",
"constraint": "must score a ticket in under 1 second on the current server",
"owner": "support operations lead (decides), ML team (builds)",
"review_date": "6 weeks after first deployment",
}
missing = [k for k in REQUIRED if not brief.get(k)]
if missing:
raise SystemExit(f"brief rejected, missing: {missing}")
width = max(len(k) for k in REQUIRED)
print("ML PROJECT BRIEF".center(70, "-"))
for key in REQUIRED:
print(f"{key.replace('_', ' '):>{width}} | {brief[key]}")
print("-" * 70)
print("one page. if it does not fit, the project is not understood yet.")---------------------------ML PROJECT BRIEF---------------------------
decision | Which support tickets get answered first each morning
prediction | Probability this ticket concerns a payment failure
unit | one support ticket, scored at the moment it arrives
metric | recall on payment tickets at 20% review budget
target value | recall >= 0.85 (rule-based triage today: 0.61)
baseline | keyword rule: contains 'refund', 'charged' or 'payment'
data source | tickets table, 14 months, ~48k rows, labels from agent tags
constraint | must score a ticket in under 1 second on the current server
owner | support operations lead (decides), ML team (builds)
review date | 6 weeks after first deployment
----------------------------------------------------------------------
one page. if it does not fit, the project is not understood yet.The walkthrough
Every line answers a question a stakeholder will ask later, in writing, in advance. "Why is recall the metric?" has an answer dated before the model existed. That timestamp is worth more than the answer.
The baseline line carries the most information per word. Today's rule scores 0.61. The target is 0.85. The project's entire value is that gap — and if the audit later shows the gap is unreachable, the brief is where the project dies cheaply.
metric includes the budget. Not "recall" but "recall at 20% review budget". A metric without its operating condition is unfalsifiable — the lesson on metrics explains why the condition matters as much as the metric.
The validation is real. Delete the owner line and the script refuses to print. A brief with an unfilled field is a disagreement postponed, and this file makes postponement impossible.
Common mistakes
Writing the brief after building the model. It becomes a justification document, and the baseline mysteriously ends up beatable. The brief's value is entirely in its timing.
Skipping the review date. Without a named date, no one ever asks "did this work?" — projects then live forever on the strength of their launch announcement. Six weeks after deployment is a reasonable default.
Committing to precise numbers before the feasibility probe. "recall >= 0.85" should follow the feasibility check, not precede it. A target invented in a meeting is a hostage to fortune.
Treating the brief as fixed. Reality updates it — a data audit finding, a constraint change. Update the file in git; the diff history is the project's decision log.
Try it yourself
Write the brief for a project you know — real or imagined. The exercise bites at exactly one line for most people: baseline. If you cannot state what the current approach scores, that measurement is your genuine next task, before any model.
What to learn next
- Getting knowledge out of a domain expert — several brief lines require an expert's input.
- Deciding a model is good enough — the review date's question, answered honestly.
- Model cards — the brief's post-deployment descendant.
Researcher — Mathematics and papers.
The brief as pre-registration
The brief is the engineering sibling of pre-registration in experimental science: hypotheses, metrics, and analysis plans declared before data collection, to prevent hypothesising after results are known (HARKing) and metric shopping. The replication-crisis literature (Nosek et al., 2018, The Preregistration Revolution, PNAS) documents the failure mode the brief guards against — in ML terms, evaluating on many metrics and reporting the flattering one, or adjusting the baseline until the model clears it.
ML-specific instantiations:
- Model cards (Mitchell et al., 2019, FAT*) — post-hoc documentation of intended use, metrics, and limitations; the brief is its pre-hoc counterpart, and a finished project should be able to generate its model card from brief plus results.
- Datasheets for datasets (Gebru et al., 2021, CACM) — the
data_sourceline, expanded to a full questionnaire. - REFORMS (Kapoor et al., 2024, Science Advances) — a consensus reporting checklist for ML-based science; a project with a maintained brief satisfies most of it by construction.
Why written commitment changes outcomes
Two mechanisms, both studied:
- Metric gaming resistance. With the metric fixed ex ante, Goodhart pressure is at least visible: any post-hoc metric change requires editing a versioned document. Auditability substitutes for enforcement.
- Coordination cost. Amershi et al. (2019, ICSE) find requirement instability among the top process problems in ML teams; a single-page shared artifact reduces the per-stakeholder divergence that verbal specs permit. The one-page limit is a forcing function against specification sprawl — bounded documents get read, and reading is the mechanism.
The review-date line and survivorship
Deployed models are rarely re-evaluated against their original claims: post-deployment monitoring tends to track operational health, not the founding hypothesis. The review_date field institutionalises a falsification checkpoint. Connect it to the deployed system's monitoring — monitoring and drift — so the review compares the promised metric on live data against target_value, closing the loop pre-registration opened.
What to learn next
- Getting knowledge out of a domain expert — several brief lines require an expert's input.
- Deciding a model is good enough — the review date's question, answered honestly.
- Model cards — the brief's post-deployment descendant.