Scoping an ML Project

Turning a vague request into a prediction task

Requests like "reduce churn" cannot be trained on directly — framing means choosing the unit, the moment of prediction, and the window that defines the label.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Framing means converting a wish like "reduce churn" into one exact question about one thing at one moment.

Think of a tailor. A customer says "make me look sharp" — a wish, not an instruction. The tailor converts it into numbers: chest 40, sleeve 24, delivery Friday. Only then can cutting begin.

Machine learning needs the same conversion. A model cannot train on "reduce churn". It can train on "for this customer, today, predict: will they order again within 60 days?".

Why it exists

Business people speak in outcomes: less fraud, happier users, lower costs. Models speak in prediction tasks: given one example, guess one value. Someone has to build the bridge, and that bridge is where most ML projects quietly succeed or fail.

A badly framed task can be trained perfectly and still be useless. "Predict monthly revenue" sounds reasonable, but nobody can act on it. "Predict which 100 customers are most likely to leave this month" hands the retention team a call list.

How it works

Every framing answers three questions:

UNIT     what is one example?         → one customer
MOMENT   when do we predict?          → the 1st of each month
WINDOW   what makes the label true?   → no order in the next 60 days

"reduce churn"  →  one row per customer, scored monthly,
                   label = ordered again within 60 days

The unit is the thing each row describes. The moment is the instant of prediction — features may only use information from before it. The window is the future period the label is read from.

Move any of the three and you have a different project. Predicting per-order instead of per-customer, or 7 days instead of 60, changes the data, the model, and who can use the output.

A real example you have seen

"Make YouTube more engaging" became, somewhere inside Google, "given this user and this video, predict watch time". The wish is vague. The task is exact. Every recommendation you see flows from that one framing choice — including its side effects.

Remember this

  • Models train on tasks, not wishes. Framing is the conversion.
  • Every framing fixes three things: the unit, the moment, and the window.
  • A perfect model of the wrong framing is useless. Framing errors outrank tuning errors.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Outputs verified with pandas 2.2.

From raw orders to a trainable table

The raw material is an event log. The framing decisions become code: a cutoff date, a feature window before it, a label window after it.

frame_churn.py
import pandas as pd

orders = pd.DataFrame({
    "customer": ["asha", "asha", "asha", "bilal", "bilal",
                 "chen", "deepa", "deepa", "deepa", "evan"],
    "order_date": pd.to_datetime([
        "2025-03-02", "2025-05-21", "2025-06-20", "2025-02-14", "2025-04-30",
        "2025-05-28", "2025-01-05", "2025-02-11", "2025-07-03", "2025-03-19"]),
    "amount": [499, 1250, 799, 2100, 460, 999, 350, 620, 410, 1899],
})

# The moment of prediction. Everything splits around this line.
cutoff = pd.Timestamp("2025-06-01")

past = orders[orders["order_date"] < cutoff]
window = orders[(orders["order_date"] >= cutoff) &
                (orders["order_date"] < cutoff + pd.Timedelta(days=60))]

# Features: only what was knowable BEFORE the cutoff.
X = past.groupby("customer").agg(
    n_orders=("amount", "size"),
    total_spent=("amount", "sum"),
    last_order=("order_date", "max"),
)
X["days_since_last"] = (cutoff - X.pop("last_order")).dt.days

# Label: did they order again in the 60 days AFTER the cutoff?
X["bought_again"] = X.index.isin(window["customer"]).astype(int)

print(X)
Output
          n_orders  total_spent  days_since_last  bought_again
customer                                                      
asha             2         1749               11             1
bilal            2         2560               32             0
chen             1          999                4             0
deepa            2          970              110             1
evan             1         1899               74             0

The walkthrough

The cutoff line is the whole design. Features come strictly from before it, the label strictly from after. Cross that line even once and you get leakage — the model reads the future during training, then fails in production where the future is unavailable.

Asha's June 20 order appears only in the label. She ordered after the cutoff, so that order proves bought_again = 1 but contributes nothing to n_orders. This is deliberate. At prediction time, that order has not happened yet.

Deepa is the row worth staring at. Silent for 110 days, then returned. Real labels contain surprises like this; a framing that seems watertight in your head meets reality in the table.

One cutoff gives one training snapshot. Production framings slide the cutoff monthly, stacking snapshots so the model sees many moments, not one.

Common mistakes

Predicting something nobody can act on. "Predict total revenue next quarter" has no per-row action. Frame around the decision: who gets a call, which order gets held, which ticket goes first.

Letting the unit drift mid-project. Features per customer, labels per order, evaluation per session — three units, one silent mess. Write the unit down and check every table against it.

A label window nobody waited for. A 60-day window means every training example is at least 60 days old, and a live prediction cannot be verified for 60 days. Teams discover this at deployment. Decide upfront whether the business can wait that long for feedback.

Using "now" as the cutoff for training data. Recent customers have not had their 60 days yet, so their labels read as 0 unfairly. Training cutoffs must be at least one full window in the past.

Try it yourself

Reframe the same table for a different wish: "increase order value". Unit stays the customer, but the label becomes the amount of their next order — a regression task. Which rows now have no label at all, and what does that force you to decide?

What to learn next

Researcher — Mathematics and papers.

Framing as estimand selection

A framing declares the estimand before any estimator is chosen. Formally, pick:

$$ (U, t, W, Y) \quad \Rightarrow \quad \text{learn } f: \mathcal{X}{U, <t} \to \hat{Y}{[t,\, t+W)} $$

Where:

  • $U$ — the unit of analysis (customer, order, session).
  • $t$ — the prediction time; the filtration $\mathcal{X}_{U,<t}$ contains only information measurable before $t$.
  • $W$ — the label horizon.
  • $Y$ — the outcome functional read from $[t, t+W)$.

The filtration constraint is the formal statement of "no future in the features". Violations are the leakage family — target leakage, temporal leakage, train-test contamination — each with its own detection pattern.

Prediction versus intervention

Most business wishes are causal ("reduce churn"), while the framed task is predictive ("who will churn"). The gap matters: the highest-risk customers are not necessarily the most persuadable ones. Targeting by risk score optimises the wrong quantity when the action's effect varies across customers.

  • Ascarza (2018), Retention Futility (Journal of Marketing Research), shows empirically that targeting high-churn-risk customers underperforms targeting high-uplift customers.
  • The corrected framing is uplift modelling: estimate $\mathbb{E}[Y(1) - Y(0) \mid X]$, the conditional average treatment effect, rather than $\mathbb{E}[Y \mid X]$. The potential-outcomes notation $Y(1), Y(0)$ denotes the outcome under treatment and under control.

Horizon choice as a bias-variance trade

Short windows give fresh labels, more training snapshots, and faster feedback, but noisier outcomes (one missed week reads as churn). Long windows stabilise the label while delaying every iteration of the project by $W$. There is no free choice: $W$ appears in both label quality and organisational learning rate. Fader and Hardie's BTYD ("buy till you die") literature (2005 onward) treats churn in non-contractual settings as a latent state, a reminder that the observable label is itself a proxy.

Process references

  • Chapman et al. (2000), CRISP-DM 1.0 — the original business-understanding → data-understanding loop; dated tooling, durable structure.
  • Amershi et al. (2019), Software Engineering for Machine Learning (ICSE), documents Microsoft's nine-stage workflow and finds problem formulation among the most error-prone stages.
  • Passi and Barocas (2019), Problem Formulation and Fairness (FAT*), show that framing choices — not model choices — carry most of the normative weight in deployed systems.

What to learn next