ML case-study interviews
Case-study rounds hand you a vague business problem and score your diagnosis — here is the eight-step script, one fully worked case, and the traps that fail candidates.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A case-study interview hands you a messy business problem and watches how you turn it into a plan — where machine learning might be one step, or none.
Think of visiting a doctor with "my stomach hurts". A good doctor does not reach for the prescription pad at once. They ask questions first — since when, after which meals, does it hurt here or there? The diagnosis is the skill; the medicine comes after.
This round works the same way. The interviewer says something vague, like "cancellations went up 20%". Candidates who prescribe a model in the first minute fail. Candidates who diagnose pass.
Why this round exists
Real ML work never arrives as "please train a classifier on this clean dataset". It arrives as a worried sentence from a manager. Someone must turn that sentence into a measurable problem, decide whether ML even helps, and prove the fix worked.
Companies once hired brilliant model-builders who could not do that translation. Their projects solved the wrong problem beautifully. This round exists to filter for the translation skill.
The shape of a good answer
"cancellations are up 20%"
│
▼
ask questions since when? which city? which users? measured how?
│
▼
slice the numbers is it everyone, or one segment dragging the average?
│
▼
name the goal one number to improve, one number that must not get worse
│
▼
try the small fix a rule, a UI change, a phone call — before any model
│
▼
model, if needed and say exactly what it predicts, for whom, to do what
│
▼
prove it worked compare against not doing itThe surprise for most engineers: reaching "no model needed, here is the smaller fix" can be a winning answer, when you show how you got there.
A real example you have seen
Food delivery apps show an estimated delivery time before you order. Behind that feature was once exactly this conversation — complaints about late orders, someone slicing the data, and a decision that a prediction shown at order time would help more than faster scooters.
Remember this
- Diagnose before prescribing. Questions first, model last.
- One success number, plus one guardrail number that must not get worse.
- "Here is the cheaper non-ML fix" is a strong answer, not a dodge.
What to learn next
- Take-home assignments that pass — the round where this thinking goes into a repo.
- Model evaluation — the offline half of every validation plan.
- A/B testing recommenders — the online half.
Developer — Code and libraries.
The eight-step script
Memorise the skeleton, not answers. It holds for any case.
- Clarify. Since when? Measured how? Absolute numbers or percentages? What changed recently — pricing, app release, a new city?
- Segment. Averages hide everything. Slice by city, platform, user age, restaurant, time of day. A 20% overall rise is usually one segment on fire.
- Define the metric. One number to move, and a guardrail: reduce cancellations without slowing delivery promises or hurting order volume.
- Hypothesise causes. List them out loud: longer delivery times, wrong time estimates shown, restaurant stock-outs, a payment bug, fraud.
- Baseline fix first. What would you do with no model at all? A rule — "stop promising 30 minutes during rain" — sets the bar any model must beat.
- Formulate the model, precisely. Target, prediction time, features available at that moment, label source. Vague model talk fails; precision passes.
- Validation plan. Offline first against the baseline — see model evaluation — then an A/B test with the guardrail watched.
- Decision economics. Actions cost money. Show the threshold where acting is worth it — worked below.
The case, walked end to end
"Cancellations on our food delivery app are up 20% this quarter. What do you do?"
Clarify. Suppose the interviewer answers: up from 5% to 6% of orders, started six weeks ago, self-reported reasons dominated by "taking too long".
Segment. Ask for cancellation rate by city, hour, and weather. Suppose evening orders in three monsoon-hit cities carry the whole rise. The problem has a shape now — it is not "cancellations", it is "rain-hour promises".
Metric. Reduce evening cancellation rate in those cities toward the old 5%, holding order volume and delivery rating steady.
Hypotheses. The app promises 30 minutes; monsoon traffic makes it 55. Customers cancel when the promise breaks, not when delivery is slow — a testable distinction, and the data above supports it.
Baseline. Add rain and hour to the promised-time rule before training anything. If showing "45–55 min" honestly halves the problem, the model's bar moved and everyone learned something for nothing.
Model, precisely stated. At order time, predict expected delivery minutes. Features available at that moment: restaurant's rolling preparation time, rider availability nearby, distance, hour, live rain. Label: actual delivery minutes from historical orders — a regression, evaluated on recent weeks, compared against the rule.
Careful with the label. "Cancellation" bundles customer cancellations, restaurant rejections, and rider no-shows. Those are three different problems with three different owners. Splitting a label into its real causes is precisely the kind of move this round rewards — it comes from feature engineering instincts applied to targets, not features.
Validation. Backtest the predictor on held-out recent weeks. Then A/B: new estimates for half the affected cities, watching cancellations and the guardrails, because honester, longer estimates might cost some orders — that trade is a business decision to surface, not to hide.
The economics step, with real arithmetic
A second mini-case shows step 8, because it wins rounds. A telecom churn model scores 10,000 customers. Marketing wants to send retention vouchers. Who gets one?
python3 --version # standard library only# 10,000 customers, grouped by the churn model's score band.
# Counts and rates are invented; the arithmetic is the real content.
bands = [
# (score band, customers, true churn rate inside the band)
("0.9-1.0", 300, 0.62),
("0.7-0.9", 700, 0.41),
("0.5-0.7", 1500, 0.22),
("0.0-0.5", 7500, 0.05),
]
offer_cost = 200 # rupees per retention voucher sent
saved_value = 2600 # margin kept when a churner stays
accept_rate = 0.30 # fraction of true churners a voucher convinces
cum_n = 0
cum_saved = 0.0
for band, n, churn_rate in bands:
cum_n += n
cum_saved += n * churn_rate * accept_rate * saved_value
spend = cum_n * offer_cost
net = cum_saved - spend
print(f"target down to {band}: contact {cum_n:5d} spend {spend:9,} net {net:10,.0f}")target down to 0.9-1.0: contact 300 spend 60,000 net 85,080 target down to 0.7-0.9: contact 1000 spend 200,000 net 168,940 target down to 0.5-0.7: contact 2500 spend 500,000 net 126,340 target down to 0.0-0.5: contact 10000 spend 2,000,000 net -1,081,160
Read the net column: profit peaks when you contact the top 1,000 customers and falls if you go deeper, because vouchers to low-risk customers are almost pure cost. Contacting everyone loses eleven lakh rupees. The model's job was never "predict churn" — it was "rank customers so this column has a peak, and tell us where the peak is".
One honest caveat to say aloud: this targets by churn risk, and the customer most likely to leave is not always the one a voucher can persuade. The researcher block makes that precise.
Common mistakes
Prescribing in minute one. "I'd build an XGBoost model" before any clarifying question is the single most common fail.
A metric with no guardrail. Cancellations can be zeroed by cancelling the app. Every success metric needs its counterweight.
Skipping the baseline. A model that cannot beat a two-line rule is negative value. Reviewers of real projects and interviewers alike check this first.
Correlation as cause. Cancellations rose while a new app version rolled out — segment by version before blaming the weather.
"Collect more data" as an answer. Say which data, how, and what decision it changes. Otherwise it is a shrug wearing a lab coat.
Try it yourself
Take the case "support tickets doubled in a month" and write the eight steps in twenty minutes, including one piece of arithmetic like the voucher table. Then read it as a sceptical interviewer and mark every claim you failed to justify.
What to learn next
- Take-home assignments that pass — the round where this thinking goes into a repo.
- Model evaluation — the offline half of every validation plan.
- A/B testing recommenders — the online half.
Researcher — Mathematics and papers.
Three pieces of theory turn a good case answer into a senior one.
The action threshold, derived
Contacting a customer with churn probability p has expected value:
E[profit] = p · a · v − c ⇒ contact when p > c / (a · v)p— calibrated probability this customer churns without intervention.a— probability the intervention works on a true churner (the accept rate).v— value retained when it works;c— cost of the intervention.
With the lesson's numbers, p > 200 / (0.3 × 2600) ≈ 0.256. Check it against the table: the 0.7–0.9 band truly churns at 0.41, above threshold, and the 0.5–0.7 band at 0.22, below it — which is exactly where the profit peak fell. Two interview-grade observations follow. First, the decision used each band's true churn rate, not its score range, because this model is miscalibrated: customers scored 0.5–0.7 churn at 0.22. The threshold rule needs a calibrated model — among customers scored 0.3, about 30% must actually churn. Second, the threshold moves when marketing changes the voucher: the model is fixed, the economics are not.
Risk is not persuadability: uplift
Targeting by churn probability optimises the wrong quantity. The voucher's true value on customer x is the uplift, a causal quantity:
τ(x) = E[ Y(1) − Y(0) | X = x ]Y(1),Y(0)— the outcome (stays or leaves) with and without the voucher; only one is ever observed.τ(x)— how much the voucher changes this customer's behaviour.
The classic quadrant: persuadables (leave without the voucher, stay with it — the only group worth paying for), sure things, lost causes, and sleeping dogs, whom the contact itself pushes to leave. A high-risk customer can be a lost cause; targeting by risk pays for them anyway. Estimating τ(x) needs randomised holdouts — deliberately not contacting a random slice — and models built for heterogeneous treatment effects. Radcliffe and Surry (2011), Real-World Uplift Modelling with Significance-Based Uplift Trees, is the practitioner's reference; Künzel et al. (2019) formalise meta-learners for the same estimand.
Prediction problems versus causal problems
Kleinberg et al. (2015), Prediction Policy Problems, draws the line every case answer should respect. Some decisions need only accurate prediction — the delivery-time estimate is one, since showing a better forecast requires no causal claim. Others need causal answers — "will this voucher change behaviour" cannot be read from observational data, because treated and untreated customers were never comparable. Naming which kind of problem the interviewer handed you, unprompted, is among the strongest senior signals this round can carry.
A last honesty note: observational product data is soaked in confounding. Customers who received past vouchers differ from those who did not; monsoon cities differ from dry ones. When the case demands a causal claim, say the words "randomised test" before anyone else does.
References
- Kleinberg, Ludwig, Mullainathan and Obermeyer, Prediction Policy Problems, American Economic Review, 2015.
- Radcliffe and Surry, Real-World Uplift Modelling with Significance-Based Uplift Trees, Stochastic Solutions white paper, 2011.
- Künzel, Sekhon, Bickel and Yu, Metalearners for Estimating Heterogeneous Treatment Effects, PNAS, 2019.
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, 2020 — the standard A/B testing reference, including guardrail metrics and novelty effects.
What to learn next
- Take-home assignments that pass — the round where this thinking goes into a repo.
- Model evaluation — the offline half of every validation plan.
- A/B testing recommenders — the online half.