ML Interview Preparation

ML system design interviews

A repeatable skeleton for the 45-minute ML system design round, walked through twice in full — a feed ranker and a UPI fraud detector.

On this page 5
  1. Why this round exists
  2. The shape of the hour
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An ML system design interview asks you to plan a whole machine-learning product on a whiteboard — data, model, serving, and what happens when it breaks.

Think about cooking for a wedding instead of cooking dinner. A cook who makes wonderful dal for four people cannot automatically feed two thousand guests. The wedding needs planning: how much rice, how many burners, what happens when a gas cylinder empties mid-service. Same dish, completely different problem.

This round tests the wedding, not the dal. Training one model in a notebook is dinner. The interview asks whether you can run the kitchen.

Why this round exists

Most ML projects that fail in companies do not fail because the model was weak. They fail around the model. The data arrived late, nobody defined what "better" meant, the model served predictions from stale features, and nobody noticed for a month.

Companies learned this the expensive way. So they added a round that tests the surroundings: where data comes from, how the model reaches users, and how you find out it is misbehaving.

The shape of the hour

Interviewers follow you through roughly this path, and you can lead it yourself:

 what exactly are we building, for whom?   (requirements)
        │
        ▼
 what number defines success?              (metric)
        │
        ▼
 how big is this?                          (rough load, rough cost)
        │
        ▼
 where do labelled examples come from?     (data)
        │
        ▼
 what model, and why this one?             (model)
        │
        ▼
 how do predictions reach users fast?      (serving)
        │
        ▼
 how do we find out it broke?              (monitoring)

Notice the model is one box out of seven. Candidates who spend forty minutes on it fail rounds they could have passed.

A real example you have seen

Open any social or shopping app. The order of posts, reels, or products was chosen by a system like the ones in this lesson. Somebody designed the whole pipeline behind that ordering — and was very likely asked to do it in an interview first.

Remember this

  • The round tests the system around the model, not the model alone.
  • Rough numbers said out loud beat perfect numbers never mentioned.
  • The strongest closing move is describing how you would detect your own failure.

What to learn next

Developer — Code and libraries.

The skeleton, and the calculator

Walk every design through the seven steps above, in order, out loud. Two habits separate strong candidates: they ask questions before drawing, and they do arithmetic in front of the interviewer.

The arithmetic is easier than people fear. This is the entire feed-ranker envelope:

bash
python3 --version   # standard library only, nothing to install
envelope.py
dau = 10_000_000                 # daily active users
feed_loads = 6                   # feed opens per user per day
candidates = 500                 # posts scored per feed open

requests_per_day = dau * feed_loads
scores_per_day = requests_per_day * candidates
peak_qps = requests_per_day / 86_400 * 3     # traffic is spiky; 3x the mean is a fair peak

ms_per_item = 0.05               # a small ranker, batched, on CPU
cpu_seconds = scores_per_day * ms_per_item / 1000
machines = cpu_seconds / 86_400  # one core gives 86,400 CPU-seconds per day

print(f"feed requests/day : {requests_per_day:,}")
print(f"items scored/day  : {scores_per_day:,}")
print(f"peak QPS          : {peak_qps:,.0f}")
print(f"scoring cores     : {machines:.0f} (single-core equivalents)")
Output
feed requests/day : 60,000,000
items scored/day  : 30,000,000,000
peak QPS          : 2,083
scoring cores     : 17 (single-core equivalents)

Every input is an assumption, and that is the point: say each assumption out loud, and the interviewer will correct any they dislike. Thirty billion scores a day is the number that shapes the whole design — it is why the expensive model cannot see every post.

Design 1: a feed ranker

Requirements to ask before drawing anything. What surface — home feed, reels, notifications? What does the business mean by "good" — taps, watch time, or longer-term return visits? How fresh must content be — is a 5-minute-old post eligible? What is the latency budget — a feed open should complete in about 200 ms. Cold start — what do brand-new users see?

Metric. Offline: ranking quality on held-out interactions, covered in ranking metrics. Online: the business metric via an A/B test, with guardrails such as reports and hides. Never present one without the other.

Architecture.

 OFFLINE
 interaction logs ─► label builder ─► training data ─► daily training ─► model store
       ▲                                                                    │
       │                                                                    ▼
 ONLINE│                                                          (models loaded by servers)
 user opens feed ─► candidate generation ─► feature fetch ─► ranker ─► business rules ─► feed
                    follows, popular,       (feature store)  scores     diversity,
                    similar-user posts:                      500 posts  freshness,
                    ~500 candidates                                     ads slots

Why two stages. Scoring millions of posts per request is impossible inside 200 ms. Candidate generation cheaply narrows millions to hundreds — typically with two-tower models, where user and post embeddings are compared by fast nearest-neighbour lookup. The expensive ranker then scores only those hundreds. Say this unprompted; it is the single most expected sentence in the round.

Data and training loop. Labels are implicit: taps, watch time, hides. Two honest problems to raise yourself. First, position bias — users tap what was shown at the top, so logs flatter whatever the old model ranked highly. Second, the feedback loop — the model learns from clicks it caused, which narrows what users see; see diversity and filter bubbles. Retrain daily; features must be computed identically offline and online, which is the reason feature stores exist.

Failure modes. Training-serving skew — a feature computed one way in training and another way live. Stale features after a pipeline delay. A celebrity post creating a hot key that hammers one cache shard. A new model that wins offline and loses online — which is why rollout goes through an A/B test at 1%, then 10%, as in A/B testing recommenders.

What interviewers probe. "Why two stages?" — cost arithmetic, above. "How do you know a new model is better?" — offline metrics gate the A/B test; the A/B test decides. "What breaks at 10x users?" — candidate generation and the feature store first; the ranker scales out horizontally.

Design 2: a UPI fraud detector

A payments company wants to catch fraudulent transactions as they happen.

Requirements to ask. Decision must be inline — the payment waits on you, so the budget is roughly 100 ms end to end. What actions exist — allow, ask for an extra OTP, or block? What does a mistake cost each way? Blocking a genuine payment burns trust; missing fraud costs money directly. The action set matters more than the model: an extra-OTP middle option makes borderline scores survivable.

The envelope, spoken aloud. Assume 50 million transactions a day — about 580 per second on average, and several thousand at festival peaks. Assume fraud is 2 in 10,000 — that is 10,000 fraud attempts a day hiding in 50 million. Flag even 1% for step-up checks and you interrupt 500,000 genuine-looking payments daily. Those three numbers drive everything below.

Metric. Accuracy is meaningless at this imbalance — predicting "never fraud" scores 99.98%. Use precision and recall at the chosen threshold, and PR-AUC for model comparison, as covered in model evaluation and imbalanced data. Business metric: money lost to fraud plus genuine payments blocked, weighted by cost.

Architecture.

 payment request ─► feature fetch ─► model score ─► threshold policy
       │            device, amount     (< 100 ms total)     │
       │            history, velocity              ┌────────┼────────┐
       │            counters                     allow   extra OTP  block
       ▼                                                    │
 all decisions logged ─► analyst case review ─► confirmed labels ─► weekly retrain
                          (labels arrive days or weeks late)

Data and training loop. The hard part is labels. Fraud is confirmed by chargebacks and complaints, which arrive days or weeks later — so today's training data describes last month's fraud. Raise this yourself. Practical answer: retrain frequently on the freshest confirmed labels, keep fast-updating rule lists for brand-new attack patterns, and treat analyst review as a labelling pipeline, not an afterthought — see data labelling.

Failure modes. Adversarial drift: fraudsters probe the system and adapt, so the data distribution shifts because you deployed — the sharpest version of monitoring and drift. Festival spikes: Diwali-night traffic looks anomalous to a model trained on ordinary weeks, and blocking honest payments at scale that night is a headline. Feature-store timeout: decide in advance whether you fail open (allow) or fail closed (extra OTP), and say the trade-off out loud.

What interviewers probe. "Where do you set the threshold?" — from the two mistake costs, not from a metric curve alone; the researcher block derives it. "How do you retrain when labels arrive three weeks late?" — freshest-confirmed data plus rules for the gap. "Fail open or fail closed?" — there is no free answer; pick per amount, and small payments fail open.

Common mistakes

Opening with the model. "I'd use XGBoost" as sentence one signals you skipped requirements. The model appears at step five.

No numbers, ever. One rough envelope, spoken aloud, outranks a beautiful diagram with no scale attached.

Designing for the wrong scale. Ask the scale first. A startup design with three components beats a planet-scale design the company will never need.

No monitoring story. A design that cannot detect its own failure is incomplete. End every design with what you watch after launch.

Silent thinking. The interviewer can only grade what you say. Narrate the trade-offs, including the ones you decide against.

Try it yourself

Set a 45-minute timer and design "predict food delivery time" out loud, alone, through all seven steps. Speaking to an empty room feels ridiculous and is the single highest-value practice available — the round tests narrated thinking, and narrated thinking is trainable.

What to learn next

Researcher — Mathematics and papers.

Senior versions of this round push into four specific pieces of mathematics.

Position bias and inverse propensity scoring

Feed logs are biased: an item's click rate depends on the position the old ranker gave it. Training on raw logs rewards the incumbent. The standard correction is inverse propensity scoring (IPS):

L_IPS = (1/n) · Σᵢ  δ(yᵢ, ŷᵢ) / pᵢ
  • δ(yᵢ, ŷᵢ) — the loss on logged interaction i.
  • pᵢ — the propensity: the probability the old system showed item i at the position where it was observed.
  • n — the number of logged interactions.

Dividing by pᵢ up-weights interactions the old policy rarely allowed, making the estimate unbiased — at the price of variance, since tiny propensities explode. Practical systems clip propensities at a floor. Joachims et al. (2017), Unbiased Learning-to-Rank with Biased Feedback, is the reference treatment.

Calibration, and why ranking scores must sometimes be probabilities

A model is calibrated when, among items scored p, the positive rate is actually p. Ranking alone does not need calibration — order survives any monotone distortion. Decisions that multiply the score by money do need it: expected fraud loss is p · amount, and ad pricing multiplies click probability by bid. State when your design needs calibrated outputs, and name the fixes: Platt scaling or isotonic regression on a held-out set, checked with reliability curves.

The threshold, derived from costs

Let c_fp be the cost of acting on a negative (blocking a genuine payment) and c_fn the cost of missing a positive (letting fraud through). Acting is optimal when expected cost of acting is lower:

p · 0 + (1−p) · c_fp   <   p · c_fn
        ⇒   act when   p > c_fp / (c_fp + c_fn)
  • p — the model's calibrated probability that this transaction is fraud.

With c_fn = 9 · c_fp, the threshold is 0.1 — far below the 0.5 beginners default to. This derivation, done on the whiteboard in thirty seconds, is one of the strongest signals available in the round. It also explains step-up actions: a middle action with its own smaller c_fp earns its own, lower threshold.

Why ROC-AUC misleads at extreme imbalance

The false positive rate has the negatives — millions — in its denominator, so a model can drown analysts in false alarms while its FPR stays microscopic and its ROC-AUC stays superb. Precision has your alerts in the denominator, which is what the analysts actually experience. Davis and Goadrich (2006), The Relationship Between Precision-Recall and ROC Curves, formalises when the two disagree. At fraud-level base rates, report PR-AUC.

Delayed feedback

Fraud labels and ad conversions share a structure: the positive label arrives with a long, heavy-tailed delay, so recent examples are censored — a "negative so far" may be a positive you have not seen yet. Chapelle (2014), Modeling Delayed Feedback in Display Advertising, models the delay distribution jointly with the conversion model instead of discarding recent data. Mentioning censoring by name marks a candidate who has shipped one of these.

Papers worth naming

  • Covington, Adams and Sargin, Deep Neural Networks for YouTube Recommendations, RecSys 2016 — the canonical two-stage retrieval-and-ranking write-up.
  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — the "the model is the small box" paper; interviewers quote it.
  • McMahan et al., Ad Click Prediction: a View from the Trenches, KDD 2013 — calibration and monitoring at production scale.
  • Joachims et al., Unbiased Learning-to-Rank with Biased Feedback, WSDM 2017.
  • Chapelle, Modeling Delayed Feedback in Display Advertising, KDD 2014.
  • Davis and Goadrich, The Relationship Between Precision-Recall and ROC Curves, ICML 2006.

What to learn next