ML Interview Preparation

Take-home assignments that pass

Reviewers spend about 25 minutes on your take-home and score process, not accuracy — here is exactly what they check, a submission checklist, and a worked structure.

On this page 5
  1. Why companies use them
  2. What actually happens to your submission
  3. The trap most people fall into
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A take-home assignment is a small ML project you complete at home, and it is scored on process and honesty far more than on model accuracy.

Think of the practical record book you kept in school science. Marks never went to the student with the flashiest experiment. They went to the neat file: aim stated, apparatus listed, steps anyone could repeat, readings honestly recorded — including the odd one that came out wrong. Same instinct here.

A take-home is your record book for one small dataset. The company is not buying your model. They are watching how you work when nobody is standing behind you.

Why companies use them

Live interviews are noisy. Nerves wreck good engineers, and confident talkers sail through. A take-home shows the actual work product: your code, your decisions, your writing.

There is a second reason candidates forget. The reviewer is a busy engineer who has perhaps twenty-five minutes and a pile of submissions. Anything that wastes those minutes — code that will not run, a report with no findings — fails regardless of the model inside.

What actually happens to your submission

 minute 0–2    reads the README
               can they tell what you found without opening code?
 minute 2–5    tries to run it, once
               one command works, or the mood sours for good
 minute 5–15   checks the evaluation
               how was data split? is there a baseline? is the
               claimed number believable?
 minute 15–25  skims code and write-up
               "would I want this person on my team?"

Everything in this lesson serves those twenty-five minutes.

The trap most people fall into

Most candidates spend nearly all their time squeezing out a slightly better score. Reviewers barely reward that. A submission with a modest score, an honest evaluation, and a clear write-up beats a higher score with a suspicious evaluation every single time.

The score is one line. The thinking around it is the submission.

Remember this

  • The reviewer has about twenty-five minutes. Design for them.
  • Process and honesty are the product. The score is a detail.
  • If it does not run with one command on their machine, little else matters.

What to learn next

Developer — Code and libraries.

What reviewers score

Typical rubric weights, from companies that publish them and reviewers who talk:

WhatRough weightWhat they actually check
Data handling30%leakage, sensible split, missing values faced honestly
Evaluation25%a baseline exists, right metric, claims match evidence
Framing20%did you answer the business question asked, or a different one
Code and reproducibility15%runs with one command, pinned dependencies, readable
Communication10%findings first, decisions explained, limits admitted

Weights vary; the ordering rarely does. Note where "model choice" sits: inside that 20% framing row, at best.

The structure that passes

churn-takehome/
├── README.md            ← findings FIRST, run instructions second
├── requirements.txt     ← pinned versions, e.g. pandas==2.2.3
├── data/                ← the given file, or a download script — never 500 MB in git
├── notebooks/
│   └── exploration.ipynb   ← the messy truth, labelled as exploration
├── src/
│   ├── prepare.py       ← raw data in, clean table out
│   ├── train.py         ← fits and saves the model, seed fixed
│   └── evaluate.py      ← THE file the reviewer runs
└── report.md            ← one page: decisions, results, what you would do next

The README opens with three sentences of findings — "churn concentrates in month-one users; a two-feature model reaches 0.81 AUC against a 0.74 baseline; the strongest lever is activation, not retention offers" — then one command:

bash
pip install -r requirements.txt
python src/evaluate.py

The one file the reviewer runs

Shrunk to twelve hand-typed rows so the mechanics are visible, this is the shape of evaluate.py — baseline printed first, always:

evaluate.py
# The reviewer runs ONE file. Make it this one, and print the baseline first.

data = [  # (hours_active, orders_last_month, churned) — 12 rows, enough to show the shape
    (1, 0, 1), (2, 0, 1), (1, 1, 1), (3, 0, 1), (2, 1, 1),
    (9, 4, 0), (12, 5, 0), (8, 3, 0), (15, 6, 0), (11, 2, 0),
    (4, 1, 1), (7, 3, 0),
]

train, test = data[:9], data[9:]     # real submissions split by time or id, never after peeking

churn_rate = sum(r[2] for r in train) / len(train)
majority = round(churn_rate)          # baseline: predict the commonest training label

def rule_model(hours, orders):
    return 1 if hours < 5 else 0      # one explainable rule; the bar any model must beat

def accuracy(rows, predict):
    hits = sum(predict(h, o) == y for h, o, y in rows)
    return hits / len(rows)

print(f"majority-class baseline : {accuracy(test, lambda h, o: majority):.2f}")
print(f"one-rule model          : {accuracy(test, rule_model):.2f}")
Output
majority-class baseline : 0.33
one-rule model          : 1.00

And here is the honesty the round rewards: that 1.00 means nothing — the test set has three rows. A real submission says so in exactly those words, next to the number, and reports a metric fit for the class balance, per model evaluation. Writing "perfect score, but n=3, so this is noise" is worth more than the score itself.

The submission checklist

Run this before sending, every time:

  • Fresh clone, fresh environment, one command — does it run end to end?
  • Is every dependency pinned to a version in requirements.txt?
  • Does a baseline appear before any model, and does the report compare against it?
  • Was every preprocessing step (scaling, encoding, imputation) fitted on training data only? See train-test split.
  • If rows have timestamps, is the split by time, with the test set entirely in the future?
  • Does the metric fit the class balance — no accuracy bragging on imbalanced data?
  • Does the README lead with findings, not with "clone this repo"?
  • Did you state your time spent, and what you would do with one more day?
  • Is there any file the reviewer does not need? Delete it.

Common mistakes

Leakage. The classic: fitting the scaler on all rows before splitting, so test statistics leak into training. The temporal version: random-splitting time-stamped data, letting the model peek at the future. Reviewers hunt for these first, and either one usually ends the review.

No baseline. An accuracy in a vacuum is uninterpretable. The majority class, or last month's value, takes five lines and makes every other number mean something.

The 40-cell notebook as the deliverable. Exploration is honest work — keep it, labelled. The deliverable is scripts plus a one-page report. A notebook the reviewer must run top-to-bottom, hoping cell 23 was not executed out of order, is a coin flip.

Over-engineering. A deep network on 800 rows, a config framework, six abstractions — these read as poor judgement, not skill. Match the tool to the data size; XGBoost or simpler wins most take-home-sized tables.

Blowing the time box. Asked for four hours, spending twenty is a negative signal with its own name in review discussions. State your hours in the README and cut scope, not honesty, to fit them.

Silent failure handling. Dropping 12% of rows without a sentence explaining why leaves the reviewer assuming the worst. Every discarded row and filled gap gets one line — see handling missing data.

Try it yourself

Take any old project and give it the twenty-five-minute treatment: hand the repo to a friend, start a timer, and watch silently. Note where they get stuck in the first five minutes — that friction, not your model, is what a reviewer would have scored.

What to learn next

Researcher — Mathematics and papers.

Four upgrades separate senior submissions from good ones.

Put an error bar on the headline number

A test-set metric is an estimate from a finite sample. For a proportion-like metric such as accuracy, the standard error is:

SE = √( acc · (1 − acc) / n )
  • acc — measured accuracy; n — number of test examples.

With acc = 0.85 on n = 200 rows, SE ≈ 0.025, so a 95% interval spans roughly ±0.05 — and a claimed improvement from 0.83 to 0.85 dissolves into noise. For ranking metrics like AUC, bootstrap instead: resample the test set with replacement about 1,000 times, recompute, report the 2.5th and 97.5th percentiles. One sentence — "0.81 AUC, 95% CI [0.77, 0.85]" — signals more statistical maturity than any model choice on the page.

Selection on the test set is quiet overfitting

Evaluate k model variants on the same test set and pick the best, and the winner's score is biased upward — you selected on the noise. The clean protocol: select on a validation split, touch the test set once, report that number whether or not it flatters you. Recht et al. (2019), Do ImageNet Classifiers Generalize to ImageNet?, measured the field-wide version: on freshly collected test sets, accuracies dropped substantially, with the drop tracking years of community reuse of the original test set. Saying "I froze the test set until the end" in a report, and meaning it, is rarer than it should be.

A leakage taxonomy worth citing

Kaufman, Rosset and Perlich (2012), Leakage in Data Mining: Formulation, Detection, and Avoidance, classify the ways future or target information seeps into features — including the subtle ones take-homes love: aggregates computed over the full dataset, identifiers correlated with the target through collection order, and features that are consequences of the label (calls-to-support after the churn decision). Their detection heuristic transfers directly: any feature suspiciously predictive on its own deserves an investigation before a celebration.

Reproducibility beyond the seed

Fixing seeds is necessary and insufficient. Results can still drift across library versions, BLAS builds, and thread counts — Raff (2019), A Step Toward Quantifying Independently Reproducible Machine Learning Research, found a meaningful fraction of published results could not be reproduced even with the authors' code. For a take-home the practical bar is: pinned versions, one entry point, no hidden state, and honest words where nondeterminism remains. A model-card-style paragraph — intended use, data provenance, known failure modes, per Mitchell et al. (2019) and the model cards lesson — closes the report with exactly the maturity reviewers hope to find.

References

  • Kaufman, Rosset and Perlich, Leakage in Data Mining, KDD 2012.
  • Recht, Roelofs, Schmidt and Shankar, Do ImageNet Classifiers Generalize to ImageNet?, ICML 2019.
  • Raff, A Step Toward Quantifying Independently Reproducible Machine Learning Research, NeurIPS 2019.
  • Mitchell et al., Model Cards for Model Reporting, FAT* 2019.

What to learn next