ML Interview Preparation

The ML interview landscape

An ML interview loop is four or five separate tests, each checking a different skill — here is what each round asks, and how to divide your preparation.

Read these first

On this page 6
  1. Why companies do it this way
  2. The rounds, in one picture
  3. What each round is really asking
  4. A real example you have seen
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An ML interview is not one test. It is four or five different tests, and each one checks a different skill.

Think of selection trials for a cricket team. The selectors never watch one match and decide. They test your batting in the nets, your bowling separately, your fielding, and your fitness. A brilliant batsman who cannot field still gets cut.

ML hiring works the same way. You face separate rounds, and each round can end your loop on its own.

Why companies do it this way

Companies kept making two opposite hiring mistakes.

Some hires could talk beautifully about models but could not write working code. Others could code anything but chose the wrong model for every problem, because they never understood the ideas.

One conversation cannot catch both failure types. So the interview split into rounds, with one skill per round. Knowing this changes how you prepare: you study for each round separately, like the cricketer who practises fielding on its own.

The rounds, in one picture

 screen        round 1        round 2       round 3          round 4
 ──────        ───────        ───────       ───────          ───────
 recruiter  →  coding      →  ML theory  →  ML system    →   behavioural
 call          (DSA and/or    and stats     design           (your projects,
               ML from        questions     ("design a       your decisions)
               scratch)                     spam filter
                                            at scale")

DSA means data structures and algorithms — the classic coding-puzzle material, arrays and hash maps and trees.

Not every company runs every round. A five-person startup might compress this into two conversations. A large company runs the full loop, and sometimes adds a take-home assignment before it.

What each round is really asking

  • Coding asks: can you turn an idea into working code, under time pressure?
  • Theory asks: do you understand why models fail, or do you only know library names?
  • Statistics asks: will you fool yourself with numbers? This round matters most for data scientist roles.
  • System design asks: can you build something real, with messy data and real users?
  • Behavioural asks: what did you decide on your past projects, and why?

Roles differ. A machine learning engineer (an engineer who builds and ships models) gets more coding and system design. A data scientist (someone who answers business questions with data) gets more statistics and case studies. A research scientist role adds paper discussions and deeper mathematics. Read the job description and weight your preparation to match.

A real example you have seen

Getting a driving licence needs a written test and a road test, separately. Passing one has never excused you from the other. Interview loops copied that idea, round by round.

The mistake most candidates make is preparing for the round they already enjoy. The loop fails you on your weakest round, not your strongest.

Remember this

  • The loop is several separate tests. Prepare for each one on its own.
  • Your weakest round decides the outcome, so spend your hours there.
  • The rounds differ by role: engineer loops lean coding, scientist loops lean statistics.

What to learn next

Developer — Code and libraries.

Setup

Nothing to install. This lesson's code is a study tool, not an algorithm — the algorithms start in ml-coding-interviews.

bash
python3 --version

Know your weakest round, honestly

The one piece of code worth running before a loop is an honest self-audit. Rate yourself per round, look at the bars, and give your next study hour to the shortest one.

readiness.py
ROUNDS = {
    "dsa coding":       "arrays, hashing, two pointers - the standard screen",
    "ml coding":        "k-means, attention, backprop from scratch",
    "ml theory":        "bias-variance, metrics, regularisation, leakage",
    "statistics":       "p-values, A/B tests, Bayes, bootstrap",
    "ml system design": "framing, data, serving, monitoring",
    "behavioural":      "your projects, told as decisions and trade-offs",
}

# rate yourself 1 (would fail today) to 5 (could pass today)
ratings = {
    "dsa coding": 4,
    "ml coding": 2,
    "ml theory": 3,
    "statistics": 2,
    "ml system design": 1,
    "behavioural": 3,
}

for name, r in sorted(ratings.items(), key=lambda kv: kv[1]):
    bar = "#" * r + "." * (5 - r)
    print(f"{name:<17} {bar}  {ROUNDS[name]}")

weakest = min(ratings, key=ratings.get)
print()
print(f"next study hour goes to: {weakest}")
Output
ml system design  #....  framing, data, serving, monitoring
ml coding         ##...  k-means, attention, backprop from scratch
statistics        ##...  p-values, A/B tests, Bayes, bootstrap
ml theory         ###..  bias-variance, metrics, regularisation, leakage
behavioural       ###..  your projects, told as decisions and trade-offs
dsa coding        ####.  arrays, hashing, two pointers - the standard screen

next study hour goes to: ml system design

The ratings above are an example profile — a strong coder who has never designed a full system. It is a common profile, and it fails loops.

What each round looks like from the inside

The DSA round. DSA means data structures and algorithms — arrays, hash maps, trees, the classic interview material. ML roles at larger companies still run this round, usually at a slightly gentler level than pure software roles. This site's /leetcode/ archive holds 1,055 worked solutions for exactly this practice.

The ML coding round. You implement a core algorithm from scratch in plain Python and NumPy. K-means, logistic regression, attention, backprop. The next lesson, ml-coding-interviews, builds all of these with you.

Theory and statistics rounds. Rapid-fire conceptual questions, spoken, no editor. The three question lessons in this section carry real questions with worked answers: ml-theory-questions, statistics-interview-questions and deep-learning-interview-questions.

System design. "Design a recommendation system for our app." One hour, a whiteboard, no code. You are graded on problem framing, data thinking and trade-offs, not on naming the fanciest model.

Behavioural. Your projects, interrogated. The strong version of every answer names a decision you made, an alternative you rejected, and why.

Common mistakes

Grinding only DSA. Software interview habits die hard. An ML loop weights the ML rounds heavier, and DSA practice does nothing for them.

Reading instead of speaking. Theory rounds are spoken. Recognising a correct answer feels identical to being able to produce one, until an interviewer is waiting. Practise answering out loud, with nothing open on screen.

Preparing projects as stories instead of decisions. "We built a churn model and accuracy was 91%" earns nothing. "We chose logistic regression over XGBoost because the team needed to explain scores to regulators" earns the offer.

Ignoring the recruiter screen. It sets which rounds you get. Ask directly: "what rounds does this loop have?" Recruiters answer this question happily, and most candidates never ask.

Try it yourself

Fill in ratings with your real, honest numbers. Then book your next five study sessions against the two shortest bars, not the longest one. Re-run it weekly — the bars move faster than you expect.

What to learn next

Researcher — Mathematics and papers.

What the evidence says about interviews as measurement

Personnel selection has a century of psychometric literature, and it maps well onto ML loops.

The meta-analytic tradition estimates the predictive validity of a selection method: the correlation $r$ between interview performance and later job performance. Schmidt and Hunter (1998), aggregating decades of studies, reported work-sample tests and structured interviews among the strongest single predictors, with unstructured interviews notably weaker. Sackett, Zhang, Berry and Lievens (2022) re-examined the corrections behind those numbers and revised most validities downward — structured interviews around $r \approx 0.42$, still near the top of the table.

Where:

  • $r$ — the Pearson correlation between the selection score and a criterion measure of job performance, across candidates.
  • Structured interview — fixed questions, anchored scoring rubrics, trained interviewers. The structure is what carries the validity.
  • Work sample — the candidate performs a slice of the actual job. The ML coding round and the take-home are work samples.

Two consequences for candidates follow directly.

First, the rounds that feel most artificial — fixed algorithm questions with rubrics — exist because structure measurably beats free-form conversation at predicting performance. Complaining about them misreads why they are there.

Second, even the best single method leaves most variance in job performance unexplained ($r \approx 0.4$ implies $r^2 \approx 0.16$). Companies stack rounds because individually weak signals combine; candidates should read a rejection the same way. A loop is a noisy estimator, and a single rejection is one draw from it.

Reliability, the quieter problem

Validity is capped by inter-rater reliability — the agreement between two interviewers scoring the same performance. Anchored rubrics ("a 3 means the candidate derived the gradient with one hint") exist to raise it. This is why memorised answer phrasings underperform understood answers: rubrics score reasoning steps, and a memorised paragraph collapses under one follow-up question.

Research-track loops differ in kind

Research scientist loops at industrial labs typically add:

  • a paper walkthrough — presenting your own publication to a hostile-but-fair audience, defending the baselines and ablations;
  • a paper-reading round — reacting to a paper you have not seen, locating the load-bearing claim and the weakest experiment;
  • depth probes into one area rather than breadth across all of them.

The grading criterion shifts from "can build" to "can generate and defend ideas".

Sources

  • Schmidt, F. L. and Hunter, J. E. (1998), The validity and utility of selection methods in personnel psychology, Psychological Bulletin 124(2).
  • Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F. (2022), Revisiting meta-analytic estimates of validity in personnel selection, Journal of Applied Psychology 107(11).
  • Huffcutt, A. I. and Arthur, W. (1994), Hunter and Hunter (1984) revisited: interview validity for entry-level jobs, Journal of Applied Psychology 79(2) — the structured-versus-unstructured decomposition.

What to learn next