ML theory questions, with answers
Twenty real ML theory interview questions with worked answers at the depth a strong candidate gives — bias-variance, regularisation, metrics, leakage and more.
- 17 min read
- 3 reading levels
- Published
Read these first
On this page 9
- Why this round exists
- How to answer, in one picture
- Q1. What is the difference between supervised and unsupervised learning?
- Q2. What is overfitting, and how would you detect it?
- Q3. Why do we split data into training and test sets?
- Q4. My model has 95% accuracy. Why might that be terrible?
- Q5. What is data leakage?
- Remember this
- What to learn next
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The theory round is a spoken exam where the interviewer keeps asking "why" until they find the edge of what you understand.
You have sat this exam before. It is the school viva: the teacher hears your first answer, nods, and asks a harder follow-up. The follow-ups continue until you reach something you cannot answer, and that is where they mark you.
So memorised one-liners fail here by design. Every answer below is written the way a strong candidate speaks: the direct answer first, then the reason, then an example, then the trade-off.
Why this round exists
Anyone can train a model by copying a tutorial. The theory round checks whether you can choose and debug — which needs the ideas, not the commands. The interviewer's real question, behind all twenty below, is: "when your model fails at my company, will you know where to look?"
How to answer, in one picture
question
│
▼
direct answer (one sentence)
│
▼
the why (two or three sentences)
│
▼
a concrete example
│
▼
the trade-off or failure case ← this line gets you hiredThis lesson carries twenty questions. The first five are asked of everyone, and are answered here without a single formula. Questions 6–15 live in the developer tab, 16–20 in the researcher tab.
Q1. What is the difference between supervised and unsupervised learning?
Supervised learning means learning from examples that come with answers attached — photos labelled "cat" or "dog". Unsupervised learning means finding structure in examples with no answers attached — grouping customers by behaviour without anyone defining the groups.
The follow-up is usually "give one real use of each". Spam filtering is supervised: old mails labelled spam or not. Customer segmentation is unsupervised: the groups did not exist until the algorithm found them. Full lessons: supervised-learning and unsupervised-learning.
Q2. What is overfitting, and how would you detect it?
Overfitting is when a model memorises its training examples instead of learning the pattern — like a student who memorised last year's question paper and fails the moment a question is reworded.
You detect it with data the model never saw. Score the model on training data and on held-out data. A big gap — excellent on training, poor on held-out — is the signature. The cures, in the order to try them: more data, a smaller model, regularisation (explained in Q7). The full story is in overfitting-and-underfitting.
Q3. Why do we split data into training and test sets?
Because a model's score on data it studied proves nothing, the same way marking your own homework proves nothing. The test set is a sealed exam paper: kept locked away during training, opened once at the end.
The strong candidate adds the discipline part: you tune on a third split (validation), and touch the test set as rarely as possible. Peek at it repeatedly while tuning and it quietly becomes training data. Full lesson: train-test-split.
Q4. My model has 95% accuracy. Why might that be terrible?
Because if 95% of the examples belong to one class, a model that always predicts that class scores 95% while learning nothing. Fraud detection, disease screening and spam are all like this — the interesting class is rare.
That model catches zero frauds. The fix is to report metrics that watch the rare class — explained properly in Q8 — and this question is usually the doorway to that one. Full lesson: model-evaluation.
Q5. What is data leakage?
Data leakage is when information from the future, or from the answer itself, sneaks into the model's inputs. The model looks brilliant in testing and collapses in production, because in production the future is not available.
The classic example: predicting hospital readmission using a column like "number of follow-up visits" — which only exists after the readmission happened. The interviewer wants one honest sentence too: leakage is the most common cause of results that look too good, so treat any suspiciously high score as a leak until proven otherwise.
Remember this
- Answer in the shape: answer, why, example, trade-off.
- The follow-ups are the real exam. Practise out loud.
- If a score looks too good, say the word leakage before the interviewer does.
What to learn next
- Statistics questions, with answers — the round that usually shares a slot with this one.
- Overfitting and underfitting — the lesson behind the most-asked question here.
- Model evaluation — precision, recall and the metric questions in full.
Developer — Code and libraries.
Setup
Questions 6–15. Two of them are best answered with ten lines of NumPy, so:
pip install numpyBoth demos run on CPU in under a second. Outputs were produced with NumPy 2.4; the same seeds give the same numbers on any recent version.
Q6. Explain the bias-variance trade-off
Bias is error from a model too rigid to capture the pattern. Variance is error from a model so flexible it captures the noise. Total error carries both, plus noise nobody can remove — so shrinking one usually grows the other, and you are choosing where to sit.
This is more convincing run than recited. Fit polynomials of increasing flexibility to fifteen noisy points from a sine curve:
import numpy as np
rng = np.random.default_rng(7)
def sample(n):
x = rng.uniform(0, 1, n)
return x, np.sin(2 * np.pi * x) + rng.normal(0, 0.25, n) # true curve + noise
x_train, y_train = sample(15) # small training set, like most real problems
x_test, y_test = sample(200) # large held-out set, so test error is stable
print("degree train error test error")
for degree in (1, 3, 10):
coef = np.polyfit(x_train, y_train, degree)
train = np.mean((np.polyval(coef, x_train) - y_train) ** 2)
test = np.mean((np.polyval(coef, x_test) - y_test) ** 2)
print(f"{degree:>6} {train:>11.3f} {test:>10.3f}")degree train error test error
1 0.146 0.284
3 0.023 0.147
10 0.016 14.636Degree 1 is a straight line through a sine wave: high error everywhere — that is bias. Degree 10 threads through every noisy point: near-zero training error, test error a hundred times worse — that is variance. Degree 3 sits in the valley between them. The interview sentence: "training error only ever falls with flexibility; test error is U-shaped, and I model-select for the bottom of the U." The formal decomposition is Q16.
Q7. L1 versus L2 regularisation — what is the practical difference?
Regularisation adds a penalty on large weights to the training objective, taxing complexity. L2 penalises squared weights and shrinks everything smoothly towards zero, but rarely to zero. L1 penalises absolute values and drives many weights exactly to zero — it performs feature selection while training.
So: L1 when you suspect most features are irrelevant and want a sparse, explainable model. L2 as the default, especially with correlated features, where L1 arbitrarily keeps one and zeroes its twin. The geometric and mathematical reason L1 zeroes weights is Q17. Both appear in practice in linear-regression and logistic-regression.
Q8. Precision versus recall — and when do you optimise for which?
Precision: of everything the model flagged, what fraction was truly positive. Recall: of everything truly positive, what fraction the model flagged. They trade off against each other through the decision threshold.
Optimise recall when a miss is expensive: cancer screening, where a missed case can be fatal and a false alarm costs one more test. Optimise precision when a false alarm is expensive: blocking payments as fraud, where each false block is a furious customer. Here is Q4's trap made concrete:
import numpy as np
y_true = np.array([0] * 950 + [1] * 50) # 5% fraud, like a real payments table
y_pred = np.zeros(1000, dtype=int) # a "model" that always says "not fraud"
accuracy = (y_pred == y_true).mean()
caught = ((y_pred == 1) & (y_true == 1)).sum()
print(f"accuracy: {accuracy:.1%}")
print(f"frauds caught: {caught} out of {y_true.sum()}")accuracy: 95.0% frauds caught: 0 out of 50
95% accurate, zero recall. Name the summary metrics too: F1 (the harmonic mean of precision and recall), PR-AUC (area under the precision-recall curve, threshold-free) for rare positives, and ROC-AUC when classes are balanced. Full treatment: model-evaluation.
Q9. How does cross-validation work, and when is one split not enough?
k-fold cross-validation cuts the data into k parts, trains k times — each time holding out a different part for scoring — and averages the k scores. One split gives you a single noisy estimate; five or ten folds give a mean and a spread.
When: small datasets, where one split's luck dominates, and any time you are comparing models or tuning. When not: huge datasets, where one split is already stable and k training runs cost real money. The honest extra: with time-ordered data, random folds leak the future into training — use time-based splits instead. That sentence has ended interviews well. See train-test-split.
Q10. Logistic regression or a tree ensemble — how do you choose?
Start with what the business needs from the model beyond accuracy. Logistic regression gives calibrated probabilities, coefficients a regulator can read, and works well when effects are roughly additive. Tree ensembles like random forests and XGBoost capture interactions and non-linear effects without feature engineering, and usually win on raw accuracy for tabular data.
The strong answer names a decision rule: regulated or high-explainability settings, or tiny datasets — logistic regression, possibly with hand-built interaction features. Accuracy-first tabular problems — gradient boosting, and keep the logistic model as the baseline that catches leakage. Saying "I would build both; the simple one is my sanity check" reads as experience.
Q11. What is the difference between bagging and boosting?
Both build many weak models and combine them; they differ in how the models relate. Bagging trains models independently, in parallel, each on a random resample of the data, and averages them — variance goes down. Random forest is bagging plus random feature choices. Boosting trains models one after another, each focused on the mistakes of the ones before, and sums them — bias goes down. XGBoost is the standard implementation.
The trade-off line: bagging is robust and hard to overfit with more trees; boosting is more accurate when tuned and more able to overfit, so it leans on early stopping and learning-rate shrinkage.
Q12. Your classes are 99-to-1 imbalanced. What do you do?
First, refuse to fix what may not be broken: if the model ranks well (good PR-AUC), the "problem" may be the metric, not the data. Then, in order of preference: pick metrics that see the rare class (Q8); use class weights, which most libraries support and which cost nothing; adjust the decision threshold to the business's cost ratio; resample only after those — oversampling duplicates rare examples, undersampling throws data away, and both distort probabilities.
The differentiating sentence: resampling changes the base rate your model believes in, so its probability outputs must be recalibrated before anyone treats them as probabilities.
Q13. How do you handle missing values, and what can go wrong?
First ask why they are missing, because the mechanism decides the fix. If a sensor dropped values at random, filling with the median is fine. If missingness carries meaning — income left blank by people with low income — then the blank is information: add a "was missing" indicator column alongside any fill.
What goes wrong: computing the fill value on the whole dataset before splitting. That leaks test-set statistics into training — leakage again (Q5). Compute fills on the training split only, inside the pipeline. Deleting rows is honest but can silently delete a biased slice of your data; say you would check what remains.
Q14. Why do we scale features, and which models need it?
Scaling puts features on comparable ranges — for instance, mean 0 and standard deviation 1. Distance-based and gradient-based models need it: k-nearest neighbours and k-means because a feature measured in thousands drowns one measured in decimals; neural networks and anything trained by gradient descent because wildly different scales make optimisation slow and unstable. Regularised linear models need it because the penalty hits all weights equally, which is only fair if features share a scale.
Tree models do not care — they split on thresholds, and any monotonic rescaling leaves the splits identical. Knowing which models are immune is what distinguishes the answer. And the fitting rule from Q13 applies: fit the scaler on training data only. See feature-engineering.
Q15. Your model was great offline and is failing in production. Walk me through your debugging.
A checklist question — the interviewer wants ordered hypotheses, not one guess.
- Data mismatch first: is production input processed identically to training input? Different tokenisation, units, encodings or defaults — this is the most common cause (training-serving skew).
- Leakage second (Q5): was the offline score ever real? A feature unavailable at prediction time explains everything.
- Drift third: has the world moved since the training data was collected? Compare production input distributions against training distributions. See monitoring-and-drift.
- Feedback loops and edge cases last: is the model seeing traffic the training data never covered?
Closing sentence: "and I would want monitoring that catches this before users do" — which is the door to the MLOps conversation.
Common mistakes in this round
Answering only the direct question. The trade-off sentence is where marks live; volunteer it.
Reciting definitions at maximum speed. One slow, structured answer outscores five rushed ones.
Bluffing at the edge. The viva structure means the interviewer will reach a question you cannot answer. "I do not know, but here is how I would reason about it" scores; confident nonsense ends loops.
Try it yourself
Cover the answers. Speak each of Q6–Q15 aloud in the four-part shape from the beginner tab, and time yourself — ninety seconds per answer is the target. Then change the degree list in bias_variance.py to (2, 5, 8, 12, 14) and predict the two error columns before running.
What to learn next
- Statistics questions, with answers — the round that usually shares a slot with this one.
- Overfitting and underfitting — the lesson behind the most-asked question here.
- Model evaluation — precision, recall and the metric questions in full.
Researcher — Mathematics and papers.
Questions 16–20: the depth probes. These arrive as follow-ups when your earlier answers were strong.
Q16. State the bias-variance decomposition formally
For squared loss, with data generated as $y = f(x) + \varepsilon$, $\mathbb{E}[\varepsilon] = 0$, $\operatorname{Var}(\varepsilon) = \sigma^2$, and $\hat{f}_D$ an estimator trained on a random dataset $D$:
$$ \mathbb{E}_{D,\varepsilon}\big[(y - \hat{f}_D(x))^2\big] = \underbrace{\big(f(x) - \mathbb{E}_D[\hat{f}D(x)]\big)^2}{\text{bias}^2}
- \underbrace{\mathbb{E}_D\big[(\hat{f}_D(x) - \mathbb{E}_D[\hat{f}D(x)])^2\big]}{\text{variance}}
- \underbrace{\sigma^2}_{\text{irreducible}} $$
Where:
- $f(x)$ — the true regression function.
- $\hat{f}_D(x)$ — the model's prediction at $x$ after training on dataset $D$; the expectation $\mathbb{E}_D$ averages over datasets drawn from the same distribution.
- $\sigma^2$ — the noise floor no model can beat.
The derivation is three lines of adding and subtracting $\mathbb{E}_D[\hat{f}_D(x)]$ and noting the cross-terms vanish. Two honest caveats a researcher-level answer includes: the clean decomposition is specific to squared loss (0-1 loss versions exist but are messier — Domingos, 2000), and modern overparameterised networks complicate the classical U-shape ("double descent", Belkin et al., 2019 — test error can fall again past the interpolation point). Citing that shows your picture is current, not textbook-frozen.
Q17. Why exactly does L1 produce sparsity?
The optimisation-theoretic answer. The L1-regularised objective $\min_w L(w) + \lambda \lVert w \rVert_1$ is non-differentiable at $w_i = 0$; the subdifferential of $|w_i|$ there is the whole interval $[-1, 1]$. So $w_i = 0$ is optimal whenever the gradient of the data term satisfies $|\partial L / \partial w_i| \le \lambda$ — the origin is "sticky", absorbing every coordinate whose evidence is weaker than $\lambda$. Under L2 the penalty term $2\lambda w_i$ vanishes at zero, so the origin is not special: any nonzero gradient pulls $w_i$ off it.
The geometric picture, for the whiteboard: the constraint region $\lVert w \rVert_1 \le t$ is a diamond with corners on the axes, and loss contours typically first touch a corner — where some coordinates are exactly zero. The L2 ball has no corners. Equivalent Bayesian framing: L2 is a Gaussian prior on weights, L1 a Laplace prior with a sharp peak at zero (Tibshirani, 1996, the lasso paper).
Q18. Generative versus discriminative classifiers — and which needs less data?
A discriminative model learns $p(y \mid x)$ directly (logistic regression). A generative model learns the joint — $p(x \mid y)$ and $p(y)$ — and classifies via Bayes' rule (naive Bayes). The generative model can generate data and handle missing features naturally; the discriminative model wastes no capacity modelling $x$ and usually wins in asymptotic accuracy.
The classic result to cite: Ng and Jordan (2001) compare exactly the naive Bayes / logistic regression pair. Naive Bayes converges to its (higher) asymptotic error with a number of examples only logarithmic in the number of features; logistic regression needs a number linear in the feature count, but converges to a lower error. So generative can win small-data regimes, discriminative wins large-data ones — with the caveat that the result depends on naive Bayes' independence assumption being tolerable.
Q19. What is the kernel trick, actually?
Many linear algorithms touch data only through inner products $\langle x_i, x_j \rangle$. Replace each with a kernel $K(x_i, x_j) = \langle \phi(x_i), \phi(x_j) \rangle$, and the algorithm runs implicitly in the feature space of $\phi$ — without ever computing $\phi(x)$, which may be infinite-dimensional, as with the RBF kernel $K(x, x') = \exp(-\gamma \lVert x - x' \rVert^2)$.
Mercer's condition says any symmetric positive semi-definite $K$ corresponds to some $\phi$, so you design similarity functions instead of feature maps. The representer theorem guarantees the solution is a weighted sum of kernel evaluations at training points. The cost sentence that shows systems awareness: the kernel matrix is $n \times n$, so kernel methods scale as $O(n^2)$ memory and worse in time — which is exactly why deep networks displaced them at scale, learning $\phi$ instead of fixing it.
Q20. What breaks in high dimensions?
Several distinct things, and naming them separately is the strong answer.
Volume: covering $[0,1]^d$ at resolution $\epsilon$ needs $(1/\epsilon)^d$ cells, so density estimation and nearest-neighbour methods need data exponential in $d$.
Distance concentration: for i.i.d. coordinates, the ratio between the farthest and nearest neighbour distances tends to 1 as $d$ grows (Beyer et al., 1999) — "nearest" neighbour loses meaning, degrading k-NN and clustering.
Geometry gets strange: nearly all the volume of a high-dimensional ball sits in a thin shell near its surface; random vectors are nearly orthogonal with high probability.
The rescue clause: real data rarely fills the ambient space — it concentrates near lower-dimensional manifolds, which is why methods with the right inductive bias (convolutions, learned representations, dimensionality reduction) work at all. Ending on the manifold sentence turns a doom answer into an insight answer.
Sources
- Domingos, P. (2000), A unified bias-variance decomposition, ICML.
- Belkin, M. et al. (2019), Reconciling modern machine-learning practice and the classical bias–variance trade-off, PNAS 116(32).
- Tibshirani, R. (1996), Regression shrinkage and selection via the lasso, JRSS B.
- Ng, A. and Jordan, M. (2001), On discriminative vs. generative classifiers, NeurIPS.
- Beyer, K. et al. (1999), When is "nearest neighbor" meaningful?, ICDT.
- Hastie, T., Tibshirani, R. and Friedman, J., The Elements of Statistical Learning, 2nd ed. — chapters 2, 3 and 7 cover Q16–Q19 fully.
What to learn next
- Statistics questions, with answers — the round that usually shares a slot with this one.
- Overfitting and underfitting — the lesson behind the most-asked question here.
- Model evaluation — precision, recall and the metric questions in full.