Deep learning questions, with answers
Twenty real deep learning interview questions with worked answers — vanishing gradients, batch norm, dropout, Adam, initialisation and the research-depth follow-ups.
- 18 min read
- 3 reading levels
- Published
Read these first
On this page 9
- Why this round exists
- The shape of the round
- Q1. Why do neural networks need activation functions?
- Q2. What does "the gradient vanishes" mean, in plain words?
- Q3. Why are CNNs so good at images?
- Q4. Your network aces training data and fails on new data. Name three cures
- Q5. What is transfer learning?
- Remember this
- What to learn next
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The deep learning round checks whether you understand the machine, or only know how to drive it.
Anyone can ride a scooter. The mechanic's interview asks a different kind of question: why does it stall on cold mornings? You can ride for years without being able to answer that. Deep learning interviews are mechanic questions about networks you may have driven for years.
Every question below is one a working engineer eventually meets as a real bug. That is why interviewers love them.
Why this round exists
Training a neural network on a tutorial dataset works first try. Training one on real data fails in strange ways: the loss freezes, or explodes, or the model aces training and flunks reality. Companies pay for people who have seen these failures and know their causes. The questions are compressed versions of those debugging sessions.
This lesson carries twenty questions. Five are answered here with no maths at all. Questions 6–15 live in the developer tab, 16–20 in the researcher tab.
The shape of the round
"what is X?" ← warm-up, definitions
│
▼
"why does X exist?" ← the real question
│
▼
"when does X fail?" ← where offers are decidedQ1. Why do neural networks need activation functions?
An activation function is the small bend applied after each layer's arithmetic. Without it, stacking layers is pointless: many straight-line steps combine into one straight-line step, so a fifty-layer network could do no more than a one-layer network.
The bend is what lets layers build on each other — edges combining into shapes, shapes into faces. Full lesson: activation-functions.
Q2. What does "the gradient vanishes" mean, in plain words?
Training works by passing a correction signal backwards through the layers, like a message whispered down a long line of people. With certain layer types, each person whispers a little softer. Thirty people deep, the message arrives silent — so the early layers never hear what to fix, and they stop learning.
That is a vanishing gradient. It is the reason very deep networks were nearly untrainable for years, and the reason several inventions in this lesson exist. The measured version is Q6.
Q3. Why are CNNs so good at images?
A CNN (convolutional neural network) looks at an image through a small sliding window, and reuses the same pattern-detector at every position. That matches two facts about pictures: nearby pixels belong together, and a cat's ear is a cat's ear whether it appears top-left or bottom-right.
Reusing one detector everywhere also means far fewer knobs to learn than connecting every pixel to every neuron. Fewer knobs, matched assumptions — that pairing is the whole answer. Full lesson: cnn.
Q4. Your network aces training data and fails on new data. Name three cures
This is overfitting wearing deep learning clothes. The three to name: get more data, including stretched and recoloured copies of what you have (augmentation); penalise or randomise the network during training so it cannot memorise (dropout, covered properly in Q9); and stop training earlier, at the point where held-out performance peaks.
The bonus cure that signals experience: start from a model pretrained on a huge dataset and adapt it — Q5.
Q5. What is transfer learning?
Borrowing a trained network instead of starting from nothing. A cook who has mastered one cuisine learns a second one fast — knife skills, timing and taste carry over. A network trained on millions of photos has already learned edges, textures and shapes; you keep those layers and re-teach the final ones for your task with a few hundred examples.
This is how nearly all real vision and language work happens now. Starting from scratch is the exception. Full story: fine-tuning.
Remember this
- Warm-up asks what; the marks are in why it exists and when it fails.
- Most questions trace back to two failure stories: gradients dying and memorising instead of learning.
- Each answer here compresses a real debugging session someone lived through.
What to learn next
- Backpropagation — the machinery underneath half of these questions.
- ML coding interviews — where these ideas get implemented under a timer.
- Transformers — the architecture the next question set is built on.
Developer — Code and libraries.
Setup
Questions 6–15. Two are answered with runnable NumPy.
pip install numpyCPU-only, under two seconds, no downloads. Outputs from NumPy 2.4; the seeds make them reproducible.
Q6. Show me the vanishing gradient problem — don't recite it
The strongest possible answer builds a 30-layer network and measures the correction signal arriving at each depth, once with sigmoid activations, once with ReLU:
import numpy as np
def sigmoid(z):
return 1 / (1 + np.exp(-z))
width, depth = 64, 30
def gradient_reaching_each_layer(kind):
rng = np.random.default_rng(0) # identical weights for both runs
h = rng.normal(size=width)
tape = [] # backprop needs W and the local slope
for _ in range(depth):
W = rng.normal(0, np.sqrt(2 / width), (width, width)) # He initialisation
z = W @ h
if kind == "relu":
h = np.maximum(z, 0)
local = (z > 0).astype(float) # relu slope: exactly 1 where active
else:
h = sigmoid(z)
local = h * (1 - h) # sigmoid slope: never above 0.25
tape.append((W, local))
grad = np.ones(width) # the gradient arriving from the loss
norms = {}
for layer in range(depth, 0, -1):
W, local = tape[layer - 1]
grad = W.T @ (grad * local) # one chain-rule step, walking down
norms[layer] = np.linalg.norm(grad)
return norms
sig = gradient_reaching_each_layer("sigmoid")
relu = gradient_reaching_each_layer("relu")
print("gradient size reaching layer: sigmoid relu")
for layer in (30, 20, 10, 1):
print(f"{layer:>27} {sig[layer]:>9.1e} {relu[layer]:>9.1e}")gradient size reaching layer: sigmoid relu
30 2.0e+00 8.5e+00
20 2.6e-05 2.0e+01
10 4.1e-10 2.4e+01
1 1.2e-14 3.2e+01With sigmoid, the signal shrinks fourteen orders of magnitude on its way down — layer 1 receives numerical silence and cannot learn. The mechanism: backprop multiplies one local slope per layer, sigmoid's slope never exceeds 0.25, and thirty numbers below 0.25 multiplied together is nearly nothing. ReLU's slope is exactly 1 on active units, so the product survives — 32 at the bottom, thirty layers down.
Then name the full toolkit, because one activation swap is not the whole cure: careful initialisation (Q16 territory), residual connections (Q14), normalisation layers (Q8), and gradient clipping for the exploding twin, where the product grows instead of shrinking — common in RNNs.
Q7. Why not initialise all weights to zero?
Because every neuron in a layer would then compute the same output, receive the same gradient, and take the same update — forever. The layer collapses into one neuron photocopied. This is the symmetry problem: identical starting points can never differentiate.
Random initialisation breaks the tie, letting different neurons specialise in different features. The follow-up — "random with what size?" — is Q16, and the demo above quietly used its answer (np.sqrt(2 / width), He initialisation).
Q8. What does batch normalisation do, and why does training differ from inference?
Batch norm re-centres and re-scales each layer's pre-activations, using the mean and spread of the current batch — the group of examples processed together. It then lets the network undo the operation via two learned parameters per channel, so nothing is forced. In practice it allows higher learning rates, makes training far less sensitive to initialisation, and speeds convergence.
The trap inside the question is train-versus-inference. At inference there may be no batch — one image arrives alone — so the layer switches to averages collected during training. Forgetting to flip that switch (model.eval() in PyTorch) is a classic real-world bug: the model behaves differently one example at a time than in batches. Mentioning that bug unprompted is the strong version. Why batch norm works is genuinely contested — that story is Q17.
Q9. Explain dropout — and why it is off at test time
Dropout randomly silences a fraction of neurons on every training step, a different set each time. No neuron can rely on a specific partner existing, so the network cannot build fragile co-dependent circuits — it must spread knowledge redundantly. It behaves like training many thinned networks that share weights, then averaging them.
At test time you want the full ensemble's stable answer, not a random thinned network — so nothing is dropped, and activations are rescaled to keep their expected size consistent with training (standard "inverted dropout" does the rescaling during training instead). The one-line summary interviewers reward: "noise at training time, deterministic average at test time."
Q10. Adam versus SGD — what does Adam actually track?
Adam keeps two running averages per parameter: the recent gradient direction (momentum) and the recent gradient magnitude. It divides one by the other, so parameters with consistently large gradients take careful steps and parameters with faint gradients take relatively bigger ones — a per-parameter learning rate, tuned automatically.
Trade-offs, which are the marks: Adam is robust to learning-rate choice and dominates in practice for transformers and sparse-gradient problems. Plain SGD with momentum, tuned well, has matched or beaten it on classic vision benchmarks and was long the choice for state-of-the-art CNNs. The senior detail: modern practice uses AdamW, which fixes how weight decay interacts with Adam's scaling — Q18 has the equations. Foundations: gradient-descent.
Q11. The loss is not going down. Walk me through your checklist
Ordered hypotheses, cheapest first:
- Overfit one batch. Train on ten examples; a healthy network memorises them in seconds. Failure here means a bug — wrong labels, wrong loss, frozen weights — not a tuning issue. This is the single highest-value trick in the round.
- Learning rate. Too high: loss spikes, oscillates, or turns
nan. Too low: loss creeps down at a crawl. Sweep powers of ten. - Data pipeline. Look at actual batches with your eyes. Shuffled labels, un-normalised inputs and duplicate-leaking splits all live here.
- Gradient flow. Print gradient norms per layer, like Q6's demo — silence at early layers or explosions at late ones each has its fix.
The framing sentence: "loss not falling is a symptom with about four common diseases; I test them in order of how cheap the test is."
Q12. How many parameters in a conv layer — and why so few?
The formula in words: channels-in times kernel-height times kernel-width times channels-out, plus one bias per output channel. Interviewers want the arithmetic done and the contrast drawn:
in_ch, out_ch, k = 3, 64, 3
conv = in_ch * k * k * out_ch + out_ch # weights + one bias per filter
print(f"conv 3x3, 3 -> 64 channels: {conv:,} parameters")
h, w = 224, 224 # a dense layer doing the same job
dense = (in_ch * h * w) * (out_ch * h * w)
print(f"dense layer on the same image: {dense:,} parameters")conv 3x3, 3 -> 64 channels: 1,792 parameters dense layer on the same image: 483,385,147,392 parameters
1,792 against 483 billion. The gap is the answer to "why convolutions": the parameter count is independent of image size, because the same 3×3 detector slides everywhere — Q3's weight sharing, quantified. Follow-up to expect: parameters are not activations. The conv layer's output is huge (64 channels at full resolution), which is why memory during training is dominated by activations, not weights.
Q13. What actually changes when I increase the batch size?
Mechanically: each update averages more examples, so the gradient estimate is less noisy, and the hardware runs more efficiently. The naive conclusion — "bigger is better" — is the trap.
What to say: larger batches take fewer, cleaner steps per epoch, and typically need a raised learning rate to compensate; the practical recipe is scaling the learning rate with the batch size, with a warm-up period. Very large batches were long associated with worse generalisation ("sharp minima"); later work showed much of the gap closes with proper tuning, but the noise of small batches does act as a mild regulariser. The honest close: "batch size is mostly an engineering choice — I pick the largest that fits memory, then retune the learning rate."
Q14. Why do residual connections let us train 100-layer networks?
A residual connection adds a layer's input directly to its output, so the layer learns only the change it should make, not the whole transformation. Two consequences, and naming both is the full answer.
For gradients: the addition creates a highway — during backprop, the correction signal flows through the identity path untouched, sidestepping Q6's shrinking product entirely. For optimisation: learning "change nothing" means pushing weights towards zero, which is easy — so extra layers can only help, never strand the network. Before ResNets, deeper networks often trained worse than shallow ones on training data itself — an optimisation failure, not overfitting — and residual connections removed it. They are also why transformers stack to enormous depth.
Q15. Why do softmax and cross-entropy appear together?
Softmax turns raw scores into probabilities; cross-entropy punishes the model by how little probability it gave the true class. Paired, something beautiful happens: the gradient at the output collapses to predicted probability minus true label — the cleanest error signal possible, large exactly when the model is confidently wrong.
The pairing is also numerically deliberate: computed separately, a confident wrong prediction can produce a log of nearly zero and overflow; fused (as every framework's CrossEntropyLoss does, via the log-sum-exp trick), it is stable. The mistake this question fishes for: putting an explicit softmax layer before a fused loss, applying it twice — a real, common, silent bug. Deriving the collapse is Q19's warm-up; the loss family lives in loss-functions.
Common mistakes in this round
Reciting cures without mechanisms. "Use ReLU, use batch norm, use residuals" without why reads as cargo cult. Every "use X" needs its failure story.
Never having watched a training curve. Q11 exposes this instantly. Train small things and break them on purpose; the pytorch-basics lesson is the place to start.
Treating contested questions as settled. Why batch norm helps, whether large batches generalise worse — these have live literatures (see the researcher tab). Presenting folklore as fact costs marks against a well-read interviewer.
Try it yourself
In vanishing_gradients.py, swap He initialisation for rng.normal(0, 0.01, ...) and predict what happens to the ReLU column before running — small weights shrink the product too, activation choice notwithstanding. Then reduce depth to 5 and see how much the sigmoid column recovers. Shallow networks are why sigmoid survived for decades.
What to learn next
- Backpropagation — the machinery underneath half of these questions.
- ML coding interviews — where these ideas get implemented under a timer.
- Transformers — the architecture the next question set is built on.
Researcher — Mathematics and papers.
Questions 16–20: the depth probes behind the developer-tab answers.
Q16. Derive the variance-preserving initialisation scale
Setup: a linear unit $z = \sum_{i=1}^{n} w_i x_i$ with i.i.d. weights and inputs, both zero-mean. Then
$$ \operatorname{Var}(z) = n \, \operatorname{Var}(w) \, \operatorname{Var}(x) $$
Where $n$ is the fan-in (inputs per neuron). For activation variance to pass through unchanged, $\operatorname{Var}(w) = 1/n$ — Xavier/Glorot initialisation (Glorot and Bengio, 2010), whose symmetric form $2/(n_{in} + n_{out})$ balances the forward and backward passes.
ReLU zeroes half of a symmetric input distribution, halving the variance: $\mathbb{E}[\max(0,z)^2] = \tfrac{1}{2}\operatorname{Var}(z)$ for symmetric $z$. Compensating gives $\operatorname{Var}(w) = 2/n$ — He initialisation (He et al., 2015), the $\sqrt{2/\text{width}}$ in this lesson's demo code. The interview kicker: mis-scaling compounds geometrically with depth — a factor of $c$ per layer becomes $c^{30}$ — so at depth 30 even the ReLU column of the Q6 demo collapses under a wrong scale. Initialisation, activation and depth interact as one product; that is the unified picture behind Q2, Q6 and Q7.
Q17. What does batch norm actually do? The contested mechanism
The operation (Ioffe and Szegedy, 2015), per channel over a batch $\mathcal{B}$:
$$ \hat{x} = \frac{x - \mu_\mathcal{B}}{\sqrt{\sigma^2_\mathcal{B} + \epsilon}}, \qquad y = \gamma \hat{x} + \beta $$
Where $\mu_\mathcal{B}, \sigma^2_\mathcal{B}$ are the batch mean and variance, $\epsilon$ a small constant preventing division by zero, and $\gamma, \beta$ learned parameters restoring representational freedom. Inference substitutes running estimates for the batch statistics.
The original explanation — reducing "internal covariate shift" — is largely rejected. Santurkar et al. (2018) showed networks with covariate shift artificially injected after BN still train well, and argued BN's real effect is smoothing the optimisation landscape (better Lipschitz behaviour of loss and gradients), permitting larger learning rates. Presenting the 2015 story as settled fact is a detectable error in front of a current interviewer; presenting the dispute is the strong move. Complete the answer with the practical taxonomy: layer norm normalises across features per example — no batch dependence, no train/test mismatch, and the choice of transformers; group norm for small-batch vision.
Q18. Write Adam's update rule and explain the bias correction and AdamW
Per parameter, with gradient $g_t$ (Kingma and Ba, 2014):
$$ m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t, \qquad v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2 $$
$$ \hat{m}_t = \frac{m_t}{1-\beta_1^t}, \qquad \hat{v}_t = \frac{v_t}{1-\beta_2^t}, \qquad \theta_{t+1} = \theta_t - \frac{\eta \, \hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} $$
Where $m_t, v_t$ are exponential moving averages of the gradient and squared gradient; $\beta_1 \approx 0.9$, $\beta_2 \approx 0.999$ their decay rates; $\eta$ the learning rate; $\epsilon \sim 10^{-8}$ for numerical safety. The bias correction terms $1 - \beta^t$ exist because both averages start at zero and are therefore biased low early in training; without correction, with $\beta_2 = 0.999$, early estimates of $v_t$ are near zero and the effective step $\eta / \sqrt{\hat v_t}$ explodes precisely when the model is least ready for it.
AdamW (Loshchilov and Hutter, 2017) is the follow-up that separates readers from users: applying L2 regularisation through the gradient lets it be divided by $\sqrt{\hat{v}_t}$, so heavily-updated parameters are barely decayed — the regularisation strength becomes accidentally per-parameter. AdamW decouples the decay, applying $-\eta \lambda \theta_t$ directly, restoring the intended penalty. It is the default optimiser of the transformer era for that reason.
Q19. Derive the softmax + cross-entropy gradient
With logits $z$, softmax $p_i = e^{z_i} / \sum_j e^{z_j}$, one-hot target $y$, and loss $L = -\sum_i y_i \log p_i$:
$$ \frac{\partial L}{\partial z_k} = p_k - y_k $$
The two-line derivation, expected on a whiteboard: $\partial p_i / \partial z_k = p_i(\delta_{ik} - p_k)$ where $\delta_{ik}$ is 1 when $i = k$ else 0; substituting into $\partial L/\partial z_k = -\sum_i (y_i / p_i) \, \partial p_i/\partial z_k$ and using $\sum_i y_i = 1$ collapses everything to $p_k - y_k$.
Two remarks convert a correct derivation into a strong one. Numerically, frameworks fuse the two operations to compute $\log p$ via log-sum-exp — $\log \sum_j e^{z_j} = z_{max} + \log \sum_j e^{z_j - z_{max}}$ — avoiding both overflow and $\log(0)$. Statistically, this loss is maximum likelihood under a categorical model, which links the whole construction to Q18 of statistics-interview-questions.
Q20. Why do giant networks generalise instead of memorising? Give the honest answer
The honest answer is that this is not fully solved, and saying so — with the evidence — is the researcher-grade response.
The provocation: Zhang et al. (2017) showed standard architectures can fit randomly labelled ImageNet perfectly. Whatever explains generalisation, it is not lack of capacity, and classical uniform bounds (VC dimension over the hypothesis class) are vacuous at these scales.
The live threads to present: implicit regularisation — SGD converges to particular minima among the many that fit, provably the maximum-margin solution in simplified settings (Soudry et al., 2018); double descent (Belkin et al., 2019; Nakkiran et al., 2019) — test error falls again beyond the interpolation threshold, with width, depth and even training epochs; flatness of minima correlating with generalisation, while being reparameterisation-sensitive (Dinh et al., 2017), so it cannot be the whole story. A closing sentence that lands: "the classical bias-variance picture from ml-theory-questions is not wrong, but it is incomplete in the overparameterised regime, and interviewers at research labs know it."
Sources
- Glorot, X. and Bengio, Y. (2010), Understanding the difficulty of training deep feedforward neural networks, AISTATS.
- He, K. et al. (2015), Delving deep into rectifiers — arxiv.org/abs/1502.01852
- Ioffe, S. and Szegedy, C. (2015), Batch normalization — arxiv.org/abs/1502.03167
- Santurkar, S. et al. (2018), How does batch normalization help optimization? — arxiv.org/abs/1805.11604
- Kingma, D. and Ba, J. (2014), Adam — arxiv.org/abs/1412.6980
- Loshchilov, I. and Hutter, F. (2017), Decoupled weight decay regularization — arxiv.org/abs/1711.05101
- He, K. et al. (2015), Deep residual learning for image recognition — arxiv.org/abs/1512.03385
- Srivastava, N. et al. (2014), Dropout, JMLR 15.
- Zhang, C. et al. (2017), Understanding deep learning requires rethinking generalization — arxiv.org/abs/1611.03530
- Belkin, M. et al. (2019), Reconciling modern machine-learning practice and the bias–variance trade-off, PNAS 116(32).
What to learn next
- Backpropagation — the machinery underneath half of these questions.
- ML coding interviews — where these ideas get implemented under a timer.
- Transformers — the architecture the next question set is built on.