ML Interview Preparation

LLM and GenAI questions, with answers

The LLM questions interviewers actually ask — hallucination, RAG versus fine-tuning, temperature, tokens, cost — with answers that show you have shipped, not memorised.

On this page 6
  1. Why this round exists
  2. What the round looks like
  3. Three questions, answered in plain English
  4. Where you have already seen all this
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An LLM interview tests whether you understand how tools like ChatGPT behave, not whether you can recite definitions.

Think of the viva after a school science practical. The examiner does not ask you to recite the textbook. They point at your experiment and ask, "why did the flame turn green?" You pass by explaining things you have watched happen with your own eyes.

LLM interview rounds work the same way. Nearly every question is secretly asking one thing: have you watched these models fail, and do you know why?

Why this round exists

Companies now ship real products on top of language models. Support bots, search summaries, writing assistants, coding tools. When the person building them holds a wrong mental model, the product fails in public.

There is a cautionary tale interviewers love. An airline's chatbot invented a refund policy, a customer relied on it, and a tribunal made the airline honour it. That is what a shaky mental model costs. So the interview probes for exactly that.

What the round looks like

 warm-up      "What is a token? What does temperature do?"
    │
    ▼
 behaviour    "Our bot invented a refund policy. Why did
    │          that happen? What would you change first?"
    ▼
 judgement    "RAG or fine-tuning for this use case?
               Defend your choice."

The interviewer is listening for mechanisms — reasons rooted in how the model works. Brand names and buzzwords score nothing.

Three questions, answered in plain English

Why does the model sometimes make things up? Because it always writes the most likely-sounding next words. There is no separate step inside that checks facts. Confident nonsense is the price of fluent text. The full story is in hallucination.

Why does the same question give different answers each time? The model rolls a weighted dice when picking each word. You can turn that randomness up or down. See temperature and sampling.

Why does it not remember yesterday's chat? Every conversation starts empty. Apps that appear to remember are quietly pasting your old messages back in. See context window.

Where you have already seen all this

  • A support chat that quotes your actual order — that is retrieval feeding the model.
  • Search results with an AI summary and links underneath — retrieval again, with citations.
  • A chat app "forgetting" the start of a long conversation — the pasted history hit its limit.

You have lived the failure modes. The interview asks you to name their causes.

Remember this

  • Interviewers score mechanisms, not memorised definitions.
  • Almost every LLM question maps to a failure you can name and explain.
  • "I watched it fail this way, and here is why" is the strongest answer shape there is.

What to learn next

Developer — Code and libraries.

The questions below appear, in some wording, in most GenAI screens. Each answer follows the shape interviewers reward: mechanism first, trade-off second, experience third.

Setup

bash
pip install numpy

One question below comes with runnable proof. It needs no GPU and finishes in under a second.

"What does temperature actually do?"

The weak answer is "it controls creativity". The strong answer: the model produces a score per candidate word, and temperature rescales those scores before they become probabilities. Low temperature sharpens the distribution; high temperature flattens it.

You can watch this happen with four fake word scores:

temperature_demo.py
import numpy as np

logits = np.array([2.0, 1.0, 0.2, -1.0])   # raw scores the model gives 4 candidate words
words = ["delhi", "mumbai", "pune", "goa"]

def softmax(x, temperature):
    z = x / temperature
    z = z - z.max()            # subtracting the max avoids overflow; result is unchanged
    p = np.exp(z)
    return p / p.sum()

for t in [0.2, 1.0, 2.0]:
    probs = softmax(logits, t)
    row = "  ".join(f"{w}:{p:.3f}" for w, p in zip(words, probs))
    print(f"T={t}:  {row}")
Output
T=0.2:  delhi:0.993  mumbai:0.007  pune:0.000  goa:0.000
T=1.0:  delhi:0.632  mumbai:0.232  pune:0.104  goa:0.031
T=2.0:  delhi:0.447  mumbai:0.271  pune:0.182  goa:0.100

At T=0.2 the top word takes 99% of the probability, so outputs barely vary. At T=2.0 the model behaves close to a random word picker.

The follow-up is always "when would you set it to zero?" Answer: extraction, classification, anything a program will parse. Then add the honest caveat: temperature 0 reduces variation, but does not guarantee bit-identical outputs across servers, batch sizes, or model updates. Saying that caveat out loud is a strong signal.

"Our chatbot hallucinates. What do you do, cheapest first?"

Interviewers ask this to see whether you reach for an ordered plan or a silver bullet.

  1. Constrain the prompt. Tell the model to answer only from provided context, and give it explicit permission to say "I don't know". Refusal must be a legal output.
  2. Add retrieval, so there is something true in front of it to copy from. That is RAG.
  3. Require citations. A wrong answer with a source label is catchable by a human in seconds. A wrong answer without one is invisible.
  4. Build the escalation path. Low retrieval confidence should route to a human, not to a guess.
  5. Measure. Build a set of real questions with known answers, and score groundedness before and after each change.

Close with the honest sentence: nothing on that list eliminates hallucination. Claiming something does is an instant credibility hit.

"RAG or fine-tuning — how do you choose?"

The rule interviewers want to hear: retrieval for knowledge, fine-tuning for behaviour.

Facts that change — prices, policies, inventory — belong in retrieval. Updating an index is cheap and instant, and you get citations for free. Style, format, tone, and domain-specific behaviour belong to fine-tuning, usually via LoRA so the cost stays sane.

The trap inside the question: fine-tuning is an unreliable way to inject facts. The model tends to learn the style of your documents rather than their contents. Also mention that the two compose — a fine-tuned model behind a retrieval layer is a normal production setup, not a contradiction.

"Costs tripled last month. Traffic was flat. Why?"

Billing is per token — the word-piece unit models read and write — in both directions. Flat request counts with growing token counts means the requests themselves got fatter.

The usual culprits: chat history is resent on every turn, and conversations got longer. Someone raised the number of retrieved chunks. A prompt template grew. The answer that gets hired: "I would measure tokens per request first, then cap history, summarise older turns, cache repeated prefixes, and route easy queries to a smaller model."

"The model must return JSON. It manages 95% of the time. Fix it."

Never parse and pray. In rising order of strength:

  • Validate every response against a schema, and on failure retry once with the error message pasted into the prompt.
  • Use the provider's structured output mode, which constrains generation to the schema. See structured output.
  • For tool arguments specifically, use function calling rather than asking for free-form JSON.

The one-line summary interviewers like: turn "usually valid" into "valid by construction, with a checked fallback".

"What is prompt injection? Does a strong system prompt stop it?"

Prompt injection is when text the model reads — a user message, a retrieved web page, a document — contains instructions that hijack its behaviour. And no, a strong system prompt does not stop it. Instructions and data travel down the same channel, so the model cannot reliably tell them apart.

Real mitigations: give the model least-privilege tools, treat all retrieved text as untrusted, filter outputs, and put a human approval step in front of consequential actions. Full lesson: prompt injection.

Common mistakes in this round

Brand names instead of mechanisms. "I'd use LangChain" answers nothing. Describe the moving parts; name tools only if asked.

Absolutes. "RAG eliminates hallucination" and "temperature 0 is deterministic" are both false, and interviewers know it.

Parameter-count worship. Quoting model sizes as proof of quality signals reading, not building. Task fit, latency, and cost decide.

Benchmark scores as product evidence. The strong version is: "we built an eval set from our own real tickets, and here is what moved it."

Try it yourself

Answer "why did our bot invent a refund policy?" out loud, in ninety seconds, and record it. Then listen back and count brand names — the target is zero. Check that a mechanism appeared inside your first two sentences.

What to learn next

Researcher — Mathematics and papers.

Senior loops add arithmetic and papers to the same questions. These five recur constantly.

"How much memory does the KV cache take?"

The KV cache stores every layer's key and value vectors so earlier tokens are not recomputed for each new token. Its size:

bytes = 2 × L × H_kv × d_head × S × p
  • 2 — one set of keys, one set of values.
  • L — number of transformer layers.
  • H_kv — number of key/value heads (fewer than query heads under grouped-query attention).
  • d_head — dimension per head.
  • S — sequence length in tokens.
  • p — bytes per value (2 for fp16/bf16).

Worked example, an 8B Llama-style model: L=32, H_kv=8, d_head=128, S=8192, p=2 gives 2 × 32 × 8 × 128 × 8192 × 2 ≈ 1.07 GB per sequence. Seventy-five concurrent long chats fill a whole 80 GB accelerator on cache alone — and the model's 16 GB of weights claimed their share first, so in practice about sixty chats hit the wall. That is why paged cache management exists (Kwon et al., 2023, vLLM), and why serving cost scales with context, not only with model size.

"Why is long context expensive, and does it help?"

Self-attention cost grows as O(S² · d) for prefill — S the sequence length, d the model width — so doubling context roughly quadruples prefill compute. Filling the window is not free accuracy either: Liu et al. (2023), Lost in the Middle, measured a U-shaped curve where facts placed mid-context are recalled measurably worse than facts at either end. A bigger window is not automatically a better answer.

"Given a training budget, how big should the model be?"

Kaplan et al. (2020) fit power laws suggesting parameters should grow faster than data. Hoffmann et al. (2022), the Chinchilla paper, corrected the methodology: compute-optimal training wants roughly 20 tokens per parameter, and a smaller model trained on more data beat larger under-trained models at equal compute.

The senior follow-up: compute-optimal is not deployment-optimal. When inference dominates lifetime cost, over-training a smaller model past the Chinchilla point buys cheaper serving forever — the trade the Llama family made deliberately.

"RLHF versus DPO on one whiteboard"

RLHF trains a reward model r(x, y) on human preference pairs, then optimises the policy with a KL-regularised objective:

maximise  E[ r(x, y) ]  −  β · KL( π ‖ π_ref )
  • π — the policy being trained; π_ref — the frozen starting model.
  • β — how hard the policy is tethered to π_ref, which limits reward hacking.
  • KL — the divergence between the two, measuring how far the policy has drifted.

DPO (Rafailov et al., 2023) showed this objective has a closed form on preference pairs directly, removing the reward model and the RL loop:

L = −E[ log σ( β·( log π(y_w|x)/π_ref(y_w|x) − log π(y_l|x)/π_ref(y_l|x) ) ) ]
  • y_w, y_l — the preferred and rejected responses in one pair; σ — the sigmoid.

Trade-off to state: DPO is far easier to run and tune. PPO-style RLHF explores beyond the preference dataset and remains common at frontier scale.

"Why does LoRA work?"

LoRA freezes a weight matrix W ∈ R^(d×k) and learns a low-rank update:

W′ = W + (α / r) · B A        B ∈ R^(d×r),  A ∈ R^(r×k),  r ≪ min(d, k)
  • r — the adapter rank; α — a fixed scaling constant.
  • B A — the product of two thin matrices, an update of rank at most r.

The premise: the change needed for adaptation has low intrinsic rank, even though W does not. Parameter arithmetic interviewers expect: for d = k = 4096 and r = 8, the adapter holds 2 × 4096 × 8 = 65,536 parameters against 16.8M in W — about 0.4% per adapted matrix.

Papers worth naming, with years

What to learn next