LLM system design interviews
How to design products built on language models under interview pressure — walked through twice in full, with a RAG support bot and a code autocomplete engine.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An LLM system design interview asks you to plan a product built on a language model — including its cost, its speed, and what happens when it confidently gets things wrong.
Think about the last time you called customer care. A recorded menu handled the easy part. A junior agent took over next, reading from a script. Only a genuinely hard problem reached the senior specialist, and anything the specialist promised you went into a logged, recorded call.
Designing an LLM product is designing that call centre. The model is the specialist: capable, expensive, and occasionally over-confident. Your design decides which questions reach it, what papers sit on its desk, and who checks its promises.
Why this round exists
The old ML design round asked you to train a model. This round asks something newer: you did not train the model, you bought or downloaded it, and the engineering lives around it.
The failures are new too. The model answers fluently from an outdated policy. A user's message tricks it into ignoring its rules. The bill triples because conversations got longer. Companies interview for people who saw these coming.
The shape of the design
question arrives
│
▼
can something cheap answer it? ──yes──► cached or scripted reply
│ no
▼
fetch the documents that matter (retrieval)
▼
model writes an answer from them (generation)
▼
checks: grounded? safe? in scope? (guardrails)
▼
user — and every answer is logged and graded laterTwo ideas carry the whole round. Give the model the right pages, because it cannot know your private, current facts — that is RAG. And never treat its output as automatically true — checking is part of the design, not an optional extra.
A real example you have seen
An online store's help chat that quotes your actual order and its actual return window is this exact diagram. So is the code suggestion that appears, grey and hopeful, while a programmer pauses typing. Same skeleton, wildly different speed requirements — which is why this lesson walks through both.
Remember this
- The model is bought, not trained — the design is everything around it.
- Right documents in, checked answers out. Retrieval and guardrails are the design.
- Cost and speed are design inputs from minute one, not launch-week surprises.
What to learn next
- ML case-study interviews — the round where the problem arrives with no shape at all.
- vLLM — the serving stack behind the numbers above.
- What is RAG? — the retrieval layer both designs leaned on.
Developer — Code and libraries.
The skeleton for LLM rounds
Adapt the ML system design skeleton with three LLM-specific steps:
- Requirements, plus risk. What may the bot do — read order status, or issue refunds? Wrong answers cost more when the model can act. See function calling.
- Model access. Hosted API or self-hosted open weights — see model deployment. APIs win on speed-to-ship; self-hosting wins on data control and high-volume unit cost.
- Context strategy. Where does truth come from — retrieval, tools, or the prompt?
- The cost and latency envelope. Done below, with arithmetic.
- Guardrails. Grounding checks, refusal paths, injection defences, human escalation.
- Evaluation. A golden set of real questions, scored before every prompt or model change.
- The feedback loop. Log, grade, and mine failures into the eval set.
The envelope calculator
python3 --version # standard library onlytickets_per_day = 20_000
chunks = 4 # retrieved passages pasted into each prompt
tokens_per_chunk = 350
prompt_overhead = 400 # system rules + the question itself
output_tokens = 250
input_tokens = tickets_per_day * (chunks * tokens_per_chunk + prompt_overhead)
out_tokens = tickets_per_day * output_tokens
in_price = 0.15 / 1_000_000 # dollars per input token — INVENTED round number
out_price = 0.60 / 1_000_000 # dollars per output token — INVENTED round number
daily = input_tokens * in_price + out_tokens * out_price
print(f"input tokens/day : {input_tokens:,}")
print(f"output tokens/day : {out_tokens:,}")
print(f"model cost/day : ${daily:,.2f}")
print(f"model cost/month : ${daily * 30:,.0f}")input tokens/day : 36,000,000 output tokens/day : 5,000,000 model cost/day : $8.40 model cost/month : $252
The prices are invented round numbers so the arithmetic is honest — real per-token prices change often, so look up your provider's current sheet and rerun. The structure is the lesson: input tokens dominate the volume, and every retrieved chunk you add taxes every single request. In the interview, write this table on the board with whatever prices the interviewer accepts.
Design 1: a RAG support bot
An e-commerce company wants a bot answering 20,000 support tickets a day from its policies and order data.
Requirements to ask. Which languages — English, Hindi, and the mixed Hinglish real customers type? What actions are allowed — read-only order lookups, or issuing refunds? What is the target — deflect 40% of tickets without a human? What must never happen — invented policies, leaked customer data?
Architecture.
OFFLINE — index build
policy docs + help pages ─► chunk ─► embed ─► vector index
(re-run on doc change; hash each doc and skip unchanged ones)
ONLINE — per question
question ─► retrieve top chunks ──► prompt = rules + chunks + order data + question
│ │
│ weak matches ▼
└──► human agent model drafts answer, cites chunks
│
grounding + safety checks ──pass──► user
│ fail
└──► human agent queueThe retrieval mechanics — chunking, embeddings, the index — are covered in RAG and vector databases; in the interview, name them and move on to what this design adds: the two escape hatches to a human, one before the model and one after.
Data loop. Log question, retrieved chunks, answer, and the user's thumbs-up or escalation. Grade a sample weekly. Every confirmed failure becomes an eval-set entry, so the golden set grows from real traffic. Measure retrieval separately from generation — when the bot is wrong, first check whether the right chunk was even fetched.
Failure modes. Retrieval miss producing a confident answer from the wrong policy — the classic silent failure. Stale index after a policy change, fixed by triggering re-indexing from document updates. Prompt injection — a user message, or even text inside a product review the bot retrieves, containing instructions. Tool misuse — refunds above a threshold need human approval by design, not by prompt.
What interviewers probe. "How do you know it is ready?" — golden-set scores, then shadow mode alongside human agents, then 5% of live traffic. "Where does your latency go?" — retrieval tens of milliseconds, generation a few seconds; stream the answer so the user sees progress. "Cost per ticket versus a human?" — pull out the envelope above and finish it.
Design 2: code autocomplete
An IDE feature suggests the next few lines while a developer pauses typing. Same skeleton, opposite physics.
Requirements to ask. Latency target — suggestions must begin appearing within roughly 200–300 ms or they arrive after the developer resumed typing, and the feature feels broken. Privacy — may customer code leave the machine at all? Some buyers require self-hosting. Success metric — acceptance rate, the fraction of shown suggestions developers keep. Not offline accuracy.
The envelope, spoken aloud. 5,000 developers, each triggering perhaps 1,500 completion requests over a working day, is 7.5 million requests concentrated in office hours — a few hundred requests per second sustained, with sharp peaks. Most requests are cancelled because the developer kept typing. Cancellation must therefore be nearly free.
Architecture.
pause in typing ─► debounce (skip if still typing) ─► context builder
current file, recent edits, imports
│
▼
small fast model, streaming
prefix cache reuses the unchanged
file context between keystrokes
│
ghost text in the editor
│
accepted or dismissed — logged either wayModel choice, defended. A small model, not a frontier one. The latency budget rules out big models, and acceptance rate improves surprisingly little with model size once suggestions are plausible — speed is the feature. Techniques that make small models viable: quantization, knowledge distillation, and serving stacks like vLLM with prefix caching, since the file prefix barely changes between requests.
Failure modes. Latency regressions are product-killers — a developer who disables the feature rarely re-enables it, so p95 latency is a launch-blocking metric. Plausible-but-wrong code that compiles erodes trust silently. Secrets in context — API keys sitting in the file must be scrubbed before leaving the machine. Cost explosion if debouncing or cancellation breaks quietly.
What interviewers probe. "Why not the best model you can buy?" — the latency and cost arithmetic, plus the acceptance-rate observation. "How do you evaluate offline?" — exact-match is wrong because many different completions are correct; execution-based checks (does it parse, do tests pass) beat string comparison. "API or self-hosted?" — per-token API pricing at 7.5 million requests a day versus fixed GPU cost; at this volume, run the numbers and self-hosting usually wins, before privacy even enters.
Common mistakes
No cost table. LLM designs without token arithmetic are fiction. One envelope, any reasonable prices, spoken aloud.
One model for everything. Strong designs route: cache first, small model for the easy 80%, big model for the hard 20%.
Guardrails as decoration. "We'll add safety later" fails the round. The escalation path is part of the first diagram.
Fine-tuning as a first resort. Reach for prompts and retrieval first; fine-tuning earns its complexity only after those plateau.
No evaluation story. If nothing is measured before a prompt change ships, the design cannot be improved safely. Say "golden set" early.
Try it yourself
Design "summarise every internal meeting into notes and action items" with the seven steps. Write the cost envelope first, before the architecture — notice how deciding the token budget early changes which model and context choices remain available to you.
What to learn next
- ML case-study interviews — the round where the problem arrives with no shape at all.
- vLLM — the serving stack behind the numbers above.
- What is RAG? — the retrieval layer both designs leaned on.
Researcher — Mathematics and papers.
Senior LLM design rounds go quantitative in four places.
Prefill versus decode, and where the ceiling is
Generation has two phases with different physics. Prefill processes the whole prompt in parallel — compute-bound, cost scaling with attention's O(S² · d) in sequence length S and width d. Decode emits one token at a time, reloading the model's weights per step — memory-bandwidth-bound. A useful single-request ceiling:
tokens/sec ≲ memory_bandwidth / bytes_per_parameter_passAn 8B model at fp16 moves roughly 16 GB per token; on a 2 TB/s accelerator that bounds a lone request near 125 tokens/sec. Throughput therefore comes from batching many requests per weight pass — which is why continuous batching and paged KV memory (Kwon et al., 2023, the vLLM paper) define modern serving. The KV cache arithmetic that limits batch size is worked in LLM interview questions.
Speculative decoding
A small draft model proposes k tokens; the large model verifies them in one parallel pass, keeping the longest accepted prefix. With per-token acceptance rate α, expected tokens per large-model pass:
E[tokens] = (1 − α^(k+1)) / (1 − α)α— probability the large model accepts a drafted token;k— draft length.
At α = 0.8, k = 4, that is about 3.4 tokens per pass instead of 1. Crucially, the output distribution is provably unchanged — this is a latency trick, not a quality trade. Leviathan et al. (2023), Fast Inference from Transformers via Speculative Decoding.
Routing and cascades, as expected cost
Route queries to a cheap model, escalating when it is inadequate. With cheap-model cost c_s, expensive-model cost c_b, and a the fraction the cheap model handles acceptably:
E[cost per query] = c_s + (1 − a) · c_bWith a = 0.8 and c_b = 10 · c_s, expected cost is 3 · c_s — a 70% saving over sending everything to the big model, bought with an adequacy classifier whose errors now become your quality floor. Chen et al. (2023), FrugalGPT, formalises the cascade. The interview follow-up is always the failure mode: the router's false "adequate" rate silently caps end quality.
Evaluating the pipeline, not the model
Benchmarks measure the model; your product is the pipeline. Separate the stages — retrieval recall@k against labelled question-chunk pairs, grounding as entailment of each claim by the retrieved context, and answer quality by pairwise preference. LLM-as-judge scales the grading but drifts, so it is calibrated against a human-labelled set; Zheng et al. (2023), Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, documents both the method and its biases. The RAG-specific metrics stack is in RAG.
Papers worth naming
- Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention, 2023 — arxiv.org/abs/2309.06180
- Leviathan, Kalman and Matias, Fast Inference from Transformers via Speculative Decoding, 2023 — arxiv.org/abs/2211.17192
- Chen, Zaharia and Zou, FrugalGPT, 2023 — arxiv.org/abs/2305.05176
- Lewis et al., Retrieval-Augmented Generation, 2020 — arxiv.org/abs/2005.11401
- Zheng et al., Judging LLM-as-a-Judge, 2023 — arxiv.org/abs/2306.05685
- Liu et al., Lost in the Middle, 2023 — arxiv.org/abs/2307.03172
What to learn next
- ML case-study interviews — the round where the problem arrives with no shape at all.
- vLLM — the serving stack behind the numbers above.
- What is RAG? — the retrieval layer both designs leaned on.