Building a demo that does not break
A live demo fails for four boring reasons — a download, a cold start, an input you never tried, and no fallback — and each one is fixed by ten lines written before anyone is watching.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A demo breaks for four boring reasons, and none of them is your model: a download, a slow first run, an untested input, and no fallback.
Cooking at home, you can walk to the shop for coriander mid-recipe. Nobody notices. Cooking for forty seated guests, that same walk is a disaster.
So caterers prep everything first. Vegetables chopped, spices measured, pots warm, before a single guest sits down. The cooking during the event is only assembly.
A demo is the seated meal. Everything that can be done before the audience arrives must be done before the audience arrives.
Why it exists
Demos fail in a specific, repeated pattern, and the pattern has nothing to do with model quality.
Somebody opens the laptop, runs the script, and it starts downloading a model — because on their machine it was already cached. The venue wifi is a shared network with three hundred people on it. Two minutes of silence, then a timeout.
Or the first prediction takes eight seconds, because the first prediction always takes longer than the rest, and nobody ever noticed while developing.
Or someone in the audience types something strange, the code raises an exception, and a stack trace fills the projector.
Or all of it works, the model returns a bad answer, and there is no graceful thing to say.
Four failures, four fixes, and every fix is written before the day.
How it works
BEFORE anyone is watching DURING the demo
download every model ──────> nothing is downloaded
load it once, keep it ──────> nothing is loaded
run one throwaway call ──────> first real call is fast
check the input first ──────> bad input gets a message
prepare a fallback ──────> a failure still shows somethingThe pattern is one sentence: do all the slow and risky work in advance, and never let an error reach the screen.
A real example you have seen
Any live TV broadcast. The presenter has a script and the studio has clips ready. When a remote link fails they cut to something prepared and keep talking.
Nobody in the audience knows what went wrong. That smoothness is not luck; it is a prepared fallback that somebody built on a quiet Tuesday.
Remember this
- Download and load everything before the demo. Venue wifi is not your wifi.
- Run one throwaway prediction so the audience never sees the slow first call.
- Check the input, catch every error, and always have something to show.
What to learn next
- Model deployment — the real service this demo is a rehearsal for.
- Model cards — what to say about the model when the demo ends.
- Quantization in practice — how to make a model small enough to carry to the venue.
Developer — Code and libraries.
Setup
pip install scikit-learn pandasVerified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, Python 3.10, CPU. Runs in about three seconds.
A prediction function that survives an audience
Everything expensive happens at import time. Everything risky happens inside a guard.
import numpy as np, pandas as pd
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(0)
n = 500
df = pd.DataFrame({"amount": rng.gamma(2.0, 50.0, n).round(2),
"city": rng.choice(["mumbai", "pune", "nagpur"], n, p=[.6, .3, .1])})
df["fraud"] = ((df.amount > 150) ^ (df.city == "nagpur")).astype(int)
X = pd.get_dummies(df[["amount", "city"]], columns=["city"])
COLUMNS = list(X.columns)
KNOWN_CITIES = sorted(df.city.unique())
MODEL = LogisticRegression(max_iter=1000).fit(X.to_numpy(dtype=float), df.fraud) # loaded ONCE
BASE_RATE = float(df.fraud.mean())
def featurise(req):
row = pd.get_dummies(pd.DataFrame([req]), columns=["city"])
return row.reindex(columns=COLUMNS, fill_value=0).to_numpy(dtype=float)
def validate(req):
if not isinstance(req.get("amount"), (int, float)):
return "amount is missing or not a number"
if req["amount"] < 0:
return "amount is negative"
if req.get("city") not in KNOWN_CITIES:
return f"unknown city {req.get('city')!r}; model knows {KNOWN_CITIES}"
return None
def safe_predict(req):
problem = validate(req)
if problem:
return {"ok": False, "reason": problem, "fallback": round(BASE_RATE, 3)}
try:
p = float(MODEL.predict_proba(featurise(req))[0, 1])
return {"ok": True, "probability": round(p, 3)}
except Exception as exc:
return {"ok": False, "reason": f"{type(exc).__name__}", "fallback": round(BASE_RATE, 3)}
safe_predict({"amount": 100.0, "city": "pune"}) # warm-up: first call is always slowest
requests = [
{"amount": 220.0, "city": "pune"},
{"amount": 40.0, "city": "mumbai"},
{"amount": 300.0, "city": "thane"},
{"amount": -5.0, "city": "pune"},
{"amount": "lots", "city": "pune"},
{"city": "pune"},
]
for req in requests:
print(f"{str(req):42s} -> {safe_predict(req)}"){'amount': 220.0, 'city': 'pune'} -> {'ok': True, 'probability': 0.832}
{'amount': 40.0, 'city': 'mumbai'} -> {'ok': True, 'probability': 0.013}
{'amount': 300.0, 'city': 'thane'} -> {'ok': False, 'reason': "unknown city 'thane'; model knows ['mumbai', 'nagpur', 'pune']", 'fallback': 0.242}
{'amount': -5.0, 'city': 'pune'} -> {'ok': False, 'reason': 'amount is negative', 'fallback': 0.242}
{'amount': 'lots', 'city': 'pune'} -> {'ok': False, 'reason': 'amount is missing or not a number', 'fallback': 0.242}
{'city': 'pune'} -> {'ok': False, 'reason': 'amount is missing or not a number', 'fallback': 0.242}Six requests, four of them hostile, zero stack traces on the projector.
The walkthrough
MODEL is built at module level, not inside safe_predict. This is the difference between a demo that answers in milliseconds and one that retrains on every click. It is also the most common mistake in demo code, because during development a function that loads its own model feels tidier.
The warm-up call before the loop is doing real work. First calls are slower for reasons that stack: imports resolve lazily, NumPy allocates its buffers, and on a GPU the CUDA context and kernels initialise. The size of the effect depends on your hardware and library versions. Measure it on the demo machine rather than trusting any published figure. The fix costs one line and always applies.
validate runs before the model, and its messages are for humans. "unknown city 'thane'; model knows ['mumbai', 'nagpur', 'pune']" is a sentence you can read aloud while it is on screen. KeyError: 'city_thane' is not.
The fallback is the base rate, and it is always present. Every failure returns a number the demo can still display and you can still narrate: "we do not have this city in the model, so it falls back to the overall rate of 24%." A demo that degrades is a demo that continues.
Catching bare Exception here is deliberate. It is poor practice in a library and correct practice at a demo's outermost boundary, where a single uncaught error costs the whole session. Log the traceback to a file, show the audience the fallback.
Being honest about downloads
If your demo uses a pretrained model from Hugging Face or similar, it downloads weights the first time. Approximate sizes, which vary by revision and file format, so check yours with huggingface-cli scan-cache rather than trusting a number in a lesson:
| Model | Rough download |
|---|---|
all-MiniLM-L6-v2 sentence embeddings | around 90 MB |
distilbert-base-uncased | around 260 MB |
whisper-small speech recognition | around 500 MB |
| A 7B language model, 4-bit quantised | around 4 GB |
| A 7B language model, float16 | around 14 GB |
On conference wifi, the last two rows are not a download; they are the end of your slot. Run the demo once on the actual machine, offline, with the network disabled, and confirm it still works. HF_HUB_OFFLINE=1 makes an accidental download fail immediately and loudly at your desk instead of silently and slowly on stage. Model compression and quantization in practice are the real fixes when the model is genuinely too large to carry.
The pre-demo checklist
Twenty minutes, done the day before, not the morning of.
- Run it with the wifi switched off. This single test catches most demo failures.
- Run it on the machine and the account you will present from, at the projector's resolution. A plot that fits your laptop can be unreadable at the back of a room.
- Try five hostile inputs: empty, enormous, wrong type, an unseen category, and a foreign-script string. Each should produce a sentence, never a traceback.
- Pin the versions. A
pip installthat resolves differently on demo day has ended more demos than any model bug. Freeze the environment and reinstall it clean once to prove the freeze works. - Record a screen capture of the demo working. If everything fails, play the video. Nobody minds, and it takes three minutes to make.
- Prepare the inputs you will actually type, and know what each one returns. Live-typing a fresh example is a coin toss performed in public.
- Decide what you will say when it is wrong — because a demo where the model is wrong once and you explain why is more convincing than one where everything works.
Common mistakes
Demonstrating from a notebook. Cell order, leftover state and one accidental shift-enter over the wrong cell. If the notebook must be used, restart and run all immediately before you start, then touch nothing above the demo cell — the reasons are in getting out of the notebook.
A demo that calls an external API. Now the demo's reliability is the network's reliability multiplied by a third party's uptime. Cache the responses for your prepared inputs and fall back to the cache.
Showing raw probabilities to a non-technical audience. "0.832" invites the question "is that good?" — which you then answer for four minutes instead of showing the product. Convert to the decision: flagged, not flagged, needs review.
Skipping validation because "I will only type valid inputs". Somebody always asks to try one. Saying yes is the best moment of a demo, and it requires the code to survive it.
No plan for a bad answer. Models are wrong sometimes. Deciding in advance how you talk about that turns the worst moment of a demo into the most credible one.
Try it yourself
Add a latency_ms field to the successful response, and print the first ten calls with the warm-up line removed. Your exact numbers will depend on the machine, so look at the shape: the first call is slower than the rest, and often by a lot. Then restore the warm-up call and watch that shape disappear.
What to learn next
- Model deployment — the real service this demo is a rehearsal for.
- Model cards — what to say about the model when the demo ends.
- Quantization in practice — how to make a model small enough to carry to the venue.
Researcher — Mathematics and papers.
Where the cold-start cost comes from
The first-call penalty is not one effect but a stack of them, and they respond to different mitigations.
- Import and linking: Python imports of NumPy, PyTorch or TensorFlow resolve dynamic libraries and register operator tables. Amortised once per process, so it argues for a long-lived process, not for per-request loading.
- Allocator warm-up: NumPy and PyTorch caching allocators grow their arenas on first use. PyTorch's CUDA caching allocator in particular reuses freed blocks, so steady-state allocation is far cheaper than the first pass.
- Device context and kernel setup: CUDA context creation and just-in-time PTX compilation for an architecture with no cached cubin are one-off costs paid at first kernel launch, and are the largest term on GPU.
- Compilation and autotuning:
torch.compile, XLA and TensorRT trace, optimise and specialise on first call, and re-trigger on any shape change. torch.compile and tf.function and retracing cover the recompilation traps, which matter here because a demo that accepts variable-length input can pay the compile cost repeatedly. - Page cache and lazy loading: weights memory-mapped from disk fault in on access, so the first forward pass reads the file for real.
The operational consequence is that a warm-up must exercise the same shapes and code paths as the real request; a warm-up on a different batch size can leave the largest term unpaid. Serving systems formalise this with readiness probes that fail until a synthetic warm-up completes, which is the production version of the single throwaway call above.
Graceful degradation as a stated policy
Returning a base rate on failure is a fallback policy, and its correctness is a decision-theoretic question rather than an engineering convenience. Under the abstention framing (Chow, 1970, On Optimum Recognition Error and Reject Tradeoff), a system may decline to predict when expected cost exceeds a threshold, trading coverage for accuracy on the covered set; the optimal reject region under known costs is a threshold on posterior confidence. The demo's validator implements a crude but sound version: reject inputs outside the model's support, where the posterior is not uncertain but undefined.
Two adjacent literatures make this precise. Selective prediction with guarantees (Geifman and El-Yaniv, 2017, Selective Classification for Deep Neural Networks, NeurIPS) supplies coverage-risk curves for choosing the abstention threshold. Conformal prediction (conformal prediction) supplies distribution-free sets with finite-sample validity, from which an "I do not know" is a set that is too large rather than a heuristic. Both are the principled version of what a demo does by hand, and both require the deployment distribution to match the calibration distribution — an assumption an audience member typing an adversarial input violates by construction.
The unseen-category case is the sharpest instance. A one-hot encoding of an unseen level yields an all-zero block, which is not a neutral input: it is the origin of that feature's subspace, and the model's response there is an extrapolation the training data never constrained. Silently predicting on it is strictly worse than abstaining, which is why handle_unknown in scikit-learn's encoders is a decision to make deliberately rather than a default to accept (encoders and scalers).
Demos as evidence, and their systematic bias
A demonstration is a sample of size $n$ from a distribution chosen by the demonstrator, which makes it the most biased evaluation instrument in common use. The literature on this is blunt: cherry-picked qualitative examples are unreliable predictors of aggregate performance, and the reproducibility programmes at major venues (Pineau et al., 2021, JMLR) exist partly because impressive examples outran measured results. Model cards (Mitchell et al., 2019, FAT*) and datasheets (Gebru et al., 2021, CACM) are the structural correction — intended use, population, and known failure modes reported alongside the artefact (model cards).
The honest demo design that follows is narrow and worth stating: show the aggregate metric first and the examples second, include at least one failure case chosen by someone other than you, and state the population the model was trained on before anyone asks whether it works on theirs. This costs a minute of the slot and is the difference between a demo that persuades and one that survives the follow-up meeting.
What to learn next
- Model deployment — the real service this demo is a rehearsal for.
- Model cards — what to say about the model when the demo ends.
- Quantization in practice — how to make a model small enough to carry to the venue.