Monitoring Models in Production
Logging every prediction
Prediction logging means saving every input, output and model version your service ever produces, because every other lesson in this section needs that record to exist first.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Prediction logging means saving a record of every input, prediction and model version your service ever produces.
The analogy you have already lived
A family doctor keeps a register. Every patient, every visit, what was prescribed, and which doctor saw them that day.
Nobody reads most pages of that register on the day they are written. Its value shows up months later. A patient reacts badly to a medicine, and the doctor needs to know exactly what was given, and when.
A prediction log is that register, kept for a machine learning service instead of a clinic.
Why it exists
Every earlier lesson in this section quietly assumed a record already existed. A record of what the model actually predicted, for real inputs, at a real time.
Without that record, measuring drift has nothing to compare against. Monitoring without labels has no predictions to check. A complaint about one wrong answer cannot be investigated, because nobody kept the answer.
Prediction logging is the unglamorous foundation everything else in this section stands on.
How it works
request arrives
|
v
[ model predicts ] ----> answer sent to the caller (fast, first)
|
v
record saved: inputs, prediction, model version, timestamp
(this happens after the answer is already on its way)The caller never waits for the logging step. It happens alongside sending the answer back, not before it.
A real example you have seen
A payments app declines a card, and the customer calls support confused. The support agent pulls up exactly what the fraud model saw and decided, for that one transaction, within seconds.
That lookup is only possible because the prediction was logged the moment it happened. Nobody could recreate it afterward from memory.
The honest part
Logging every prediction sounds free. It is not quite free — storage costs money, and a poorly built logger can slow down every single request.
The trade a good design makes: log fast and log everything, over trying to be clever about which predictions "seem important enough" to keep. You cannot go back and log a moment that already passed.
Remember this
- A prediction log records inputs, prediction, model version and timestamp, for every request.
- Every other monitoring technique in this section needs this record to already exist.
- Logging should never slow down the answer the caller is waiting for.
What to learn next
- Monitoring by segment — the first thing you can build once predictions are logged.
- Alerts people do not ignore — turning this log into something that pages a human.
- Model serving — the service this logging code plugs into.
Developer — Code and libraries.
Setup
No install needed. sqlite3 and uuid ship in Python's standard library.
A minimal logger
sqlite3 is in the Python standard library, which makes it a reasonable starting point for a log you will query with SQL later.
import sqlite3
import json
import time
import uuid
conn = sqlite3.connect(":memory:") # use a real file path in production
conn.execute("""
CREATE TABLE predictions (
request_id TEXT PRIMARY KEY,
logged_at REAL,
model_version TEXT,
features TEXT,
prediction INTEGER,
probability REAL
)
""")
def log_prediction(model_version, features, prediction, probability):
conn.execute(
"INSERT INTO predictions VALUES (?, ?, ?, ?, ?, ?)",
(str(uuid.uuid4()), time.time(), model_version,
json.dumps(features), prediction, probability),
)
conn.commit()
# Three requests a loan-scoring service might have handled today.
log_prediction("v3", {"income": 60.0, "years": 7.0, "age": 34}, 1, 0.9941)
log_prediction("v3", {"income": 8.0, "years": 0.5, "age": 22}, 0, 0.0037)
log_prediction("v3", {"income": 45.0, "years": 3.0, "age": 29}, 1, 0.6120)
rows = conn.execute(
"SELECT model_version, prediction, probability, features FROM predictions ORDER BY logged_at"
).fetchall()
for r in rows:
print(r)
total, approved, avg_p = conn.execute(
"SELECT COUNT(*), SUM(prediction), AVG(probability) FROM predictions"
).fetchone()
print(f"\nlogged today: {total} approved: {approved} mean probability: {avg_p:.3f}")('v3', 1, 0.9941, '{"income": 60.0, "years": 7.0, "age": 34}')
('v3', 0, 0.0037, '{"income": 8.0, "years": 0.5, "age": 22}')
('v3', 1, 0.612, '{"income": 45.0, "years": 3.0, "age": 29}')
logged today: 3 approved: 2 mean probability: 0.537This is exact, deterministic output — every input here is fixed, and nothing in this script involves randomness. Your own log will fill with real request data instead.
Line-by-line walkthrough
request_id is a fresh UUID per row, the handle you would hand a support agent to look up one exact prediction.
features is stored as a JSON string, not as separate columns. New feature columns can then arrive without an ALTER TABLE, at the cost of needing json.dumps / json.loads to read them back.
conn.commit() after every insert is deliberate here for clarity. A real high-traffic service batches commits, covered under common mistakes below.
Wiring it into the serving endpoint
Add one line to the /predict handler from model serving:
@app.post("/predict")
def predict(applicant: Applicant):
row = pd.DataFrame([applicant.model_dump()])[STATE["columns"]]
probability = float(STATE["model"].predict_proba(row)[0, 1])
result = {"repaid": int(probability >= 0.5), "probability": round(probability, 4)}
log_prediction("v3", applicant.model_dump(), result["repaid"], probability) # added
return resultNo output block here — this snippet only makes sense running inside the full service from that lesson, not on its own.
Common mistakes
Logging synchronously, on the same path the caller is waiting on. A slow disk write should never make a prediction slower. Push logging onto a background thread, a queue, or an async task.
Committing to disk after every single row. Under real traffic this is a large fraction of total latency. Batch commits every few hundred rows or every few seconds instead.
Logging the prediction but not the model version. Without a version column, a bad week of predictions cannot be traced to the retrain that caused it.
Storing raw features forever with no retention policy. Feature logs can carry personal data. Redacting personal data from LLM logs covers this for text; the same discipline applies to structured features like income or address.
Try it yourself
Add a latency_ms column, measured with time.perf_counter() around the model call, and log it alongside every prediction. That one column turns this log into a source for latency monitoring too, not only accuracy monitoring.
What to learn next
- Monitoring by segment — the first thing you can build once predictions are logged.
- Alerts people do not ignore — turning this log into something that pages a human.
- Model serving — the service this logging code plugs into.
Researcher — Mathematics and papers.
What belongs in the schema
A production prediction log typically carries more than the minimal example above:
| Field | Why |
|---|---|
request_id | Correlates a prediction with a downstream complaint or trace |
model_version / model_hash | Attributes a bad period to a specific deployed artifact |
feature_vector (post-transform) | Lets skew checks compare the exact values the model saw |
raw_input (pre-transform) | Lets you reconstruct what the feature pipeline actually did |
prediction, probability / logits | The full output, not only the thresholded class |
latency_ms | Feeds latency monitoring from the same pipeline |
served_by (host / pod / region) | Isolates issues to specific infrastructure |
label (nullable, backfilled) | Joined in later, once delayed labels arrive |
Throughput and storage design
A synchronous write-per-request design bounds a service's throughput by disk fsync latency, typically single-digit milliseconds per commit on spinning or networked storage. The standard fix is a write-behind log: predictions go onto an in-memory queue or a local append-only file immediately, and a separate consumer batches them into durable storage (a database, or object storage as Parquet) on its own schedule. This decouples serving latency from storage latency entirely.
At scale, the common architecture is: service emits an event to a message queue (Kafka, Kinesis, or a managed equivalent) synchronously but non-blockingly, and a stream processor or scheduled job lands those events into a queryable store. This also naturally provides the durability and replay properties needed for replaying production traffic.
Sampling versus full logging
Full logging is $O(n)$ in storage for $n$ requests, which is cheap for structured features and can be expensive for large payloads (images, long text). A common compromise: log 100% of metadata (features, prediction, latency) and a sampled fraction of large raw payloads, chosen either uniformly or biased toward low-confidence and edge-case predictions, where investigation value is highest per byte stored.
Papers and systems
- Kreps, Narkhede and Rao, Kafka: A Distributed Messaging System for Log Processing, NetDB 2011 — the architecture underlying most write-behind logging pipelines described above.
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — names prediction logging under "monitoring" as one of the most commonly under-invested pieces of a production ML system.
What to learn next
- Monitoring by segment — the first thing you can build once predictions are logged.
- Alerts people do not ignore — turning this log into something that pages a human.
- Model serving — the service this logging code plugs into.