Observability for LLM Applications
Capturing user feedback
User feedback is a direct signal from the person who actually received an answer, and capturing it well means tying every rating back to the exact request and prompt version that earned it.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
User feedback is a direct signal from the person who actually received an answer, telling you whether it actually worked.
The analogy you have already lived
A food delivery app asks for a rating after every order. Five stars, or one, with an optional comment.
Nobody has to guess whether the food was good. The person who ate it said so directly, tied to that exact order, at that exact time.
The thumbs-up and thumbs-down buttons under a chatbot's reply are doing the same job, for an answer instead of a meal.
Why it exists
Everything earlier in this section — traces, cost, evals, a judge model — infers quality from the outside. Feedback is different: it comes straight from the one person qualified to say whether an answer actually helped them.
Nothing else in this section can replace that. A judge model can guess well. Only the real user actually knows.
How it works
answer shown to the user
|
v
[ thumbs up / down button ] ----> explicit feedback
[ did they click "regenerate"? ] ----> implicit feedback
|
v
tied back to: the exact request, and the exact prompt version behind itFeedback is only useful once it is linked to the request and the prompt version that produced the thing being rated. A rating floating on its own, unlinked, cannot be acted on.
A real example you have seen
A messaging app's "regenerate response" button is feedback too, even without a single star or click on "thumbs down". Someone rejecting an answer and asking for another one is telling you the first one did not work.
The honest part
Very few users leave explicit feedback. A thumbs-up rate of even a few percent of all traffic is normal, not broken.
That is why implicit signals — regenerating, copying an answer, abandoning a conversation — matter so much. They fill in for the silent majority who will never click a button either way.
Remember this
- Feedback is the one signal that comes directly from the real user.
- It is only useful once tied to the exact request and prompt version it is about.
- Most users never leave explicit feedback — implicit signals cover the rest.
What to learn next
- Redacting personal data from LLM logs — feedback comments often contain exactly the personal data that needs handling carefully.
- Replaying production traffic — using badly-rated requests as a targeted test set for the next fix.
- Feedback loops in production — the risk of training directly on the feedback captured here.
Developer — Code and libraries.
Setup
No install needed. sqlite3 ships in Python's standard library.
Tying feedback back to a prompt version
Reusing the request_id idea from prediction logging and the version hashes from versioning prompts, attach both explicit ratings and one implicit signal to real requests.
import sqlite3
conn = sqlite3.connect(":memory:")
conn.execute("""
CREATE TABLE feedback (
request_id TEXT,
prompt_version TEXT,
signal_type TEXT, -- 'explicit_rating' or 'implicit_regenerate'
value INTEGER -- +1 thumbs up, -1 thumbs down/regenerate
)
""")
# request_id ties back to a specific logged prediction, the way prediction
# logging tags every response. prompt_version ties back to the exact
# wording that produced it, the way prompt versioning tags every response.
events = [
("req-101", "1acb71eb", "explicit_rating", 1),
("req-102", "1acb71eb", "explicit_rating", 1),
("req-103", "1acb71eb", "explicit_rating", 1),
("req-104", "1acb71eb", "implicit_regenerate", -1),
("req-201", "20acf79f", "explicit_rating", -1),
("req-202", "20acf79f", "explicit_rating", -1),
("req-203", "20acf79f", "implicit_regenerate", -1),
("req-204", "20acf79f", "explicit_rating", 1),
]
conn.executemany("INSERT INTO feedback VALUES (?, ?, ?, ?)", events)
conn.commit()
print("Per prompt version:")
rows = conn.execute("""
SELECT prompt_version,
COUNT(*) AS total_signals,
SUM(CASE WHEN value = 1 THEN 1 ELSE 0 END) AS positive,
SUM(CASE WHEN value = -1 THEN 1 ELSE 0 END) AS negative,
AVG(value) AS net_score
FROM feedback
GROUP BY prompt_version
""").fetchall()
for version, total, pos, neg, net in rows:
print(f" {version} signals={total} +{pos}/-{neg} net score: {net:+.2f}")
print("\nBy signal type, version 20acf79f only:")
by_type = conn.execute("""
SELECT signal_type, COUNT(*), AVG(value)
FROM feedback WHERE prompt_version = '20acf79f'
GROUP BY signal_type
""").fetchall()
for signal_type, count, avg in by_type:
print(f" {signal_type:<22} count={count} avg value: {avg:+.2f}")Per prompt version: 1acb71eb signals=4 +3/-1 net score: +0.50 20acf79f signals=4 +1/-3 net score: -0.50 By signal type, version 20acf79f only: explicit_rating count=3 avg value: -0.33 implicit_regenerate count=1 avg value: -1.00
Exact output from this fixed, hand-written dataset. The two prompt versions from the versioning lesson show sharply opposite net scores here — exactly the kind of result that should trigger rolling one version back.
Line-by-line walkthrough
prompt_version is the same eight-character hash a real system would attach at generation time, from the registry built in the versioning lesson.
signal_type keeps explicit and implicit feedback in one table, distinguishable, rather than two disconnected systems that are hard to reason about together.
The GROUP BY prompt_version query is the entire point of tying feedback to a version in the first place — it turns individual ratings into a per-version comparison, automatically.
Common mistakes
Storing a rating with no link back to the request or prompt version. An unlinked "thumbs down" tells you something went wrong, and nothing about what or when. Always store the request_id alongside a rating.
Treating silence as a neutral or positive signal. Most users do not rate at all. Silence is missing data, not agreement — do not fold it into an average as if it were a middling score.
Only capturing explicit feedback. As shown above, implicit signals like implicit_regenerate catch dissatisfaction the vast, silent majority never bothers to click a button about.
Never sampling for feedback proactively. Waiting passively for a user to click a button under-collects. Occasionally, and sparingly, asking "was this helpful?" on a small random sample of conversations yields much more data than waiting.
Try it yourself
Add a copied_response signal type with value +1, standing in for a user copying an answer to use elsewhere — a strong positive implicit signal. Add two of them to version 1acb71eb, and recompute its net score.
What to learn next
- Redacting personal data from LLM logs — feedback comments often contain exactly the personal data that needs handling carefully.
- Replaying production traffic — using badly-rated requests as a targeted test set for the next fix.
- Feedback loops in production — the risk of training directly on the feedback captured here.
Researcher — Mathematics and papers.
Feedback as a label source, and its biases
User feedback is tempting to treat as ground truth for retraining or evaluation, and it carries well-documented biases that make that risky without correction:
- Selection bias — people who rate are not a random sample of all users. Extremely satisfied and extremely dissatisfied users rate far more often than the moderately satisfied majority, a pattern well documented in review-platform research generally (Hu, Pavlou and Zhang, 2009, on the J-shaped distribution of online ratings).
- Presentation bias — where and how a feedback control is shown changes response rates and even response valence, independent of the underlying answer quality.
- Feedback-loop risk — if a model is retrained directly on its own user feedback, the feedback loop failure mode applies here too: early feedback shapes what gets shown, which shapes future feedback.
Explicit versus implicit signal reliability
Explicit ratings are lower-volume but higher-precision per instance — a thumbs-down is an unambiguous statement of dissatisfaction. Implicit signals (regeneration, abandonment, copy, dwell time) are higher-volume but noisier — a regeneration could mean the answer was wrong, or only that the user wanted a different style, an ambiguity explicit ratings do not carry. Production systems typically weight implicit signals lower per-instance while relying on them for volume, and use explicit signals as the higher-trust calibration set.
Combining feedback with the earlier signals in this section
A mature quality dashboard joins four signal families on request_id: the trace (what happened), token/cost telemetry (what it cost), an eval or judge score (an automated quality estimate), and user feedback (the real outcome). Correlating judge scores against real user feedback, on the overlap where both exist, is the empirical way to answer "is our judge model actually any good?" raised in LLM-as-judge in production — treating feedback as the calibration target the judge is being validated against.
Sample size for a trustworthy per-version comparison
Given a baseline positive-feedback rate $p_0$ and a minimum detectable difference $\delta$, the sample size needed per version for a two-proportion test at power $1-\beta$ and significance $\alpha$ follows the standard formula:
$$n = \frac{(z_{\alpha/2} + z_\beta)^2 \left[p_0(1-p_0) + p_1(1-p_1)\right]}{\delta^2}$$
Where $p_1 = p_0 + \delta$. This is why the eight-event toy comparison in the developer tab is illustrative only — a real per-version feedback comparison at typical feedback rates (often under 5% of traffic) needs a substantially larger volume of requests before a difference like the one shown is statistically trustworthy rather than noise.
Papers
- Hu, Pavlou and Zhang, Overcoming the J-Shaped Distribution of Product Reviews, Communications of the ACM 2009.
- Joachims et al., Accurately Interpreting Clickthrough Data as Implicit Feedback, SIGIR 2005 — foundational work on the gap between implicit signals and true relevance judgments.
- Chaney, Stewart and Engelhardt, How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility, RecSys 2018 — the feedback-loop risk in a retraining context, referenced above.
What to learn next
- Redacting personal data from LLM logs — feedback comments often contain exactly the personal data that needs handling carefully.
- Replaying production traffic — using badly-rated requests as a targeted test set for the next fix.
- Feedback loops in production — the risk of training directly on the feedback captured here.