Observability for LLM Applications

Capturing user feedback

User feedback is a direct signal from the person who actually received an answer, and capturing it well means tying every rating back to the exact request and prompt version that earned it.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

User feedback is a direct signal from the person who actually received an answer, telling you whether it actually worked.

The analogy you have already lived

A food delivery app asks for a rating after every order. Five stars, or one, with an optional comment.

Nobody has to guess whether the food was good. The person who ate it said so directly, tied to that exact order, at that exact time.

The thumbs-up and thumbs-down buttons under a chatbot's reply are doing the same job, for an answer instead of a meal.

Why it exists

Everything earlier in this section — traces, cost, evals, a judge model — infers quality from the outside. Feedback is different: it comes straight from the one person qualified to say whether an answer actually helped them.

Nothing else in this section can replace that. A judge model can guess well. Only the real user actually knows.

How it works

   answer shown to the user
        |
        v
   [ thumbs up / down button ]  ---->  explicit feedback
   [ did they click "regenerate"? ]  ---->  implicit feedback
        |
        v
   tied back to: the exact request, and the exact prompt version behind it

Feedback is only useful once it is linked to the request and the prompt version that produced the thing being rated. A rating floating on its own, unlinked, cannot be acted on.

A real example you have seen

A messaging app's "regenerate response" button is feedback too, even without a single star or click on "thumbs down". Someone rejecting an answer and asking for another one is telling you the first one did not work.

The honest part

Very few users leave explicit feedback. A thumbs-up rate of even a few percent of all traffic is normal, not broken.

That is why implicit signals — regenerating, copying an answer, abandoning a conversation — matter so much. They fill in for the silent majority who will never click a button either way.

Remember this

  • Feedback is the one signal that comes directly from the real user.
  • It is only useful once tied to the exact request and prompt version it is about.
  • Most users never leave explicit feedback — implicit signals cover the rest.

What to learn next

Developer — Code and libraries.

Setup

No install needed. sqlite3 ships in Python's standard library.

Tying feedback back to a prompt version

Reusing the request_id idea from prediction logging and the version hashes from versioning prompts, attach both explicit ratings and one implicit signal to real requests.

feedback.py
import sqlite3

conn = sqlite3.connect(":memory:")
conn.execute("""
    CREATE TABLE feedback (
        request_id     TEXT,
        prompt_version TEXT,
        signal_type    TEXT,   -- 'explicit_rating' or 'implicit_regenerate'
        value          INTEGER -- +1 thumbs up, -1 thumbs down/regenerate
    )
""")

# request_id ties back to a specific logged prediction, the way prediction
# logging tags every response. prompt_version ties back to the exact
# wording that produced it, the way prompt versioning tags every response.
events = [
    ("req-101", "1acb71eb", "explicit_rating", 1),
    ("req-102", "1acb71eb", "explicit_rating", 1),
    ("req-103", "1acb71eb", "explicit_rating", 1),
    ("req-104", "1acb71eb", "implicit_regenerate", -1),
    ("req-201", "20acf79f", "explicit_rating", -1),
    ("req-202", "20acf79f", "explicit_rating", -1),
    ("req-203", "20acf79f", "implicit_regenerate", -1),
    ("req-204", "20acf79f", "explicit_rating", 1),
]
conn.executemany("INSERT INTO feedback VALUES (?, ?, ?, ?)", events)
conn.commit()

print("Per prompt version:")
rows = conn.execute("""
    SELECT prompt_version,
           COUNT(*) AS total_signals,
           SUM(CASE WHEN value = 1 THEN 1 ELSE 0 END) AS positive,
           SUM(CASE WHEN value = -1 THEN 1 ELSE 0 END) AS negative,
           AVG(value) AS net_score
    FROM feedback
    GROUP BY prompt_version
""").fetchall()

for version, total, pos, neg, net in rows:
    print(f"  {version}   signals={total}   +{pos}/-{neg}   net score: {net:+.2f}")

print("\nBy signal type, version 20acf79f only:")
by_type = conn.execute("""
    SELECT signal_type, COUNT(*), AVG(value)
    FROM feedback WHERE prompt_version = '20acf79f'
    GROUP BY signal_type
""").fetchall()
for signal_type, count, avg in by_type:
    print(f"  {signal_type:<22} count={count}   avg value: {avg:+.2f}")
Output
Per prompt version:
  1acb71eb   signals=4   +3/-1   net score: +0.50
  20acf79f   signals=4   +1/-3   net score: -0.50

By signal type, version 20acf79f only:
  explicit_rating        count=3   avg value: -0.33
  implicit_regenerate    count=1   avg value: -1.00

Exact output from this fixed, hand-written dataset. The two prompt versions from the versioning lesson show sharply opposite net scores here — exactly the kind of result that should trigger rolling one version back.

Line-by-line walkthrough

prompt_version is the same eight-character hash a real system would attach at generation time, from the registry built in the versioning lesson.

signal_type keeps explicit and implicit feedback in one table, distinguishable, rather than two disconnected systems that are hard to reason about together.

The GROUP BY prompt_version query is the entire point of tying feedback to a version in the first place — it turns individual ratings into a per-version comparison, automatically.

Common mistakes

Storing a rating with no link back to the request or prompt version. An unlinked "thumbs down" tells you something went wrong, and nothing about what or when. Always store the request_id alongside a rating.

Treating silence as a neutral or positive signal. Most users do not rate at all. Silence is missing data, not agreement — do not fold it into an average as if it were a middling score.

Only capturing explicit feedback. As shown above, implicit signals like implicit_regenerate catch dissatisfaction the vast, silent majority never bothers to click a button about.

Never sampling for feedback proactively. Waiting passively for a user to click a button under-collects. Occasionally, and sparingly, asking "was this helpful?" on a small random sample of conversations yields much more data than waiting.

Try it yourself

Add a copied_response signal type with value +1, standing in for a user copying an answer to use elsewhere — a strong positive implicit signal. Add two of them to version 1acb71eb, and recompute its net score.

What to learn next

Researcher — Mathematics and papers.

Feedback as a label source, and its biases

User feedback is tempting to treat as ground truth for retraining or evaluation, and it carries well-documented biases that make that risky without correction:

  • Selection bias — people who rate are not a random sample of all users. Extremely satisfied and extremely dissatisfied users rate far more often than the moderately satisfied majority, a pattern well documented in review-platform research generally (Hu, Pavlou and Zhang, 2009, on the J-shaped distribution of online ratings).
  • Presentation bias — where and how a feedback control is shown changes response rates and even response valence, independent of the underlying answer quality.
  • Feedback-loop risk — if a model is retrained directly on its own user feedback, the feedback loop failure mode applies here too: early feedback shapes what gets shown, which shapes future feedback.

Explicit versus implicit signal reliability

Explicit ratings are lower-volume but higher-precision per instance — a thumbs-down is an unambiguous statement of dissatisfaction. Implicit signals (regeneration, abandonment, copy, dwell time) are higher-volume but noisier — a regeneration could mean the answer was wrong, or only that the user wanted a different style, an ambiguity explicit ratings do not carry. Production systems typically weight implicit signals lower per-instance while relying on them for volume, and use explicit signals as the higher-trust calibration set.

Combining feedback with the earlier signals in this section

A mature quality dashboard joins four signal families on request_id: the trace (what happened), token/cost telemetry (what it cost), an eval or judge score (an automated quality estimate), and user feedback (the real outcome). Correlating judge scores against real user feedback, on the overlap where both exist, is the empirical way to answer "is our judge model actually any good?" raised in LLM-as-judge in production — treating feedback as the calibration target the judge is being validated against.

Sample size for a trustworthy per-version comparison

Given a baseline positive-feedback rate $p_0$ and a minimum detectable difference $\delta$, the sample size needed per version for a two-proportion test at power $1-\beta$ and significance $\alpha$ follows the standard formula:

$$n = \frac{(z_{\alpha/2} + z_\beta)^2 \left[p_0(1-p_0) + p_1(1-p_1)\right]}{\delta^2}$$

Where $p_1 = p_0 + \delta$. This is why the eight-event toy comparison in the developer tab is illustrative only — a real per-version feedback comparison at typical feedback rates (often under 5% of traffic) needs a substantially larger volume of requests before a difference like the one shown is statistically trustworthy rather than noise.

Papers

  • Hu, Pavlou and Zhang, Overcoming the J-Shaped Distribution of Product Reviews, Communications of the ACM 2009.
  • Joachims et al., Accurately Interpreting Clickthrough Data as Implicit Feedback, SIGIR 2005 — foundational work on the gap between implicit signals and true relevance judgments.
  • Chaney, Stewart and Engelhardt, How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility, RecSys 2018 — the feedback-loop risk in a retraining context, referenced above.

What to learn next