From thumbs-down to eval set: turning user complaints into test cases
When a client sends you a file full of "dissatisfied" clicks, the first job is to read the traces and label them, not to count the clicks.

In brief
- A thumbs-down tells you which trace to read first. It does not tell you where the system went wrong.
- Read each trace, write down the first thing that went wrong, and label it pass/fail in a spreadsheet before adding it to the eval set.
- Stop reading when no new failure types appear. Start a new round when users repeat the same complaint.
Three weeks into rolling out an internal chatbot for a client, the head of operations sends you an export from the system. It has a few hundred rows, each one a user clicking thumbs-down. The accompanying message is a single line: “The bot gets too many answers wrong, please fix it.”
Many engineers respond on reflex. They tweak the prompt, rerun a few questions from the file, see that things look fine and deploy. Two weeks later the thumbs-down count is unchanged, and nobody can say whether the fix broke anything else.
The right approach is to treat the thumbs-down file as raw material for an eval set, not as a list of bugs to patch one by one. If you want to work as an FDE, practise this skill before a file like that lands in your inbox.
Is a thumbs-down a signal or a label?
Braintrust defines user feedback as explicit signals, such as thumbs up/down or ratings, captured from end users in production and attached to traces. The key phrase is “attached to traces”. A click is only useful when you can open the full input, the retrieved context and the output behind it.
But a click does not tell you where the error is. A user might click because the bot cited the wrong document, because the answer was too long, or simply because they did not like hearing the truth.
That is why Hamel Husain and Shreya Shankar list “sorting by user feedback” among advanced sampling techniques. Feedback helps you find failing traces faster. Working out what the failure actually is remains the reader’s job.
LangSmith describes a common workflow that follows the same logic. Every datapoint with negative feedback goes into an annotation queue for human review instead of straight into a dataset. Datapoints with positive feedback, by contrast, are often filtered out and moved directly into a dataset.
A worked example: a policy Q&A chatbot
Imagine a chatbot that answers employees’ questions about the client’s leave policy. You have a week of logs, and some of the traces received a thumbs-down. Step one: do not take only the thumbs-down traces.
Langfuse describes the first step of error analysis as drawing a representative sample from production traffic. Eugene Yan likewise starts product evals by sampling inputs and outputs from real LLM requests.
The reason is practical. The people who click thumbs-down are a particular group of users, and if you only look at them you will never see many of the errors. Mix the thumbs-down traces with a similar number of random traces.
The code below only illustrates the idea. It assumes you have already exported the traces as a list of dicts:
import csv, random
def build_review_sheet(traces, n_neg=30, n_rand=30, out="review.csv"):
neg = [t for t in traces if t.get("feedback") == "thumbs_down"]
rest = [t for t in traces if t.get("feedback") != "thumbs_down"]
sample = random.sample(neg, min(n_neg, len(neg))) + \
random.sample(rest, min(n_rand, len(rest)))
with open(out, "w", newline="") as f:
w = csv.writer(f)
w.writerow(["trace_id", "feedback", "question", "answer",
"first_error_note", "category", "pass_fail"])
for t in sample:
w.writerow([t["id"], t.get("feedback", ""), t["input"],
t["output"], "", "", ""])
The last three columns are left blank on purpose, because filling them in is human work. In the open coding step, you read each trace and write free-form notes, rather like keeping a journal.
Langfuse offers a rule worth following: record only the first thing that went wrong in each trace. The first error often causes the ones after it, so noting all of them means counting one root cause several times.
Once you have finished reading, your spreadsheet might look like this (the rows below are hypothetical):
| trace_id | feedback | first_error_note | category | pass_fail |
|---|---|---|---|---|
| t-014 | thumbs_down | Answered using the old leave rules; the new document was not retrieved | Wrong document version retrieved | fail |
| t-022 | thumbs_down | Answer was correct; the user disagreed with the policy | Not a system error | pass |
| t-031 | (none) | Invented the number of leave days for seasonal contracts | Fabricated information | fail |
| t-047 | thumbs_down | Did not ask a follow-up when the question omitted the contract type | Failed to clarify the question | fail |
Rows t-022 and t-031 show why you cannot skip reading the traces. One thumbs-down trace turns out to pass, while a trace nobody clicked on is the most serious failure in the table. If you had used the clicks as labels, both would have been classified wrongly.
Pass/fail is enough, and so is a spreadsheet
Eugene Yan recommends using only binary labels, pass/fail or win/lose, and starting simply with a spreadsheet. That advice fits the table above. Each row needs to answer one question: “Does this trace pass?” To answer it consistently, you have to write down the pass criteria for each error category.
With a category column in place, count the traces in each group. The largest group becomes the first thing to fix, and its failing traces become test cases. Braintrust describes exactly this move: traces that receive a thumbs-down are added to the eval set. The difference in this workflow is that a trace enters the eval set only after someone has read and labelled it.
What about traces that got a thumbs-up? You can add them to a dataset, as LangSmith describes, and use them as regression tests that catch a fix breaking something that already worked. Skim them first, though. Users also give a thumbs-up to answers that sound convincing but are wrong.
When do you stop reading?
Hamel Husain and Shreya Shankar call the stopping point theoretical saturation. You reach it when reading more traces no longer surfaces new failure types or changes the existing categories.
As a rough guide, you are close when about twenty traces in a row all fit into categories you have already named. Treat that as a reference point, not a fixed threshold. The more varied the system, the longer you should keep reading.
An eval set is never finished. Langfuse treats recurring complaints as a monitoring signal to start another round of error analysis. When the client writes “it’s doing the same thing as last week”, pull fresh traces into the spreadsheet and read again.
Common traps
The first trap is treating the thumbs-down count as a quality metric. The number rises and falls with how willing users are to click. Use it to choose traces to read, and do not report it to the client as a metric.
The second trap is fixing the list of error categories before you read. A list drawn up in advance usually reflects the engineer’s worries rather than what users actually run into. Let the categories emerge from the free-form notes.
The third trap is handing all the trace reading to someone else. Most of what you learn about a system comes from reading its traces. And in front of a client, an FDE who has read the traces explains failures far more convincingly than one who has only looked at a dashboard.
One round of trace reading, one line on your CV
The fastest way to practise is to run a full round yourself on the logs of any LLM application you can access, side projects included. Take a mixed sample, note the first error, group the notes, label pass/fail, then rerun the eval set after your next fix.
As you go, record how many traces you read, which error categories you found, and the fail rate of the largest category before and after the fix.
Those numbers are what goes on your CV. A line such as “read and labelled production traces, built a pass/fail eval set for retrieval errors, re-measured after every prompt change” shows FDE recruiters that you can work with real data. A generic line like “improved chatbot quality” does not.
By the twentieth trace, you will most likely find your system failing in ways nobody has clicked thumbs-down to tell you about.
5 sources
- Q: Why is error analysis so important in AI evals, and how is it performed? (Hamel Husain & Shreya Shankar) · 2025-06-27
- Product Evals in Three Simple Steps · 2025-11
- Encyclopedia Evalica / Tracing and instrumentation / User feedback (Braintrust)
- LangSmith: Production Monitoring & Automations · 2024-04-02
- Error analysis to evaluate LLM applications (Langfuse) · 2025-08-29