FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

LLM-as-judge in practice: calibrating a judge against the client's expert pass/fail labels

A judge with 90% agreement can still miss every wrong answer. If you want a client to trust your eval numbers, measure the judge against the person on their side who knows the domain best.

In brief

  • Treat building a judge like a small ML project: the expert's labels are ground truth, and the judge prompt is what you tune over many rounds.
  • Grade pass/fail and measure TPR and TNR separately. When failures are rare, an overall agreement rate is easily misleading.
  • Calibration is never one-and-done: have a human review 50 to 100 traces each week, prioritising hard cases and edge cases.
ShareLinkedInFacebookX
GraphicSame judge: 90% agreement but 0% TNR
Overall agreement rateTPR and TNR measured separately
Result on 100 answers90%, sounds respectableTPR 100%, TNR 0%
The 10 answers the expert failedHidden inside the overall numberJudge catches 0 of 10, exposed at once
When failures are rareEasily misleadingPass and fail cases measured separately
Question it answersWhat percentage of the time does the judge match the human?How many of the expert's errors does the judge catch?

On 100 answers with 10 failures, a judge that passes everything still scores 90% agreement; only splitting out TPR and TNR shows it catches no errors at all.

Graphic: FDE Times

Suppose you have 100 answers from an agent. The client’s expert marks 10 of them as fail. Your judge marks all 100 as pass. Agreement is 90%, which sounds respectable, yet the judge has not caught a single wrong answer.

Hamel Husain names this trap plainly: when failures are rare, agreement can be misleading. Anthropic, in its guidance on evals for agents, also says LLM-based judges must be closely calibrated with human experts.

The five steps below take you from collecting the expert’s labels to a judge reliable enough for the client to base decisions on.

For anyone aiming to become an FDE, this is a skill worth learning early. There is unlikely to be a public benchmark for one particular insurer’s claims-assessment process. The most valuable yardstick is the judgement of the best person on the client’s side, and your job is to turn that judgement into a judge that runs automatically.

What you will build, and what you need

By the end you will have three things: a file of pass/fail labels with critiques, a binary judge prompt, and a script that measures the judge’s TPR and TNR against the human labels. Evidently AI describes building an LLM judge as a small ML project. The framing is right: there is labelled data, a model, and a loop of evaluation and iteration.

You need Python, access to an LLM through whatever API you already use and, most importantly, time from a domain expert. The code below is trimmed for illustration. The call_llm function is a placeholder for your real client, not the API of any particular library.

Step 1: Find the one person who gets to say “good enough”

Hamel Husain advises identifying a principal domain expert and bringing that person in as early as possible. Pick one person, not a committee. Picture a chatbot that answers customer questions about insurance terms: the right person might be the head of the underwriting team, the one everyone in the department asks when a case gets hard.

Then sample traces. Braintrust suggests starting with 50 to 100 traces a week, prioritising important samples and edge cases. Do not pull a random batch of easy cases, because the judge will fail precisely where you did not look.

When failures are rare, deliberately oversample traces you suspect will fail, such as answers customers complained about or questions touching complex policy terms. A test set of a few dozen traces with only one or two failures will give a TNR that swings wildly with each case, or one you cannot compute at all.

Step 2: The expert grades pass/fail and writes a critique

Hamel recommends dropping elaborate scoring scales and keeping a single, clear pass or fail decision. Evidently also observes that binary scores tend to be more stable and consistent, for LLMs and human graders alike. A 1–10 scale sounds more granular, but the expert’s “6” and the judge’s “6” rarely mean the same thing.

The most valuable part is the critique. According to Hamel, critiques should be detailed enough to drop straight into the judge’s few-shot prompt. Store each trace as one JSONL line (the examples here come from a Vietnamese insurance chatbot; the critique says the answer said “yes” but skipped the waiting-period condition, so the customer would misunderstand their benefits):

{"id": "t017", "input": "Hợp đồng có chi trả khi ...?", "output": "Có, ...", "label": "fail", "critique": "Trả lời 'có' nhưng bỏ qua điều kiện thời gian chờ; khách sẽ hiểu sai quyền lợi."}

Check after this step: reread 10 random critiques. If you cannot tell why an answer failed, neither will the judge. Go back to the expert straight away, while they still remember.

Step 3: Write the judge prompt around a rubric

Anthropic recommends building clear, structured rubrics for each dimension of a task. Combined with the binary principle, each dimension becomes a yes/no question, and the final verdict is pass or fail. Save the trimmed prompt below as judge_prompt.txt. It tells the model it is grading an insurance assistant, asks two yes/no criteria (does the answer state the clause’s conditions correctly; does it avoid promising benefits outside the contract), supplies graded examples, and requests JSON output:

Bạn chấm câu trả lời của trợ lý bảo hiểm.
Tiêu chí (mỗi tiêu chí: CÓ/KHÔNG):
1. Nêu đúng điều kiện áp dụng của điều khoản?
2. Không hứa quyền lợi ngoài hợp đồng?
Ví dụ đã chấm:
{few_shot_critiques}
Câu hỏi: {input}
Câu trả lời: {output}
Trả về JSON: {"critique": "...", "label": "pass" | "fail"}

Make the judge write its critique before giving a label. When the judge gets it wrong, its critique tells you which criterion it misread. Any trace used as a few-shot example must be removed from the test set, just as you separate train and test data; otherwise the numbers will look artificially good.

Step 4: Put the judge’s labels next to the human’s

Now you measure. Run the judge on the traces not used as few-shot examples, record its label on each line, then compute two separate numbers instead of one overall agreement rate. This is the alignment Evidently describes, comparing judge output against hand-labelled ground truth, and Braintrust likewise treats human scores as the ground truth for calibrating LLM scorers.

import json

def call_llm(prompt):  # thay bằng client thật, trả về chuỗi JSON
    raise NotImplementedError

TEMPLATE = open("judge_prompt.txt", encoding="utf-8").read()
FEW_SHOT = open("few_shot.txt", encoding="utf-8").read()

def run_judge(rows):
    for r in rows:
        prompt = (TEMPLATE.replace("{few_shot_critiques}", FEW_SHOT)
                          .replace("{input}", r["input"])
                          .replace("{output}", r["output"]))
        result = json.loads(call_llm(prompt))
        r["judge"] = result["label"]
        r["judge_critique"] = result["critique"]
    return rows

def rate(hit, total):
    return hit / total if total else None  # tránh chia cho 0

def metrics(rows):
    tp = sum(r["label"] == "pass" and r["judge"] == "pass" for r in rows)
    tn = sum(r["label"] == "fail" and r["judge"] == "fail" for r in rows)
    pos = sum(r["label"] == "pass" for r in rows)
    neg = sum(r["label"] == "fail" for r in rows)
    return {"TPR": rate(tp, pos), "TNR": rate(tn, neg)}

rows = [json.loads(line) for line in open("test.jsonl", encoding="utf-8")]
print(metrics(run_judge(rows)))

(The comments read “replace with a real client, returns a JSON string” and “avoid division by zero”.) The script uses replace rather than str.format because the prompt already contains JSON curly braces. A production version should catch cases where the model returns something that is not JSON; that is omitted here for brevity. If TNR comes back as None, the test set has no failing cases: that is a signal to return to Step 1 and oversample failures, not evidence of a perfect judge.

TPR is the share of cases the expert passed that the judge also passed. TNR is the share of cases the expert failed that the judge also failed. Back to the opening example: a judge that passes all 100 answers has a TPR of 100% and a TNR of 0%, and the problem hidden by the 90% agreement rate is exposed at once.

Step 5: Iterate the prompt until nothing surprises you

In each round, filter out the cases where the judge disagrees with the expert and read the judge_critique for each. Fix exactly the criterion that was misunderstood, or add a few-shot example of that specific error type, then rerun the whole test set.

Evidently describes this as tuning the judge exactly as you would tune a product prompt, while Anthropic acknowledges that model-based grading often needs careful iteration before its accuracy can be verified.

If endless prompt fixes still leave results fluctuating, look at the model. Confident AI notes that with traditional metrics such as GEval, weaker models struggle to produce reliable results. Another option is DeepEval’s DAG metric, described as fully deterministic thanks to its decision-tree structure executed by an LLM.

A rubric made of several yes/no criteria is already close to that approach.

Three common mistakes when doing this with clients

The most common mistake is reporting only a single agreement number, as shown above. The second is letting engineers label in place of the expert because it is “faster”. The judge then ends up calibrated to you, not to the person the client trusts.

The third is treating calibration as a one-off. The client’s business changes, new types of question appear, and a judge that once matched well will gradually drift.

Anthropic says LLM-based rubrics for subjective tasks, such as research agents, need frequent calibration against expert judgement; SuperAnnotate describes human reviewers as having the final say on ambiguous cases and continually refining the grading criteria.

So schedule a weekly review of 50 to 100 traces with the expert from the very first meeting.

What does this skill look like on a CV?

When reading FDE or solutions engineer job descriptions, look for phrases such as “evals”, “LLM-as-judge”, “human-in-the-loop” and “work with domain experts”. On your CV, do not write “experience with evals”.

Write something like: built a pass/fail judge for a domain chatbot, calibrated against the head underwriter’s labels on N traces, raised TNR from X to Y over K rounds.

A line like that shows a recruiter you can do three things an FDE needs: work with the client’s people, measure the right thing, and iterate until the numbers can be trusted. At your first working session with a client, do not open with a judge demo. Ask who everyone in the department turns to when a case gets hard.

6 sources
Read next on the roadmap · Stage 6: MeasurementMeasuring ROI for customers: count work that meets the bar, subtract checking and reworkWhen the customer's CFO asks whether the project was worth the money, "users say it feels faster" will convince nobody.