Hands-on: selecting, balancing and ordering few-shot examples from real client data
Ten examples drawn at random from a client's ticket backlog can push a model towards the majority label, and changing only their order can shift results dramatically.

In brief
- Skewed examples produce a model skewed towards the majority label, so sample per class.
- Example order alone can move accuracy from near chance to near state-of-the-art, so test several orders.
- Beyond accuracy, compare the predicted label distribution with the true one and measure recall on the rare class.
Illustrative figures: accuracy varies by a few points, but predictions of the rare class range from 2 to 9.
Graphic: FDE Times
In 2021 the authors of “Calibrate Before Use” reported an uncomfortable finding: reordering the examples in a prompt was enough to move a model’s accuracy from close to random guessing to close to state-of-the-art. The examples were the same and so were the labels. Only the order had changed.
At a client site you rarely get to pick examples under clean conditions. The data is real tickets, emails and meeting notes, and the labels are almost always imbalanced: one class dominates, while a few others are rare but important. Grab a handful of examples at random and paste them into the prompt, and the model learns that imbalance too.
The workflow below runs on a laptop. You measure the data, sample per class, remove duplicates, shuffle the order and check whether the model is leaning towards one label.
What you will build, and what you need
Take a hypothetical case. The client is a telecoms company that hands you 1,000 labelled tickets: 700 “information request”, 200 “complaint” and 100 “service cancellation”. The task is to classify new tickets with an LLM. By the end you will have a script that selects a few-shot example set, generates several orderings and measures how stable the results are.
You need Python 3 and a tickets.csv file with two columns, text and label. The code below uses only the standard library. The model-calling function is left empty on purpose, because every client uses a different provider. All of the code is simplified to teach the workflow and is not production code.
Step 1: run zero-shot first, before jumping to few-shot
Zero-shot means a prompt with no examples. The DAIR.AI guide recommends adding examples only when zero-shot falls short. The zero-shot score is your baseline: if few-shot cannot beat it, the example set is doing harm rather than helping.
Before running anything, measure the data:
import csv
from collections import Counter
rows = list(csv.DictReader(open("tickets.csv", encoding="utf-8")))
print(Counter(r["label"] for r in rows))
# Counter({'hoi_thong_tin': 700, 'khieu_nai': 200, 'huy_dich_vu': 100})
(The labels are Vietnamese: hoi_thong_tin is information request, khieu_nai is complaint, huy_dich_vu is service cancellation.)
Check: the class counts must add up to the number of rows, and there must be no stray labels caused by typos, such as “khieu nai” mixed in with “khieu_nai”. Client data often has this problem. No prompt can fix an error in the labels.
Step 2: sample per class, because random sampling copies the imbalance
With a 70/20/10 split, ten randomly drawn examples will on average include about seven “information request” tickets. Learn Prompting describes exactly what follows: when the example distribution is skewed towards one class, the model also leans towards predicting that class.
For the client, that means “service cancellation” tickets, the ones that most urgently need a retention response, get filed with the harmless ones.
The first debiasing method Learn Prompting offers is to balance the number of examples across classes:
import random
def sample_per_class(rows, k_per_class, seed=0):
rng = random.Random(seed)
by_label = {}
for r in rows:
by_label.setdefault(r["label"], []).append(r)
picked = []
for label, items in by_label.items():
picked += rng.sample(items, min(k_per_class, len(items)))
return picked
balanced = sample_per_class(rows, k_per_class=3)
Do not treat balancing as the final answer, though. DAIR.AI also notes that in experiments with random labels, drawing labels from the true distribution worked better than drawing them uniformly. So build a second set that follows the real proportions for comparison, and let the numbers on the client’s test set decide.
Check: print a Counter of balanced; each class should have exactly 3 examples.
Step 3: examples should resemble real data, but not each other
According to DAIR.AI, both the label space and the distribution of the input text in the examples affect few-shot results. So take examples from real tickets, keeping the end customers’ spelling mistakes and abbreviations, rather than tidy sample sentences you wrote yourself.
Learn Prompting warns, however, that examples which are too similar can lead the model to overgeneralise, and the context window also limits how many examples you can fit. The simplified version below removes near-duplicate pairs with difflib. In a real project you would replace it with embedding similarity:
from difflib import SequenceMatcher
def dedupe(examples, threshold=0.8):
kept = []
for ex in examples:
if all(SequenceMatcher(None, ex["text"], k["text"]).ratio() < threshold
for k in kept):
kept.append(ex)
return kept
If a class runs short after deduplication, sample more for that class. Learn Prompting also mentions two automated approaches worth exploring once the data grows.
KNN picks the examples most similar to the input query. Vote-K picks diverse, representative examples from unlabelled data, which helps when the client only has raw logs.
Step 4: shuffle the order and measure; never trust a single run
Learn Prompting suggests that a random ordering of examples usually works better. Here is a hypothesis to test on the client’s own data: if you put all three “service cancellation” examples at the end of the prompt, does the model predict that label more often than with an even shuffle?
Rather than guess, add this deliberate ordering as a configuration to compare.
Zhao and colleagues found that order-driven variance is large enough that you have to treat order as a variable to measure.
def build_prompt(examples, query):
head = ("Phân loại ticket vào một trong: hoi_thong_tin, khieu_nai, huy_dich_vu.\n"
"Đánh giá từng ticket độc lập, không ưu tiên nhãn nào.\n\n")
shots = "".join(f"Ticket: {e['text']}\nNhãn: {e['label']}\n\n" for e in examples)
return head + shots + f"Ticket: {query}\nNhãn:"
def call_model(prompt):
raise NotImplementedError # gọi API của nhà cung cấp khách đang dùng
for seed in range(5):
order = balanced[:]
random.Random(seed).shuffle(order)
# chạy build_prompt + call_model trên tập test, ghi lại kết quả theo seed
The prompt’s first line asks the model to classify each ticket into one of the three labels. The second line, “Evaluate each ticket independently, without favouring any label”, has a purpose of its own: Learn Prompting notes that you can explicitly instruct the model not to be biased. (The code comments say, in order: call the API of the client’s provider; run build_prompt and call_model on the test set and record results per seed.)
The Ticket: ... / Nhãn: ... format (“Nhãn” means “Label”) must also be identical in every example. DAIR.AI points out that format has a significant effect on performance, and Learn Prompting regards a fixed output structure as the biggest benefit of few-shot.
Check: print one complete prompt and read it. A single extra space or a missing blank line in one example is enough to break consistency.
Step 5: look at the prediction distribution, not just accuracy
“Calibrate Before Use” explains that this instability comes from the model’s bias towards certain answers, and that calibrating against the context improves accuracy.
The technique is called contextual calibration, and it is the term to search for next. Unlike the steps above, which adjust the example set, it targets the model’s bias towards particular answers directly, so it is worth trying when the examples are already balanced but the model still leans towards one label.
You do not need a full calibration implementation to spot the symptom. For each seed, print three things: accuracy, how many times the model predicted each label, and recall on the rare class:
from collections import Counter
def report(seed, y_true, y_pred, rare="huy_dich_vu"):
acc = sum(t == p for t, p in zip(y_true, y_pred)) / len(y_true)
hit = sum(t == p == rare for t, p in zip(y_true, y_pred))
total = sum(t == rare for t in y_true)
print(f"seed {seed}: acc={acc:.2f} pred={Counter(y_pred)} "
f"recall_{rare}={hit}/{total}")
# Kết quả minh họa, không phải số đo thật:
# seed 0: acc=0.81 pred=Counter({'hoi_thong_tin': 78, 'khieu_nai': 18, 'huy_dich_vu': 4}) ...
# seed 1: acc=0.74 pred=Counter({'hoi_thong_tin': 86, 'khieu_nai': 12, 'huy_dich_vu': 2}) ...
# seed 2: acc=0.84 pred=Counter({'hoi_thong_tin': 72, 'khieu_nai': 19, 'huy_dich_vu': 9}) ...
These are illustrative results, not real measurements. Read them against a hypothetical test set of 100 tickets split 70/20/10. With seed 1 the model predicts “service cancellation” only twice. Even if both are correct, at least 8 of the 10 customers who want to leave are missed, while an accuracy of 0.74 still looks acceptable because the majority class props it up.
That recall figure is what the client cares about. And the gap between the best and the worst seed is what you must report, not just the best number.
The most common mistakes
| Mistake | Symptom | Fix |
|---|---|---|
| Drawing examples at random from imbalanced data | The model piles its predictions onto the majority class | Sample per class, and compare against a set in the true proportions |
| Writing “clean” examples yourself | Works in the demo, fails on real tickets | Take examples from client data and keep their original wording |
| Many near-duplicate examples | The model overgeneralises and wastes context | Deduplicate and favour diverse examples |
| Running only one ordering | Results cannot be reproduced between runs | Run several seeds and report the variance too |
Keep expectations in check as well. DAIR.AI acknowledges that standard few-shot is not a perfect technique, especially for tasks that require reasoning. If the client’s labels depend on a multi-step chain of reasoning, however carefully you balance the examples, you will solve only part of the problem.
At a client site, the deliverable is a table, not just a prompt
The table states which example set was used, how many orderings were run, the range accuracy moved within, and how often the rare class was missed. When the client asks why the model worked well yesterday and badly today, that table is the answer.
When reading FDE job descriptions, look for phrases such as “evaluation”, “prompt reliability” or “working with customer data”. On your CV, do not write “experienced in prompt engineering”. Write that you measured label skew, fixed it, and showed that results stayed stable when the example order changed.
Anyone can write a prompt that works in a demo. FDEs are paid to show that the prompt still works on the thousandth ticket, including the one the majority class does not represent.
6 sources
- Few-Shot Prompting - Prompt Engineering Guide (DAIR.AI)
- Zero-Shot Prompting - Prompt Engineering Guide (DAIR.AI)
- Prompt Debiasing: Ensuring Fair and Balanced LLM Outputs (Learn Prompting) · 2024-08-07
- Introduction to Few-Shot Prompting Techniques (Learn Prompting) · 2025-03-25
- Technique #3: Examples in Prompts: From Zero-Shot to Few-Shot (Learn Prompting) · 2025-03-06
- Calibrate Before Use: Improving Few-Shot Performance of Language Models · 2021-02-19