Zero-shot, few-shot, CoT or ReAct: test all four on one ticket set to learn when a heavier technique is needed
Newer models have not made prompt engineering obsolete. They have made the habit of starting with the heaviest technique expensive.
In brief
- Start with zero-shot: instruction tuning has made newer models good at many simple tasks without examples.
- Few-shot is still worth using if the examples are few, diverse and correctly formatted. CoT helps with reasoning problems, but you must verify the results yourself.
- ReAct underpins today's agent loops. Once you have an agent, choosing what goes into the context matters more than the wording of the prompt.
In September 2025 Anthropic wrote that context engineering is the natural progression of prompt engineering. Yet in the same piece it still made room for a very old technique, few-shot prompting, describing examples as the “pictures” worth a thousand words to an LLM.
Asking which technique is obsolete is therefore the wrong question. For a forward deployed engineer, the practical question is when to move to a heavier technique. Every step adds tokens, latency and maintenance, so you should move only when the numbers show the previous step is not enough.
This guide shows you how to answer that question on your own laptop. You will build a small test set, run four techniques in turn on the same task, and decide on the basis of scores rather than instinct.
What do you need?
Imagine an e-commerce client asks you to classify support tickets into four labels: hoan_tien (refund), giao_hang (delivery), tai_khoan (account) and khac (other). The scenario is hypothetical, but it closely resembles the work you will meet in your first week at a client.
The kit is 20 hand-labelled tickets, API access to any model, and a call_model(prompt) function you write yourself to wrap that API. The code in this article is simplified: call_model and the tools are your own functions, not API names from any provider.
eval/
tickets.jsonl # 20 dòng: {"text": ..., "label": ...}
prompts/
zero_shot.txt
few_shot.txt
cot.txt
run.py # gọi call_model, so với label, in độ chính xác
Step 1: zero-shot is often enough for simple tasks
Zero-shot means a prompt with no examples or demonstrations. It works because the model has been through instruction tuning and RLHF. Instruction tuning has been shown to improve zero-shot ability, so newer models usually handle simple tasks well without being shown a pattern.
Phân loại ticket hỗ trợ sau vào đúng MỘT nhãn:
hoan_tien, giao_hang, tai_khoan, khac.
Chỉ trả về tên nhãn, không giải thích.
Ticket: {text}
(The prompt asks the model to assign exactly one of the four labels and return only the label name, with no explanation.)
Check: run all 20 tickets and record how many are correct. That number is your baseline. If the result already meets the level the client will accept, stop here. Read each wrong answer individually, because the errors are what tell you whether you need step 2.
Step 2: few-shot is still worth using, if the examples are chosen carefully
Few-shot prompting is a way to produce in-context learning: the examples in the prompt steer the model’s answer. Anthropic advises choosing a diverse, canonical set of examples that clearly illustrate the desired behaviour, rather than stuffing every edge case into the prompt.
Phân loại ticket vào đúng MỘT nhãn: hoan_tien, giao_hang, tai_khoan, khac.
Ticket: Đã trả hàng tuần trước mà chưa thấy tiền về.
Nhãn: hoan_tien
Ticket: Đơn báo đã giao nhưng chưa nhận được hàng.
Nhãn: giao_hang
Ticket: Không đăng nhập được dù đã đổi mật khẩu.
Nhãn: tai_khoan
Ticket: {text}
Nhãn:
(The three examples: goods returned last week but no money back yet; an order marked as delivered but not received; unable to log in despite changing the password.)
The most common mistake is choosing examples lazily. Research by Min and colleagues (2022), as summarised in the literature, shows that both the format and the label distribution of the examples significantly affect results. Consider: if all three examples carry the hoan_tien label, the model will tend to lean towards that label.
If one example says “Nhãn:” and another says “Label -”, the output will be just as inconsistent.
IBM also notes that the effectiveness of few-shot prompting depends largely on the quality of the prompt design. Check: compare the score with the baseline and look at exactly which tickets changed result. If the score does not rise, swap the examples before concluding that few-shot is useless.
Step 3: CoT is for reasoning problems, on condition that you verify it
Now the client also wants to know which refund tickets are still within the deadline. At this point few-shot is often no longer reliable. This is exactly the kind of reasoning problem that chain-of-thought was invented for.
Chính sách: khách được hoàn tiền trong 7 ngày kể từ ngày nhận hàng.
Ngày nhận hàng: {received}. Ngày gửi yêu cầu: {requested}.
Hãy suy luận từng bước, rồi ghi kết luận cuối cùng trên một dòng
dạng: KET_LUAN: hop_le hoặc KET_LUAN: khong_hop_le
(The policy: refunds are allowed within 7 days of receiving the goods. The model is given the receipt and request dates, asked to reason step by step, and to finish with a single line: KET_LUAN: hop_le for valid or KET_LUAN: khong_hop_le for invalid.)
IBM warns that CoT can produce reasoning chains that sound plausible but are wrong. Try a hypothetical case: goods received on 3 October, refund requested on 12 October. The gap is 9 days, so the conclusion must be invalid. A smooth argument that miscounts the days can still reach the opposite conclusion.
So whatever can be computed in code should be computed in code. The snippet below is simplified:
from datetime import date
def check_refund(received: date, requested: date, days: int = 7) -> bool:
return (requested - received).days <= days
# so với KET_LUAN của model; đếm số lần lệch
assert check_refund(date(2026, 10, 3), date(2026, 10, 12)) is False
Check: count how often the model’s conclusion disagrees with the check function. If it disagrees often, do not try to write a better prompt. Move the calculation out of the model, which is also the reason to go on to step 4.
Step 4: ReAct underpins agents, but it has a cost
ReAct lets an LLM generate both reasoning traces and concrete actions in an interleaved way, and it is the foundation of today’s agent loops. In our running example, instead of counting the days itself, the model calls a tool to look up the order and then calls the check_refund function from step 3.
# Bản đơn giản hoá: call_model, parse_action và TOOLS đều do bạn tự viết
history = [f"Ticket: {text}"]
for _ in range(5): # luôn đặt giới hạn số vòng
step = call_model(REACT_PROMPT + "\n".join(history))
history.append(step) # Thought + Action
action = parse_action(step)
if action.name == "finish":
break
observation = TOOLS[action.name](**action.args)
history.append(f"Observation: {observation}")
What you gain is interpretability: every Thought, Action and Observation step can be read back. What you lose is flexibility, because ReAct’s rigid structure constrains how the model builds its reasoning steps. Do not use ReAct for the classification in step 1. Save it for tasks that genuinely need external data.
Check: print the full history of the tickets that went wrong. Because every step is readable, first see whether the model called the wrong tool or misread an observation, and only then consider changing the wording of the prompt.
Once you have an agent, the content of the context matters more than the wording
Elastic separates the two ideas neatly. Prompt engineering is how you communicate with the model. Context engineering is the information the model has in hand when it answers, including retrieved documents, tool results and conversation history, together with actively curating what goes into the context window.
Look again at the loop in step 4: history grows with every iteration. So the next worthwhile job is not tweaking a few more words in REACT_PROMPT. It is deciding which observations to keep verbatim, which only need a summary, and which section of the policy to retrieve for each ticket.
How does this skill help with clients?
In your first week at a client, what persuades people is a tickets.jsonl file with 20 rows labelled from their own data, plus a score table comparing the techniques. A flashy agent demo does less. That test set also answers the question the client will put to you: why not use something simpler?
If you are preparing to apply for FDE roles, read job descriptions closely for phrases such as “evaluation”, “prompt and context design” or “agent reliability”. On your CV, replace “proficient in prompt engineering” with a line that has numbers in it, for example: built a 20-case eval set, showed zero-shot was sufficient for the classification step, and moved the refund-deadline calculation from CoT into verification code.
The stronger the model, the further zero-shot will take you. What remains of the FDE’s job is knowing which step to stop at, and having the numbers to show the client why.