LLM basics for FDEs: tokens, context windows, temperature and why models make things up
When a client's AI assistant invents a clause that does not exist, these four concepts tell you where the fault lies. Without them, all you can do is edit the prompt and hope.
In brief
- An LLM generates one token at a time from the prompt and what it has just written. It does not look up facts.
- Everything in a request takes up space in the context window, and a longer context can make the model less accurate (context rot).
- Models make things up partly because training and evaluation reward guessing over admitting uncertainty, so the system has to be designed so that "I don't know" counts as a valid answer.
Picture the third demo at an insurance company. The head of legal asks the internal assistant about the deadline for filing a claim. The model answers fluently and even cites “Clause 14.3”. The client’s documents contain no Clause 14.3.
Many engineers react by adding “do not make things up” to the prompt and running it again. Sometimes that works. But you don’t know why it worked, so when the error comes back next week you have no idea where to start looking.
An FDE has to answer a harder question: what did the model see, how was it configured to choose its words, and why did it guess instead of saying “I don’t know”?
Tokens, context windows, temperature and hallucination are the tools for that diagnosis. This article explains each one in the order you will use them when debugging at a client site.
The model doesn’t look things up. It keeps writing
The Hugging Face Transformers documentation describes LLMs this way: the model is trained to generate the next token, based on the initial text (the prompt) and on what it has generated so far.
This comes first because “Clause 14.3” was not pulled from any database. The model wrote “Clause”, then saw that “Clause” is usually followed by a number, then a full stop, then another number. To the model the sequence sounds plausible. It never checks whether it is true.
How the model picks the next token is a parameter you control. The default for generate() in Transformers is greedy search, which always picks the most probable token. That suits tasks that need to stay close to the input, such as translation or transcription, but it does poorly on creative tasks.
Sampling gives more varied output, and temperature sets how unpredictable the chosen token is.
Hugging Face suggests a temperature above 0.8 for creative tasks and below 0.4 for tasks that need careful reasoning. Temperature only has an effect when sampling is switched on. A policy-lookup assistant running at a high temperature takes on extra risk for no benefit.
The context window is a desk, not a library
The Claude API documentation defines the context window as all the text the model can refer to while generating a response, including the response itself. It is working memory, and quite separate from training data. Anything that is not on this desk, the model can only half-remember from training, or guess.
What newcomers tend to underestimate is how much sits on the desk. According to the same documentation, everything in the request takes up space: the system prompt, every message in messages (including tool results, images and documents), tool definitions, and the model’s output.
In a multi-turn conversation, the output of one turn becomes the input of the next, so the history keeps growing until it hits the limit.
At that point, adding more documents just to be safe backfires. The Claude documentation is explicit: as the number of tokens grows, accuracy and recall fall, a phenomenon called context rot. A bigger context does not automatically give better results, so choosing what goes in is a real engineering decision.
Taking apart the “Clause 14.3” case
Back to the insurance company. The first step is to log the exact request that was sent, not just the user’s question. Then count the tokens in each part with the token counting API, which the Claude documentation recommends for estimating usage before you send.
The trick is to count several times, adding one component each time, and subtract to see how much each part takes:
import anthropic
client = anthropic.Anthropic()
MODEL = "TEN_MODEL_DU_AN" # thay bằng model dự án đang dùng
def dem(messages, tools=None):
kw = dict(model=MODEL, system=SYSTEM, messages=messages)
if tools:
kw["tools"] = tools
return client.messages.count_tokens(**kw).input_tokens
hoi = {"role": "user", "content": QUESTION}
hoi_kem_tai_lieu = {"role": "user", "content": DOCS + QUESTION}
nen = dem([hoi])
co_tool = dem([hoi], TOOLS)
co_lich_su = dem(HISTORY + [hoi], TOOLS)
day_du = dem(HISTORY + [hoi_kem_tai_lieu], TOOLS)
print("system + câu hỏi:", nen)
print("định nghĩa tool: ", co_tool - nen)
print("lịch sử hội thoại:", co_lich_su - co_tool)
print("tài liệu retrieval:", day_du - co_lich_su)
print("tổng input: ", day_du)
The comment means “replace with the model the project uses”. The printed labels read, in order: system + question, tool definitions, conversation history, retrieval documents, total input. Suppose that in this case the output looks like this (illustrative figures):
system + câu hỏi: 1850
định nghĩa tool: 3200
lịch sử hội thoại: 41600
tài liệu retrieval: 2400
tổng input: 49050
Reading the table: the history from the previous twenty demo turns takes 41,600 of the 49,050 tokens, more than 84% of the budget. The three passages fetched by retrieval take only 2,400 tokens, and when you open them, none of them mentions the claims deadline.
That gives you two suspects. The information the model needed was never on the desk, and most of the desk was taken up by irrelevant history.
Next, open the configuration file. If the assistant is running at temperature 0.9 because someone copied it from a marketing-email demo, that is the third suspect. All three can be fixed with ordinary engineering: trim the history, fix retrieval and lower the temperature.
But even with all three fixed, one question remains: why didn’t the model simply say the documents don’t contain this information?
Why models prefer guessing to saying “I don’t know”
The paper “Why Language Models Hallucinate” by Kalai, Nachum, Vempala and Zhang gives a fairly blunt answer. The authors argue that training and evaluation reward guessing rather than admitting uncertainty.
Models are optimised to be good test-takers, and on a test, guessing when unsure still improves the score.
They also argue that hallucination is nothing mysterious. It comes from ordinary binary classification errors: failing to tell a true statement from a false one. That makes the FDE’s job at a client site fairly clear.
If your eval suite only rewards “correct answer” and counts “I don’t know” as wrong, you are recreating the incentive that makes the model guess.
The sensible approach is to make “not in the documents” a valid answer at both ends. In the prompt, state plainly that the model may decline and must name the clause it relies on. In the eval, include questions that deliberately have no answer, and give a high score to a well-timed refusal.
Common mistakes
The most common mistake is treating the context window like a hard drive and stuffing in as much as possible. It is working memory, and it gets worse when overloaded. The second is forgetting that tool definitions and output take space too, so a request that works in a three-turn demo breaks when a real user reaches turn thirty.
The third is setting temperature by feel or copying it from another project. The fourth is the hardest to spot: building an eval suite that only counts correct answers, so it neither catches nor prevents confident, invented ones.
You can avoid all four by debugging in the order shown in the diagram: what is on the desk first, decoding parameters next, how the system rewards guessing last, and then turning the failure into a test case so it cannot quietly come back.
That process, not a line saying “proficient in prompt engineering”, is what belongs on a CV. For example: “reduced fabricated answers using evals with unanswerable questions and a per-component token budget”.
Clients won’t remember you explaining what a token is. They will remember the time their assistant answered “the current documents do not specify this deadline” instead of inventing Clause 14.3. That is when they start to trust the system.