What happens, technically
A language model does not store facts as records that can be checked. It predicts how text continues: at each step it chooses the next fragment based on statistical patterns learned in training. For this process, plausibility and truthfulness are different things, and it is plausibility that gets optimised.
That is why “I don’t know” is a statistically rare continuation: in the texts the model learned from, a question is almost always followed by an answer. The model reproduces this pattern and confidently fills in whatever is missing, the way it completes the end of a word.
Why it is not called a bug
The same mechanism gives useful properties: the ability to paraphrase, to generalise, to handle tasks that were not in the training data, to fill in meaning from an incomplete description. A model physically unable to say anything beyond the texts it absorbed would be useless as an assistant.
The practical conclusion is not to resign yourself to it, but to stop looking for a magic setting and start designing a system in which an invented answer finds it hard to survive.
Five types of hallucination that behave differently
- An invented fact: a policy clause, an article of law or a version number that does not exist. The most dangerous type in corporate tasks.
- An invented source: a link or a document title that does not exist. This can be checked automatically — the existence of a source is easy to confirm.
- Mixed sources: two real documents glued into one answer, so that a condition from one is applied to the subject of the other.
- Outdated knowledge presented as current: the model answers as things stood at the time of its training.
- Invented reasoning: a correct answer with a wrong justification. It occurs more often than people think and is almost never caught by checking the final answer.
What actually reduces invented answers
Give the model a source and require a citation
An answer assembled from retrieved fragments, with a mandatory reference to them, is the baseline measure. It does not remove the problem, but it turns it into a form that can be checked: if there is no citation, or the citation does not support the claim, the answer is rejected before a person sees it.
Check that what the model named exists
Document numbers, part numbers, identifiers, links, table names — all of these can be checked against real systems in code. Such a check catches a whole class of hallucinations with no human involved, and it costs next to nothing compared with the consequences.
Allow refusal and reward it
The system needs an explicit path for “this is not in the sources”. It has to be tested on purpose: the set of test questions must include ones where the correct answer is a refusal. Without them, the model will get good scores for confident invention.
Narrow the task
The broader the question, the more freedom the model has to fill in. Splitting the work into narrow steps with a checkable result reduces invented answers more than any instruction in the prompt.
Deterministic constraints on the output
Answer format, allowed values, ranges, required fields — these are checked by code, not by a request to the model. Anything critical must be impossible to violate, not merely undesirable to violate.
What only creates a feeling of control
- “Only tell the truth” in the prompt: the instruction affects style, not the prediction mechanism.
- Asking for confidence as a percentage: the model makes that up too — the numbers are not linked to the real probability of error.
- Checking the answer with the same model on the same data: it tends to agree. You need an independent source of truth — a document, a calculation, a query to a system.
- A larger model: it lowers the frequency but makes the errors more convincing. Checking becomes harder.
- Lowering the temperature: it makes the output more stable, but a consistently wrong answer is still wrong.
How to measure
Without measurement, any measure is a matter of faith. A minimal working set of metrics on a fixed set of questions with reference answers:
- the share of answers fully supported by the sources they cite;
- the share of answers that cite a non-existent or irrelevant source;
- the share of correct refusals on questions whose answer is not in the data;
- the share of wrong refusals — the other side, which needs watching, or the system becomes uselessly cautious;
- the share of correct answers with a wrong justification — caught only by spot manual checks of the reasoning.
These figures are taken again at every change of model, prompt or retrieval rules. Model behaviour changes between versions, and a system that worked yesterday can quietly degrade after an update.
What it looks like in a project
- Before development: collect 30–50 real questions with reference answers, including ones where the correct answer is a refusal.
- In the architecture: mandatory citation, checks of identifiers in code, deterministic format constraints, a decision log.
- In the pilot: human confirmation where the cost of error is high; metrics measured on the real flow.
- In operation: a regression run of the set at every change, and monitoring of the refusal rate — a rise in it is usually the first sign that something has broken in retrieval.