AI EngineeringLLM · Quality · Evaluation

LLM hallucinations: the mechanism and the engineering measures against them

A hallucination is not a model failure but the model’s normal mode of operation, carried through to a result that is inconvenient for us. Understanding the mechanism matters more than any list of “prompts against hallucinations”: it shows which measures work and which only create a feeling of control.

What happens, technically

A language model does not store facts as records that can be checked. It predicts how text continues: at each step it chooses the next fragment based on statistical patterns learned in training. For this process, plausibility and truthfulness are different things, and it is plausibility that gets optimised.

That is why “I don’t know” is a statistically rare continuation: in the texts the model learned from, a question is almost always followed by an answer. The model reproduces this pattern and confidently fills in whatever is missing, the way it completes the end of a word.

Why it is not called a bug

The same mechanism gives useful properties: the ability to paraphrase, to generalise, to handle tasks that were not in the training data, to fill in meaning from an incomplete description. A model physically unable to say anything beyond the texts it absorbed would be useless as an assistant.

The practical conclusion is not to resign yourself to it, but to stop looking for a magic setting and start designing a system in which an invented answer finds it hard to survive.

Five types of hallucination that behave differently

What actually reduces invented answers

Give the model a source and require a citation

An answer assembled from retrieved fragments, with a mandatory reference to them, is the baseline measure. It does not remove the problem, but it turns it into a form that can be checked: if there is no citation, or the citation does not support the claim, the answer is rejected before a person sees it.

Check that what the model named exists

Document numbers, part numbers, identifiers, links, table names — all of these can be checked against real systems in code. Such a check catches a whole class of hallucinations with no human involved, and it costs next to nothing compared with the consequences.

Allow refusal and reward it

The system needs an explicit path for “this is not in the sources”. It has to be tested on purpose: the set of test questions must include ones where the correct answer is a refusal. Without them, the model will get good scores for confident invention.

Narrow the task

The broader the question, the more freedom the model has to fill in. Splitting the work into narrow steps with a checkable result reduces invented answers more than any instruction in the prompt.

Deterministic constraints on the output

Answer format, allowed values, ranges, required fields — these are checked by code, not by a request to the model. Anything critical must be impossible to violate, not merely undesirable to violate.

What only creates a feeling of control

How to measure

Without measurement, any measure is a matter of faith. A minimal working set of metrics on a fixed set of questions with reference answers:

These figures are taken again at every change of model, prompt or retrieval rules. Model behaviour changes between versions, and a system that worked yesterday can quietly degrade after an update.

What it looks like in a project

  1. Before development: collect 30–50 real questions with reference answers, including ones where the correct answer is a refusal.
  2. In the architecture: mandatory citation, checks of identifiers in code, deterministic format constraints, a decision log.
  3. In the pilot: human confirmation where the cost of error is high; metrics measured on the real flow.
  4. In operation: a regression run of the set at every change, and monitoring of the refusal rate — a rise in it is usually the first sign that something has broken in retrieval.

Frequently asked questions

Can hallucinations be eliminated completely?

No. The mechanism that produces them is the same mechanism that makes the model useful. A realistic goal is to bring the frequency down to a level the process can accept, and to make the remaining errors visible and reversible.

Does RAG solve the hallucination problem?

Partly. It gives the model something to rely on and the ability to cite, but if retrieval brought the wrong fragment, or a fragment cut off in the middle of a condition, the model will confidently fill in the missing part. That is why retrieval quality is measured separately from answer quality.

Which data quality metric has the strongest effect on hallucinations?

The completeness and currency of the corpus: if the answer is not in the data or is out of date, the system has two paths left — refuse or invent. Next in influence comes the correctness of document chunking: a fragment without the condition under which it applies invites the model to fill in.

Does multi-agent checking help?

Only when the checker has an independent source of truth: a test run, a comparison against a document, a recalculation. Checking with the same model on the same data produces agreement, not control.

How we help with this

Read next

Let us talk about your task

If your task is similar, tell us what needs solving. We will say so plainly if it can be solved more simply than it looks.

Write to us