AI EngineeringRAG · Document search · LLM engineering

RAG system: what it is, when you need it and where it breaks

RAG is a way to make a language model answer from your documents rather than from its memory of the internet. The idea is simple and the failure is predictable: almost every problem in these systems arises not in the model but in how the documents are chunked, retrieved and passed to it.

What RAG is, in plain words

RAG (retrieval-augmented generation) is search and a language model working together. Before answering, the system finds the fragments of your knowledge base that relate to the question and puts them into the model’s context. The model answers from what was found, not from general notions about the world.

An everyday analogy: an employee is asked a question about an internal policy. They do not try to remember it; they open the right pages and answer from them, citing the clause. RAG works the same way, and that is exactly why its answers can be checked.

When you need RAG and when you do not

RAG makes sense when knowledge lives in documents and changes: policies, contracts, technical documentation, a support knowledge base, project history. Update a document and the system answers differently the same day, with no retraining.

RAG is not needed, and only adds complexity, if:

How a RAG system works

A minimal working pipeline has six parts, and each of them is a source of failure.

  1. Collecting documents: exports from storage, email, wikis, file shares. This is also where you decide what to do with scans and tables.
  2. Chunking: each document is cut into pieces that the model will receive whole. Chunks that are too small lose their meaning; chunks that are too large blur the search.
  3. Vectorisation: each chunk is turned into a set of numbers that reflects its meaning and goes into a vector database.
  4. Retrieval: the question is turned into the same kind of vector, and the database returns the closest chunks. Good systems add ordinary full-text search over terms.
  5. Context assembly: the selected chunks are filtered, reordered and put into the prompt together with the instruction.
  6. Generation with citations: the model answers and states which chunks the answer is based on.

Where RAG breaks on real documents

Chunking

The most common cause of bad answers. A policy cut in the middle of a clause yields a chunk without the condition under which the clause applies — and the model answers confidently and wrongly. Chunking must follow the structure of the document: sections, clauses, tables, not a fixed number of characters.

Tables, scans and appendices

A table flattened into plain text loses the link between header and value. A scan without text recognition never reaches the search at all. Corporate archives usually hold more of this material than anyone expects at the start — it is a separate line of work, not a detail.

Versions and duplicates

A contract, its amended version and a draft all sit in the base at once. Search by meaning will find any of them just as readily. Without a date, a status and a version priority, the system will answer from a revoked document — and be formally “correct”.

Meaning versus terms

Vector search handles rephrasing well and exact identifiers badly: part numbers, GOST standard numbers, version names. The cure is hybrid search: vector plus full-text, with the results merged.

Access rights

If filtering by access rights is not built into the search itself, sooner or later the system will quote to an employee a document they should not see. Access rights are part of the index, not post-processing of the answer.

How to tell that the system works

A feeling that “it answers reasonably well” is not acceptance. You need a set of questions with known correct answers and references to the source — it is put together with domain experts before launch. Against this set you measure the things that actually decide the outcome:

Measuring retrieval and generation separately is essential: if the error is in retrieval, changing the prompt is pointless, and vice versa. Without that separation, improving the system turns into guesswork.

RAG, fine-tuning or a knowledge graph

The three approaches solve different problems, and they are regularly confused.

In practice, strong systems combine them: the graph handles structure and relations, vector search handles phrasing, fine-tuning handles style and format. Choosing between a graph and vector search is covered in a separate article.

Where to start

  1. Take one process with recurring questions and a measurable cost per answer — support, product selection from documentation, contract review.
  2. Collect 30–50 real questions with reference answers. This is the most valuable work of the stage, and it cannot be handed over to a contractor entirely.
  3. Build a prototype on a limited set of documents and measure it on your questions.
  4. Only then expand the corpus. A bad system on a large corpus costs many times more to debug.

Frequently asked questions

How is RAG different from ordinary document search?

Ordinary search returns a list of documents and leaves the reading to a person. RAG finds fragments and formulates an answer, stating where it came from. Search still sits inside it — if search works badly, the answer will be bad whatever the model.

Can a RAG system be built inside a closed perimeter?

Yes. Both the vector database and the model can run inside your perimeter, with no calls to external providers. Local models fall behind the best cloud models on complex reasoning, so the scenario is matched to what the model can do — this is discussed before the start, not discovered at the end.

How many documents does it take for this to make sense?

The sense comes not from the number of documents but from how often the questions repeat. If people turn to the same policy every day and the cost of an error is noticeable, a hundred documents is enough. If the requests are one-offs, it is cheaper to keep ordinary search.

Why does the system sometimes answer confidently and wrongly?

Because the language model formulates its answer from what it was given. If retrieval brought the wrong chunk, or a chunk cut off in the middle of a condition, the model fills in the missing part itself. The cure is good chunking and retrieval, a requirement to cite the source and the ability to decline to answer.

How we help with this

Read next

Let us talk about your task

If your task is similar, tell us what needs solving. We will say so plainly if it can be solved more simply than it looks.

Write to us