What RAG is, in plain words
RAG (retrieval-augmented generation) is search and a language model working together. Before answering, the system finds the fragments of your knowledge base that relate to the question and puts them into the model’s context. The model answers from what was found, not from general notions about the world.
An everyday analogy: an employee is asked a question about an internal policy. They do not try to remember it; they open the right pages and answer from them, citing the clause. RAG works the same way, and that is exactly why its answers can be checked.
When you need RAG and when you do not
RAG makes sense when knowledge lives in documents and changes: policies, contracts, technical documentation, a support knowledge base, project history. Update a document and the system answers differently the same day, with no retraining.
RAG is not needed, and only adds complexity, if:
- the answer is computed from structured data — that needs a database query, not search by meaning;
- there are few documents and they fit into the model’s context in full;
- the task calls for actions in systems rather than knowledge — that is an agent’s job, and RAG would only be one of its tools;
- you need a strictly deterministic answer by formula or policy — a rule in code is more reliable.
How a RAG system works
A minimal working pipeline has six parts, and each of them is a source of failure.
- Collecting documents: exports from storage, email, wikis, file shares. This is also where you decide what to do with scans and tables.
- Chunking: each document is cut into pieces that the model will receive whole. Chunks that are too small lose their meaning; chunks that are too large blur the search.
- Vectorisation: each chunk is turned into a set of numbers that reflects its meaning and goes into a vector database.
- Retrieval: the question is turned into the same kind of vector, and the database returns the closest chunks. Good systems add ordinary full-text search over terms.
- Context assembly: the selected chunks are filtered, reordered and put into the prompt together with the instruction.
- Generation with citations: the model answers and states which chunks the answer is based on.
Where RAG breaks on real documents
Chunking
The most common cause of bad answers. A policy cut in the middle of a clause yields a chunk without the condition under which the clause applies — and the model answers confidently and wrongly. Chunking must follow the structure of the document: sections, clauses, tables, not a fixed number of characters.
Tables, scans and appendices
A table flattened into plain text loses the link between header and value. A scan without text recognition never reaches the search at all. Corporate archives usually hold more of this material than anyone expects at the start — it is a separate line of work, not a detail.
Versions and duplicates
A contract, its amended version and a draft all sit in the base at once. Search by meaning will find any of them just as readily. Without a date, a status and a version priority, the system will answer from a revoked document — and be formally “correct”.
Meaning versus terms
Vector search handles rephrasing well and exact identifiers badly: part numbers, GOST standard numbers, version names. The cure is hybrid search: vector plus full-text, with the results merged.
Access rights
If filtering by access rights is not built into the search itself, sooner or later the system will quote to an employee a document they should not see. Access rights are part of the index, not post-processing of the answer.
How to tell that the system works
A feeling that “it answers reasonably well” is not acceptance. You need a set of questions with known correct answers and references to the source — it is put together with domain experts before launch. Against this set you measure the things that actually decide the outcome:
- whether the right chunks were found at all — retrieval quality, separately from answer quality;
- whether the answer rests on what was found or the model added something of its own;
- whether the answer carries a reference to the source and whether it points to the right place;
- whether the system can decline when the documents hold no answer — this is a separate skill, and it has to be tested on purpose.
Measuring retrieval and generation separately is essential: if the error is in retrieval, changing the prompt is pointless, and vice versa. Without that separation, improving the system turns into guesswork.
RAG, fine-tuning or a knowledge graph
The three approaches solve different problems, and they are regularly confused.
- RAG — when knowledge changes and has to be cited. Updating the data costs next to nothing.
- Fine-tuning — when you need to change the model’s behaviour and style or teach it a narrow format. It is poor at giving the model knowledge of facts, and it goes stale along with the data.
- Knowledge graph — when the relations between entities matter and the answer takes several steps of reasoning: who is linked to whom, what depends on what, which part sits in which assembly.
In practice, strong systems combine them: the graph handles structure and relations, vector search handles phrasing, fine-tuning handles style and format. Choosing between a graph and vector search is covered in a separate article.
Where to start
- Take one process with recurring questions and a measurable cost per answer — support, product selection from documentation, contract review.
- Collect 30–50 real questions with reference answers. This is the most valuable work of the stage, and it cannot be handed over to a contractor entirely.
- Build a prototype on a limited set of documents and measure it on your questions.
- Only then expand the corpus. A bad system on a large corpus costs many times more to debug.