What a knowledge graph is
A knowledge graph is data recorded as entities and the relations between them: “a part belongs to an assembly”, “an employee is responsible for a process”, “a topic builds on another topic”. The meaning is stored in the relations themselves, so a graph can answer questions that take several steps: not “what is similar to this” but “what happens to production if this machine goes down”.
What vector search is
A vector database stores text fragments as numbers that reflect their meaning. A query is turned into the same kind of set of numbers, and the database returns the closest fragments. This makes search robust to rephrasing: “how to book a holiday” will find the policy clause on “provision of annual paid leave”, even though not a single word matches.
The fundamental difference
- Vector search answers the question “what is similar to this”. A graph answers the question “what is connected to this, and through what”.
- A vector knows nothing about structure: two fragments with the same words are close to it even if they belong to different contracts.
- A graph does not tolerate imprecise wording: if an entity has not been extracted and a relation has not been laid down, the question stays unanswered.
- A vector index is cheap to fill — chunking the documents is enough. A graph needs an ontology: a decision on which types of entities and relations exist.
Which questions each approach can handle
Vector search is stronger when
- knowledge lives in text: policies, instructions, correspondence, a support knowledge base;
- users ask in their own words, not in the system’s terms;
- the corpus is large and mixed, and there is nobody to label it;
- you need a fast start and documents updated on the fly.
A knowledge graph is stronger when
- the answer requires a path: cause → effect, part → assembly → product, topic → prerequisite topic;
- complete traversal matters: list everything that depends on an object, missing nothing;
- you need constraints and rules: what is incompatible with what, which order is mandatory;
- the same entities appear in different documents and have to be merged into one.
How they are combined
In working systems this is not an either-or choice. A typical split of duties looks like this:
- The vector index takes the user’s question in free form and finds an entry point — fragments and the entities they mention.
- From that point the graph builds out the context: related objects, dependencies, constraints, the current version.
- The model receives both the text fragments and the structured context — and answers with references to the sources.
This pairing is noticeably more robust than pure semantic search on questions where completeness matters, and noticeably cheaper to fill than an attempt to force absolutely everything into the graph.
Where a graph breaks
An ontology “for every occasion”
The temptation to describe the whole domain leads to a schema that cannot be filled and that nobody understands. The approach that works is a minimal ontology for specific questions, extended as new questions appear.
Extracting entities from text
Building a graph automatically with a language model gives a draft, not a result: one entity arrives in three spellings, relations are duplicated, some relations are made up. You need normalisation, deduplication and spot checks by a human — this is work, not a box to tick.
Keeping it current
A graph that is not updated along with its sources starts to lie more convincingly than search does: it looks structured and authoritative. The update rule has to be designed together with the graph.
How to choose in four questions
- Is the answer to a typical question a fragment of text or a traversal of relations? The first means vector search, the second means a graph.
- Does an incomplete answer have a cost? If missing one item from a list is expensive, you need a graph.
- Who will maintain the structure? If nobody, the graph will not take off — start with vector search.
- How often do the documents change? The more often, the more cheap updates matter, which means vector search or a hybrid.
Our experience
We run both technologies in our own products: a graph database for the domain model and the relations between topics, a vector database for searching by the user’s phrasing. We arrived at the split of duties between them not from articles but from debugging our own systems, where pure semantic search stopped coping exactly where the answer had to be complete.