Knowledge SystemsKnowledge graphs · Vector databases · Data architecture

Knowledge graph or vector search: which one fits your task

Both approaches answer the request “find me what is relevant”, but they understand relevance differently: vector search looks for what is similar in meaning, a graph for what is connected by structure. Confusing the two costs projects months of work.

What a knowledge graph is

A knowledge graph is data recorded as entities and the relations between them: “a part belongs to an assembly”, “an employee is responsible for a process”, “a topic builds on another topic”. The meaning is stored in the relations themselves, so a graph can answer questions that take several steps: not “what is similar to this” but “what happens to production if this machine goes down”.

What vector search is

A vector database stores text fragments as numbers that reflect their meaning. A query is turned into the same kind of set of numbers, and the database returns the closest fragments. This makes search robust to rephrasing: “how to book a holiday” will find the policy clause on “provision of annual paid leave”, even though not a single word matches.

The fundamental difference

Which questions each approach can handle

Vector search is stronger when

A knowledge graph is stronger when

How they are combined

In working systems this is not an either-or choice. A typical split of duties looks like this:

  1. The vector index takes the user’s question in free form and finds an entry point — fragments and the entities they mention.
  2. From that point the graph builds out the context: related objects, dependencies, constraints, the current version.
  3. The model receives both the text fragments and the structured context — and answers with references to the sources.

This pairing is noticeably more robust than pure semantic search on questions where completeness matters, and noticeably cheaper to fill than an attempt to force absolutely everything into the graph.

Where a graph breaks

An ontology “for every occasion”

The temptation to describe the whole domain leads to a schema that cannot be filled and that nobody understands. The approach that works is a minimal ontology for specific questions, extended as new questions appear.

Extracting entities from text

Building a graph automatically with a language model gives a draft, not a result: one entity arrives in three spellings, relations are duplicated, some relations are made up. You need normalisation, deduplication and spot checks by a human — this is work, not a box to tick.

Keeping it current

A graph that is not updated along with its sources starts to lie more convincingly than search does: it looks structured and authoritative. The update rule has to be designed together with the graph.

How to choose in four questions

  1. Is the answer to a typical question a fragment of text or a traversal of relations? The first means vector search, the second means a graph.
  2. Does an incomplete answer have a cost? If missing one item from a list is expensive, you need a graph.
  3. Who will maintain the structure? If nobody, the graph will not take off — start with vector search.
  4. How often do the documents change? The more often, the more cheap updates matter, which means vector search or a hybrid.

Our experience

We run both technologies in our own products: a graph database for the domain model and the relations between topics, a vector database for searching by the user’s phrasing. We arrived at the split of duties between them not from articles but from debugging our own systems, where pure semantic search stopped coping exactly where the answer had to be complete.

Frequently asked questions

Can you get by with a vector database alone?

Yes, if users’ questions come down to “find and explain” over text. As soon as questions like “list everything that depends on X” or “in what order is this done” appear, pure vector search starts to miss parts of the answer — and that shows up on a set of test questions.

Does a knowledge graph have to be built by hand?

No, but it cannot be built fully automatically either. The scheme that works: a model extracts candidates, rules normalise and deduplicate them, a human checks a sample and the disputed cases. The share of manual work falls as the ontology stabilises.

Which databases do you use?

In our systems, a graph database for relations and a vector database for semantic search. The specific product matters less than the data model and the update pipeline: replacing an engine costs less than reworking a badly designed ontology.

How we help with this

Read next

Let us talk about your task

If your task is similar, tell us what needs solving. We will say so plainly if it can be solved more simply than it looks.

Write to us