What we are talking about
A multi-agent system is an architecture in which a task is solved by several autonomous participants, each with its own role, tools and permissions: one plans, another searches, a third checks. A single agent is one language model with a set of tools and one decision-making loop.
The argument between the two is almost always settled not by ideology but by three measurable quantities: how many tasks are carried through to the end without a human, how much one task costs and how long it takes.
Where multi-agent designs lose
The task fits in one context
If all the data and tools you need fit into the context of one model, splitting the work between agents adds nothing but retelling. Every handover between agents is compression: the next participant receives not the original data but someone else’s summary of it, and works with losses.
Errors compound along the chain
This is the main failure mechanism. Say each step is done correctly nine times out of ten — that sounds decent. Five sequential steps already give roughly six successful outcomes out of ten, because the probabilities multiply. A single agent with one accountable step makes mistakes less often simply because there are fewer steps.
Coordination costs more than the work itself
Messages between agents are extra model calls, each with its own context. Cost and latency grow faster than the benefit of specialisation: often half the budget goes on agents telling each other what they have done.
The system cannot be debugged
When an answer is wrong, you need to find out which participant made the mistake: the planner broke the task down badly, the executor misunderstood it, the reviewer let it through. Without logging every step with its inputs and outputs, debugging turns into reading tea leaves, and improvements into random prompt edits.
Agents agree with each other
The “one does, another checks” pattern works only if the checker has an independent source of truth: a test, a document, a calculation. A checker on the same model with the same information tends to confirm rather than refute — and creates an illusion of control.
Non-determinism multiplies
Each agent adds variability. A system of five agents gives a noticeably less reproducible result than one agent does — and reproducibility is what a business is willing to pay for.
When a multi-agent design is actually needed
- Different access rights: an agent that works with personal data and an agent that writes to an external system should not be the same actor.
- Independent verification: the checker has its own source of truth — a test run, a comparison against a document, a recalculation.
- Real parallelism: ten subtasks of the same kind run at the same time, and that saves time.
- Different models for different steps: a cheap, fast one for routine work, a strong, expensive one for reasoning.
- Context isolation: a subtask needs so much data that it would otherwise push everything else out of the context.
How to decide
- Build a single agent and measure it on a set of real tasks. This takes days and becomes your baseline.
- Find where it fails: not enough context, mixes up roles, loses steps, lacks permission to act.
- Add a split only to address a specific failure — and measure again.
- Compare not only quality but also cost per task and latency. A quality gain of a few per cent at double the cost usually does not pay off.
What to measure in both cases
- the share of tasks carried through to the end without human intervention;
- the share of answers that needed correcting;
- the cost of one task in model calls;
- the time from assignment to result;
- reproducibility: how many times out of ten the same task is solved the same way.
Our experience
We develop and run our own agent orchestration system and work in it every day. The most useful thing this position gives us is the habit of first trying to solve a task with one agent and proving the need for a second, not the other way round. A multi-agent design is a tool against specific constraints, not a sign of a mature system.