Metaintelligence ArchitectureMetaintelligence · Consensus · Fault tolerance

Consensus without a centre: deciding together without taking anyone’s word for it

As long as a system has a master node, the question of a shared decision does not arise: the master decides, the rest carry it out. As soon as there is no master, or it cannot be trusted unconditionally, a problem appears that ordinary agent set-ups simply do not deal with.

Metaintelligence Architecture · part 5 of 6

Why “let the coordinator decide” is not a solution

A coordinator is convenient while the system lives in one process on one machine. As soon as the participants are distributed, the coordinator acquires three properties that are expensive in production: it is a throughput bottleneck, a single point of failure and the only one everyone is obliged to trust.

The third property is usually underestimated. “Trust” here is not about malice: a software bug, stale state or a partial loss of connectivity is enough for the coordinator to start handing out decisions that contradict what the others see. And the system will accept them — because trust is built into the architecture.

What the Byzantine fault model means

The usual fault model assumes that a node either works correctly or does not work at all. The Byzantine model allows for worse: the node is running but behaves arbitrarily — it sends different versions to different participants, replays old messages, signs with someone else’s name, votes twice.

For a network of autonomous agents this is not paranoia but an honest statement of the problem. A participant may be on someone else’s machine, run a different version of the code, suffer from clock drift or be compromised. The protocol has to stay correct under these conditions, not under ideal ones.

How one consensus cycle works

  1. The round’s leader sends out a signed proposal: which change to the shared state it proposes to commit.
  2. Each replica checks the proposal on its own — the validity of the state transition, the sequence number, the signature — and replies with its own signed vote.
  3. The leader collects a quorum of unique votes and sends out a commit certificate.
  4. Each replica checks every signature in the certificate, writes it durably to its log, and only then applies the change.

The arithmetic of fault tolerance

For a committee of n participants, the tolerable number of arbitrarily behaving replicas and the quorum size are calculated as follows:

f = floor((n - 1) / 3)
quorum = 2f + 1

In practice: a committee of four survives one failed or malicious participant, provided the other three can communicate. The formula also carries an unpleasant truth — tolerance is paid for in participants: to survive two, you need seven. Consensus is expensive, and that is an argument for applying it selectively.

What the protocol must reject

All of this is checked before a message reaches the consensus logic: the transport verifies the sender, the recipient, the timestamp and the nonce. The separation of responsibilities is fundamental here — consensus code must not deal with network security, otherwise a bug in one place destroys both properties.

Leader change and catch-up synchronisation

The leader may disappear or fall silent. The participants then vote to change the round, and leadership passes on. This is not an emergency procedure but a routine part of the protocol: the system has to keep working without human intervention.

A replica that was unavailable catches up from certificates: it requests every committed decision after its last number and applies them, checking the signatures. It does not need to trust whoever sent the history — the certificates are self-sufficient.

Where an agent system needs this

Not everywhere. Consensus is justified where a decision is shared and irreversible:

Everything else is the participants’ local decisions and pairwise agreements, which do not need the consent of the whole group. Trying to push every action through a committee turns a living system into a slow one: the price of consensus is paid for every decision, while the benefit appears only where divergence is genuinely dangerous.

Code and artefacts

The “Metaintelligence Architecture” series

  1. Metaintelligence: why the next level is not the model but the environment
  2. The agent as a bounded local world
  3. A link as a first-class object, not a message queue
  4. Living graph: when topology is a consequence of work, not a design
  5. Consensus without a centre: deciding together without taking anyone’s word for it
  6. Agent economics: why autonomy needs a budget

Frequently asked questions

Is this a blockchain?

No. This is a fixed, known committee and a classical Byzantine agreement protocol, not an open network with economic incentives. The only thing they have in common is the requirement to stay correct when some of the participants behave arbitrarily.

Why do agents need this complexity?

It is needed exactly when the participants are distributed and the decision is irreversible. For one process on one machine, an ordinary sequential write is enough. We build consensus into the environment because the environment is designed to work across machines, not because it looks elegant.

How many nodes are needed?

At least four, to survive one participant behaving arbitrarily. The formula is strict: each additional fault to be survived requires three more nodes, and this has to be taken into account in design — fault tolerance is not free.

How we help with this

Read next

Let us talk about your task

If your task is similar, tell us what needs solving. We will say so plainly if it can be solved more simply than it looks.

Write to us