The problem of the unbounded agent
An agent without an account is limited only by the patience of whoever pays the provider’s bills. It can go off into a long chain of reasoning, call a tool a hundred times in a row, spawn copies of itself. In a demo this looks like persistence; in operation, like uncontrolled spending.
Instructions do not work here: asking it to “be economical” does not create a constraint, it creates a style. A constraint appears when the resource runs out.
Which resources to count
In our environment a participant holds several kinds of resource, and they are not a single abstract “energy” but things that each run out in their own way:
- credits — the universal means of settlement between participants;
- compute time;
- model tokens — what thinking is paid for with;
- storage;
- bandwidth.
The separation matters because the shortages differ. A participant can be rich in credits and poor in tokens — and then it pays to buy a ready result from a neighbour rather than think for itself. It is exactly from asymmetries like this that a division of labour arises which nobody assigned.
Thinking as a paid operation
A call to a language model is built into the participant’s cycle as an ordinary operation on its account, not as an external call made “somewhere off to the side”. The order is as follows:
- the participant’s state and goal are serialised into a request;
- tokens and, if needed, credits are reserved — before the call, not after;
- the request goes through a provider-neutral interface, so the model vendor can be replaced;
- actual consumption is debited as incurred, and the unused reservation is returned;
- validated changes are applied to the participant’s state, and the requested actions go to a handler.
Why both sides of a trade reserve
When participants exchange resources, the naive scheme of “agree now, settle later” breaks at the first failure: the seller has already promised the same thing to two parties, the buyer has spent credits it no longer has.
So both sides reserve at once: the seller’s resources and the buyer’s credits go into escrow when the order is placed. From there, simple rules apply:
- the trade executes atomically — the resource and the payment change hands at the same time;
- a price improvement goes back to the buyer rather than staying with the seller;
- cancellation and expiry return the unused escrow;
- balances, orders, trades and the journal survive a restart.
The main consequence is that the same resource cannot be offered twice. That is what separates a working economy from a table of numbers: a promise is backed before anyone has even heard it.
What changes in the system’s behaviour
- Useless activity stops by itself: it does not pay off, and the resource is finite.
- A natural priority appears: a participant spends where it expects the greater return.
- Exchange arises: it is cheaper to buy a result than to produce it — or the other way round, depending on the shortage.
- Costs become visible not at the end of the month on the provider’s bill, but at the level of each participant and each task.
- Limiting autonomy stops being a question of trust: even a fully autonomous participant will not go beyond its account.
How this carries over to ordinary products
Not every project needs the full economy, but two of its elements pay off almost always:
- a per-task budget with a hard ceiling — an agent that has used it up stops and hands the work to a person instead of carrying on for “just a little longer”;
- cost accounting in business units — the cost of a completed task, not the cost of a thousand tokens. Only that figure can be compared with the value of the result.
A practical observation: as soon as a team starts seeing the cost of one solved task, architectural arguments about which model to choose are over within a day. Numbers settle things faster than arguments do.