A definition that is actually useful
An AI agent is a system that receives a goal, not a command. It decides for itself which steps to take, calls external tools, looks at the result and adjusts its behaviour until the goal is reached or it runs into the limit of its authority.
A formula that makes this easy to check: goal → plan → action → observation → memory → adjustment. If a system lacks at least observation and adjustment, it is not an agent but a text generator in a wrapper.
How an agent differs from neighbouring concepts
From a chatbot
A chatbot answers. An agent acts: it files a ticket, finds a document, recalculates, sends the result to your system. An agent may have no chat window at all — it can work on an event rather than on a message.
From classic automation
An automation script runs a sequence described in advance. An agent builds the sequence itself for the case at hand. That is its strength on varied inputs — and its weakness where exactly the same order of actions is required every time.
From “a model with a prompt”
A prompt sets behaviour within one exchange. An agent lives longer than one exchange: it remembers the context of the task, sees the results of its actions and is responsible for the final result, not for a single reply.
What it is made of
- Model — decides on the next step. The same agent architecture can run on different models.
- Tools — access to databases, APIs, search, code. They are what turn reasoning into action.
- Memory — the context of the task, the history of steps, long-term knowledge about the domain and the user.
- Planner — breaks the goal into steps; in simple agents the model itself plays this role.
- Control loop — limits of authority, points where a human confirms, an action log.
- Evaluation — a set of test tasks and metrics, without which you cannot tell whether things got better.
Which tasks to give it
A good candidate for an agent meets four conditions at once:
- The task repeats — otherwise preparing the tools and checks will not pay off.
- The result is verifiable — there is a way to confirm it was done right without relying on an impression.
- The inputs are varied — otherwise an ordinary automation script is cheaper.
- The cost of an error is known — it determines how much control to build in.
Bad candidates: unique one-off tasks, processes with no criterion of correctness, operations with irreversible consequences and no way for a human to confirm them.
Autonomy is a scale, not a switch
A sensible rollout goes up in steps, and each next step is switched on only after the metrics have confirmed the previous one:
- The agent proposes — a human executes. Useful already at this step: you see the quality with no risk.
- The agent executes — a human confirms every action.
- The agent executes on its own; a human confirms only risky operations.
- The agent works autonomously; a human handles exceptions and reviews reports.
Jumping straight to the fourth step is the most common reason pilots get rolled back: without accumulated statistics, nobody can tell what exactly the system is doing wrong.
What to measure
- the share of tasks carried through to the end without human intervention;
- the share of results that needed correcting;
- the cost of one solved task, not of one model call;
- the time from assignment to result, compared with the manual process;
- the share of correct refusals — when the agent rightly decided that the task was not its job.
A common implementation mistake
Starting by choosing a framework. In practice it is more useful to start by describing one scenario down to the level of “what counts as a successful result”, and by checking that the agent has access to the data and systems it needs at all. A framework can be changed in a week, while a wrongly chosen scenario wastes a quarter.