"Our agent can handle your collections end to end." The salesperson is not lying. The demo backs him up, the proposal repeats it, the price makes it tempting. But that sentence describes capability, not authority. And confusing those two words is the fastest way to turn a good AI project into an incident with a name of its own.
Medicine solved this problem a century ago. A newly graduated resident may know more pharmacology than the head of surgery, and even so nobody hands him an operating room on the first day. First he observes, then he assists, then he executes under supervision, then he operates alone within limits. Supervision is withdrawn with evidence, not with sympathy. No hospital calls that distrust. It calls it responsibility.
Your AI Employee deserves the same treatment. Not because it is weak, but because authority over your accounts, your customers and your brand does not come in the box. You grant it. And there is a disciplined way of granting it: a ladder.
Capability is not authority
The market erases this distinction every day, so it is worth engraving it: the vendor declares capability; the organization grants authority. When the salesperson says that his agent "can" handle your collections, he makes a technical claim, and probably a true one. The authority to touch your accounts and write to your customers does not come with the product: you grant it, in stages, against evidence, with a designed withdrawal.
A model can answer perfectly in the test and behave differently when the data changes, the volume rises or the exception appears that its training did not contemplate. A good answer in the laboratory does not grant authority over production, just as passing the licensing exam does not grant patients. The standard that accompanies the book fixes it in a single line, clause HWF-31: "Authority must be explicit, limited and revocable." And it is written before the resource operates, never inferred afterward.
This also changes the underlying question. "Can I trust AI?" forces a choice between faith and rejection, and neither of the two is manageable. The correct question is another one: what level of autonomy has this resource demonstrated it can handle, within the risk of this role? With that turn, trust stops being a binary emotion and becomes an administrative decision: gradual, conditional, reversible and based on evidence.
The five rungs
The ladder has five levels, and it is worth knowing them as much for what the resource can do as for what it still cannot touch.
Level 1: observe. The resource records and compares without intervening. It is shadow mode institutionalized: it produces evidence, not effects. Nothing it does reaches a customer or modifies a system.
Level 2: recommend. It proposes the action and shows the evidence it used. A human executes. The value is in the verb to show: the book warns that "a recommendation without its evidence is an order disguised as a suggestion". If you cannot see why it recommends, you are not supervising, you are obeying.
Level 3: execute with approval. It acts only after explicit human approval, case by case. Here the organization learns the pattern of the exceptions while keeping complete control.
Level 4: act by exception. It executes within rules and the manager intervenes only when something leaves the defined path. This is the point where the economics of automation start to materialize, and where the quality of escalation becomes the critical variable.
Level 5: autonomous within limits. It operates alone, keeps traceability and escalates what it cannot resolve. The limits never disappear. They are managed.
What each level demands in writing
A level without a written definition is a label, not a control. For each rung define three things.
First, entry conditions: what minimum quality, how many representative cases, what rate of exceptions is acceptable, what incidents block the promotion. Second, the corresponding authority: what data it consults, what systems it modifies, what amounts and what communications it is allowed. Autonomy without an authority matrix is an empty word. Third, rollback triggers: if quality drops, a policy changes, human intervention increases or a new risk appears, the resource returns to the previous level.
That last rule prevents a promotion from turning into an acquired right: authority stays conditioned on performance and on context. And there is a brutal test: what event sends the resource back a level, and who has the power to send it back? If you cannot answer the second question with a name, the ladder does not exist yet.
Two ceilings that are not negotiable
No responsibility starts above level 3, whatever the vendor's assessment says. And the risk class puts its own ceiling: in a Critical responsibility, every action ends in a final human decision: the high rungs are not available, however well the resource behaves.
The first ceiling protects against the initial enthusiasm: the demo does not replace your operation, with your data and your exceptions. The second protects against accumulated enthusiasm: there are roles where the cost of the error is so severe or so irreversible that the question is never whether the system is good, but who signs the final decision. And the risk class is not decided by taste: it comes out of a formal assessment of the role, the same one that grades how much approval an action demands and how much depth an audit has.
The opposite error also has a cost
The ladder protects in both directions. The visible error is granting too much, too soon. The less visible error, and almost equally expensive, is keeping the AI Employee asking for authorization on every detail for months. If a person has to review everything, autonomy does not exist and the human cost stays hidden: you pay for the platform and you also pay the full-time supervisor that the operation does not report.
The objective is not to accumulate approvals: it is to withdraw supervision where the evidence allows doing it without losing control. Neither faith nor paralysis: evidence.
Reversible authority: control changes shape, it does not disappear
In human management there are promotions, improvement plans and suspensions. An artificial role needs the same lifecycle movements: it can gain scope, lose permissions, return to shadow mode or be suspended. Authority is not a definitive ceremony; it is a reviewable operational state.
This idea frees the manager from two fears. The first is believing that granting autonomy means losing control forever. It does not: it means changing the shape of the control, from prior approval to supervision by exception with a withdrawal available. The second is feeling that withdrawing authority proves that the implementation failed. It does not either: the organization learns, the risks change, and adjusting the role is a normal function of management.
Going down a rung is an administrative action, not a confession.
The kill switch is rehearsed, not declared
Reversibility is designed before the incident, and the correct register is that of aviation, the industry that turned the emergency into a procedure. No pilot executes his first engine shutdown on a burning engine: he practiced it dozens of times in the simulator.
The kill switch of an AI Employee deserves the same discipline. Who can reduce permissions. What work goes back to humans and to whom. How the context of the cases in flight is preserved. What evidence is reviewed afterward. And drills with a calendar, because a stopping mechanism that has never been exercised is a hypothesis, not a control. The standard says it in institutional register: "a suspension mechanism that has not been exercised does not constitute an effective operational control". The cadence is declared, never longer than twelve months, and shorter for High and Critical risk roles.
In the dojo it works the same way: the technique you never practiced does not exist on the day you need it. The first real pull cannot be the first pull.
Frequently asked questions
It is the level of operational authority that an organization grants to an AI system over a concrete role, after observing its performance with evidence. It is not a characteristic of the product: the vendor declares capability, but the authority to touch accounts, customers or systems is granted by the organization, in stages and with a designed withdrawal. The useful question is never "can I trust AI?", but "what level of autonomy has this resource demonstrated it can handle, within the risk of this role?".
Level 1, observe: it records and compares without intervening, it produces evidence and not effects. Level 2, recommend: it proposes the action and shows the evidence it used; a human executes. Level 3, execute with approval: it acts only after explicit human approval, case by case. Level 4, act by exception: it executes within rules and the manager intervenes when something leaves the defined path. Level 5, autonomous within limits: it operates alone, keeps traceability and escalates what it cannot resolve. Each level demands entry conditions, authority in writing and rollback triggers to the previous rung.
No. No responsibility starts above level 3 (execute with explicit human approval, case by case), whatever the vendor's assessment says. Autonomy is earned in production, not in the sales proposal. Besides, the risk class of the role puts its own ceiling: in a Critical responsibility, every action ends in a final human decision, so the high levels are not available however well the system behaves. Before any rung with real effects, the resource goes through shadow mode.
It is the mechanism that allows reducing permissions, suspending the system and returning its work to humans in an orderly way: who can activate it, what work goes back and to whom, how the context of the cases in progress is preserved and what evidence is reviewed afterward. It must be rehearsed with scheduled drills because a stopping mechanism that has never been exercised is a hypothesis, not a control. The standard of the book demands exercising it with a declared cadence, never longer than twelve months, and shorter for High and Critical risk roles.
The first rung of the ladder has a method of its own: read Shadow mode: testing the AI Employee before handing it the operation. And no level promotion is decided with impressions: it is decided with the AI Employee scorecard, written before the pilot.
Want the full method? Read AI Employee. For executive AI consulting or keynotes and workshops.
Go deeper
Want to bring your team to the next belt?
Book a discovery call or explore the full book.