Master Joe Phillips
Gobierno y autonomía10 min read

Shadow Mode: Testing the AI Employee Before Giving It the Operation

Shadow mode AI: real cases without touching the operation. What evidence it gives, how long it lasts, and why skipping it is the costliest way to know it.

The demo went flawlessly. The team wants to connect the agent on Monday, the vendor offers to accompany the launch, and the business case promises to recover the investment in one quarter. Everything pushes in the same direction: turn it on.

Stop for a moment. There are two ways to really get to know your AI Employee. One is to observe it for weeks processing your real cases with no authority to touch anything, and to study every difference between what it decided and what your people decided. The other is to discover its limits in production, with real clients receiving its errors. Both teach exactly the same thing. The second one charges you for it in money, clients, and internal credibility.

The book gives the first one a name: shadow mode. It is the medical residency of the artificial resource, and chapter 6 defines it in one line: "it executes the real work, with real cases, without its actions touching the operation yet".

First the baseline, then the comparison

The sequence matters more than the tool. Before turning anything on, the human baseline is established: quality, time, cost, errors, and exceptions of the process as it works today. This step is non-negotiable and it usually gets skipped out of anxiety. The book is blunt: "without a baseline, any later comparison is an opinion with charts".

With the baseline in hand, the resource processes the same cases in parallel. Where your analyst decided A, what would the system have decided? Every match adds evidence. Every difference opens an investigation: did the system see something the human did not see, or was it missing context that the human had in her head? Out of that comparison comes something no demo can teach you: where the design, the context, or the model need adjustments before their errors cost money or clients.

That is what shadow mode produces: compared evidence, case by case, about your operation and not about the vendor's benchmark. A record of which types of case the system matches the human on, which ones it beats the human on, and which ones it still cannot be trusted with on its own.

Least privilege or theater

Throughout the whole journey an elementary security rule applies: the resource receives only the access and the power necessary for the level it is being tested at. The complete progression can begin with simulation, continue with shadow mode, move to human approval case by case, and reach controlled delegation. At each stage, the permissions go with the stage, not with the ambition of the project.

Check the permissions before boasting about prudence

A resource in shadow mode does not need write permissions in the billing system. If it has them, your shadow mode is theater with real risk: you boast about observing while you left the door to production open. Least privilege is not paranoia; it is the difference between observing and exposing yourself.

And if there is nobody to compare against?

A question appears as soon as you try to apply the method: what if the responsibility is new and there is no human performance to compare against? It happens more than it seems. A service that was never offered, an hourly coverage nobody covered, an analysis that was not done because there were no hands.

The temptation is to skip the step. It does not get skipped: it gets built forward. A prospective baseline is put together with samples that an expert resolves by hand and that stay as a reference, with a short manual pilot, with double human review over the first cases, with data from an analogous process that does exist. And above all with acceptance criteria written before turning anything on: what counts as a correct answer, what counts as a serious error, which case must always escalate.

The difference with the historical baseline is one of nature, not of quality. That one compares against what used to happen; this one compares against what you decided should happen. And there is a rule worth engraving because the error goes in the direction contrary to intuition: "the absence of a baseline is a reason to demand more prudence, never to lower the quality standard". When you do not know what "good" looked like before, your uncertainty is greater, not smaller. That calls for a lower rung of autonomy and a longer observation window, not a looser bar.

How long does it last?

There is no universal number, and distrust whoever sells you one. Validation time depends on three variables: the risk of the role, the frequency of the cases, and their variety. A daily low-impact process can gather enough evidence in weeks, because the volume accumulates representative cases fast. An infrequent, high-cost decision can require months. And some responsibilities, however well the assistance works, should never leave human hands.

What does exist is a universal temptation: declaring success after a few favorable examples. Ten cases resolved well feel like proof and they are not, just as one good week does not turn a new salesperson into a key account manager. The book finishes it off in one sentence: "Resisting it is the difference between validating and celebrating".

The exit signal is not a date on the calendar. It is evidence accumulated against the entry conditions of the next level of autonomy: sustained minimum quality, enough representative cases, an exception rate within what was agreed, zero incidents of the kind that block promotion. All of that written before starting, so that enthusiasm does not draft the verdict.

Skipping it is the most expensive way to know it

Shadow mode does not eliminate risk. It makes it observable before transferring custody of the work, which is different and is enough. It lets the organization learn without turning every hypothesis into an irreversible bet.

Think about it from the cost of learning. Everything shadow mode teaches you cheaply, production teaches you expensively: the policy the system misinterprets, the type of client that confuses it, the exception nobody documented because it lived in the analyst's head. In shadow mode, each one of those findings is a row in a table of differences. In production, each one is a wrong email, an account touched improperly, or an apology somebody has to sign.

And there is a less visible cost: the internal credibility of the project. The team that sees incidents in the first weeks stops trusting the system, and that trust costs twice as much to recover. Shadow mode also protects that: when the resource finally touches the operation, nobody is guessing anymore how good it is. There is a file.

Autonomy is not a property of the model. It is a decision of the manager, based on evidence and conditioned by the role.

AI Employee

Frequently asked questions

It is the stage in which an AI system executes the real work, with real cases, without its actions touching the operation yet: it processes in parallel the same cases as the human team and its decisions are compared against the human decisions, without reaching clients or modifying systems. It works as the medical residency of the AI Employee: it produces evidence about where the design, the context, or the model need adjustments before their errors cost money or clients. It does not eliminate risk, but it makes it observable before transferring custody of the work.

There is no universal number. The duration depends on the risk of the role, the frequency of the cases, and their variety. A daily low-impact process can gather evidence in weeks; an infrequent, high-cost decision can require months, and some responsibilities should never leave human hands. The exit is not marked by a date but by the evidence: sustained minimum quality over representative cases, an exception rate within what was agreed, and zero blocking incidents, all defined in writing before starting. The temptation to declare success with a few favorable examples is the one that has to be resisted.

A prospective baseline gets built, the step never gets skipped. It is put together with samples that an expert resolves by hand and that stay as a reference, a short manual pilot, double human review over the first cases, data from an analogous process that does exist and, above all, acceptance criteria written before turning anything on: what counts as a correct answer, what is a serious error, which case always escalates. Unlike the historical baseline, which compares against what used to happen, this one compares against what you decided should happen. And the absence of a baseline demands more prudence: a lower rung of autonomy and a longer observation window.

The minimum: read-only on the data necessary to process the cases it is observing. The principle of least privilege applies across the whole progression (simulation, shadow mode, approval case by case, controlled delegation): the resource receives only the access and the power necessary for the level it is being tested at. A system in shadow mode does not need write permissions in billing, CRM, or email; if it has them, the shadow mode is theater with real risk, because a failure or an attack could touch production while the organization believes it is only observing. Permissions grow with the evidence, level by level.


Shadow mode is the first rung of a complete progression: read The AI autonomy ladder. And the evidence it produces is judged against metrics defined before the pilot: read The AI Employee scorecard.

Want the full method? Read AI Employee. For executive AI consulting or keynotes and workshops.

Go deeper

Want to bring your team to the next belt?

Book a discovery call or explore the full book.

FAQ

Frequently asked questions

Detailed answer in the article body. See the relevant section.

Detailed answer in the article body. See the relevant section.

Detailed answer in the article body. See the relevant section.

Detailed answer in the article body. See the relevant section.

Keep training