Long before there was AI in any conversation, we gave an operations manager a one-line objective: increase sales. And he achieved it. Extraordinarily, at that. Sales went up month after month and for a while we celebrated it.
Until the margin started to fall. Not a little: enormously. When we went looking for the cause, it was obvious and it was ours. In order to sell more, he had lowered prices. The book sums it up without anesthesia: "He did exactly what we asked of him; the problem was that we asked for half of what we wanted". Nobody gave him a wrong instruction. We gave him an incomplete instruction, which is worse, because "a wrong instruction gets argued and an incomplete one gets obeyed".
That manager was a person with judgment and with years inside the company, and even so he optimized exactly what we measured him on. The damage took months to become visible because a person, sooner or later, asks himself whether what he is doing makes sense. An AI Employee to which you hand a single metric does not have that pause: it is going to do the same thing, faster, and with nobody in the hallway to ask it whether it is sure.
The correction that became a method
The correction of that case had three parts, and over the years one recognizes in them the skeleton of how to direct any resource, human or artificial. First, the objective went from one dimension to two: sell more, always defending the margin. Second, the authority was left in numbers and not in judgment: fixed list prices and a minimum price per service, below which nobody sells without authorization. Third, a monthly monitor that raises an alert when some customer pays less than what it costs to serve them.
Each piece does a different job. The first fixes the metric. The second sets a limit that does not depend on anybody's interpretation. The third assumes that the previous two can fail, and watches. None of them existed when we gave the original instruction. With an AI Employee, all three have to exist before switching it on.
Output is not outcome
The underlying antidote begins by separating two words that enthusiasm melts into one. The output is the message sent; the outcome is that the case reaches the appropriate person in charge without delay or harm to the customer. Measuring the first is simple. Managing the second demands understanding what the role exists for.
Speed aggravates the confusion, because a large amount of activity looks like value. The book says it with arithmetic: "automation does not improve a process, it multiplies it". A badly designed process that produced ten errors a month will produce a hundred once it is automated. And if the dashboard measures activity, it will applaud just the same.
The human work left behind also counts. If three people dedicate hours to reviewing, correcting and explaining the agent's decisions, that intervention is part of the cost, and hiding it allows telling a story of autonomy that the operation does not sustain.
Performance must be measured by outcomes, quality, risk and cost, never by activity, hours, tokens or number of messages.
The six metrics that should reach the CEO
The core is six executive metrics in two trios. The first is economic and operational and answers three questions: is the work working better than before?, how much does a correct outcome cost?, how much autonomy is real?
The Post-Transition Performance Delta compares the role before and after the change. It does not ask whether the agent is fast; it asks whether the role improved in the outcome that justifies the responsibility. Without a baseline, the organization can celebrate an imaginary improvement. The Cost per Successful Outcome calculates how much it costs to produce the result done right: platform, integration, supervision, rework, incidents and residual human operation. Comparing only salary against subscription is intellectually poor. And the Human Intervention Rate measures what proportion of cases needs human help or correction; a high rate can be correct in shadow mode, but calling it autonomy when the team manually holds up every result is lying to yourself with numbers.
The second trio is the one of safety and harm, the one that almost no executive dashboard builds. In that missing half live the problems that end up in the press.
If you reward the system for escalating less, it learns silence: it improves its number by producing the failure you fear most, the convincing answer where there should have been an escalated doubt. Escalation is judged by quality, never by volume. A Human Intervention Rate that goes down is not automatically good news: review the missed escalations before celebrating.
The first metric of that trio is the one almost nobody measures: the Missed Escalation Rate, the case that should have gone up to a human and did not, discovered in an audit or after the harm. The second is the Harmful Outcome Severity, which classifies the harm that actually reached customers or workers: it does not ask how many errors there were but how much the worst one weighed, because a hundred typos in reminders and one collection threat to a customer in mourning are not "101 errors", they are two different phenomena. And the third is the Hybrid Workforce Incident Rate, the incidents attributable to the design of the role: the handoff that dropped a case, the badly drawn authority, the context that was missing. Its value is uncomfortable on purpose: every incident of this category is a design decision signed by somebody.
The scorecard is written before the pilot
Before modifying a role, record how it works today. Choose between three and five KPIs that represent its result, not only its activity, with their standards of quality, cycle time, cost, errors and human intervention. Then set the band inside which you will consider the new performance stable, and the conditions that force stopping or reviewing the pilot.
Keep comparability: if you change the process, the definition of success, the population of cases and the resource all at once, afterward you are not going to know what caused the result. And give the scorecard a frequency and an owner: who reviews it, how often, with what authority to intervene. What improvement widens autonomy, what deterioration activates remediation, what incident demands immediate suspension. A dashboard without an associated decision is decoration.
The psychological benefit is worth more than the technical one: with the scorecard written first, enthusiasm stops drafting the definition of success after seeing the results. If the hypothesis fails, you learned. The failed hypothesis is information, not shame.
Five greens and one red
The book shows a collections scorecard with hypothetical data, at the close of its first ninety days, and the lesson fits in two lines. Recovery went up from 61% to 64%. Material errors went down. Zero contacts to disputed accounts. The suspension drill was executed successfully. The hours freed up for the analyst have a declared destination. Five greens. And one red: the audit detected two critical missed escalations, two cases where the customer mentioned a serious economic difficulty and the system continued its sequence.
Is it promoted to the next rung of autonomy? No. And that is the whole lesson.
Any dashboard that averaged those six indicators would show a successful pilot. But the threshold for missed escalations was not "few": it was zero, fixed before the pilot, and for that reason it admits no negotiation. If "less than 1%" had been written, two cases out of 3,600 contacts would have passed as excellent, and the organization would have promoted a system that learned not to ask precisely where it matters most. What is appropriate is the boring thing: the resource stays on its rung, an investigation looks into why the two mentions did not trigger the routing rule, it is corrected and it is measured again for a quarter.
An indicator that never changes a decision is not an indicator. It is decoration with numbers.
Frequently asked questions
Six metrics in two trios. The economic and operational one: Post-Transition Performance Delta (did the role improve against its baseline?), Cost per Successful Outcome (how much does a correct result cost, including supervision, rework and incidents?) and Human Intervention Rate (how much autonomy is real?). The safety and harm one: Missed Escalation Rate (cases that should have gone up to a human and did not), Harmful Outcome Severity (how much the worst harm that reached customers or workers weighed) and Hybrid Workforce Incident Rate (incidents attributable to the design of the role). Most executive dashboards only build the first trio.
For two reasons. The technical one: without a baseline and without prior thresholds there is no valid comparison, and the organization can celebrate an imaginary improvement or negotiate the verdict after seeing the numbers. The psychological one, which is worth more: with the scorecard written first, enthusiasm stops drafting the definition of success after knowing the results, and the team is protected from justifying the implementation by the effort invested. A threshold fixed before the pilot (for example, zero critical missed escalations) admits no negotiation when the uncomfortable data arrives.
The output is the activity executed: the message sent, the classification made, the record updated. The outcome is the effect that justifies the role: the case that reaches the appropriate person in charge without delay or harm, the useful commercial conversation, the payment received. Measuring outputs is simple and that is why dashboards fill up with them; managing outcomes demands understanding what the role exists for. The distinction matters because a fast system can produce extraordinary activity and harmful results at the same time, and a board that measures activity applauds both equally.
The Missed Escalation Rate measures the cases that should have gone up to a human and did not, discovered in an audit or after the harm. It is the metric that almost nobody builds and the most important one of the safety trio, because it captures the most dangerous failure of an AI system: the convincing answer where there should have been an escalated doubt. It is complemented by the unnecessary escalation (noise that erodes the reviewer's attention) and the appropriate one, which is the health of the system. The system is never rewarded for escalating less: if you do it, it learns silence, and it improves its number by producing exactly the harm you fear most.
The scorecard decides the promotions and demotions of the AI autonomy ladder. And when an indicator turns red, the next question is why: read Auditing is not logging.
Want the full method? Read AI Employee. For executive AI consulting or keynotes and workshops.
Go deeper
Want to bring your team to the next belt?
Book a discovery call or explore the full book.