Master Joe Phillips
La segunda fuerza laboral11 min read

Buying AI Software Because of the Demo: The Error That Stays

An astonishing demo says almost nothing about the capacity to hold a role. The questions you have to ask the vendor before you buy a piece of AI software.

A mid-sized distributor attends an industry trade show and sees the demonstration of a conversational platform: an assistant that handles customer inquiries with astonishing naturalness. Three months later the platform is bought, integrated into the website and answering questions. Meanwhile, the company's real bottleneck remains intact: quotes take days because the pricing exceptions live in the head of the commercial manager, who resolves them one by one over WhatsApp.

The case comes from my book AI Employee, and the most uncomfortable part is not the error. It is that nobody did anything stupid: each individual step seemed reasonable. The demo was genuinely good. The platform works. The assistant contributes something. The problem was in the order: the solution arrived before the diagnosis, and the company automated the most eye-catching part of its operation while leaving untouched the part that was bleeding.

If you are evaluating whether to buy AI software, this article is the vaccine against that order.

The error that does not explode

It is better not to close the distributor's story with a clean ending, because the real ones do not have one. That company keeps paying the license. Canceling the assistant feels like admitting a three-month error and a budget already spent, so nobody cancels it. Quotes still take days. Most likely, a year from now the assistant will still be there, neither indispensable nor retired, and the bottleneck will still be the same.

Bad decisions rarely explode

The real cost of buying the demo is not a visible disaster. It is a company that spent its transformation budget and its internal credibility on the part of the problem that did not hurt. When somebody proposes the next AI initiative, the answer will be "we already tried that." The money comes back; the internal credibility to transform, much more slowly.

And there is a late signal that confirms the diagnosis: the organization begins to serve the tool it bought so that the tool would serve it. Integrations nobody asked for, parallel processes that duplicate work, people working around the system instead of being freed by it. The book describes it as a demanding guest: it has to be fed with data, justified in meetings, defended from skepticism.

Why a demo fools intelligent people

Chapter 3 of the book opens with a comparison that everyone who has hired staff recognizes. A candidate arrives at the interview perfectly prepared: he knows the company's history, answers with confidence, elegantly solves the case he is given. Weeks later, already inside the role, he needs instructions for every exception, loses context between conversations and does not know when to stop and ask for help. Nothing he did in the interview was false. We simply confused a one-off act with recurring performance.

The interview and the demo share the same asymmetry, and naming it explains why both fool intelligent people: "in the interview the answers shine; in the operation the doubts protect." Whoever builds a demo selects clean examples, provides complete context and shows a brief execution where everything goes well. Real operation lasts every day: it receives incomplete information, runs into exceptions, competes for resources, affects customers and produces consequences.

And if the result of the demo is good, we do the rest of the work: we project onto the system a continuity, a judgment before exceptions and a response under pressure that nobody demonstrated. The fluency of the conversation works as an endorsement of capabilities that were never on the screen.

The question that does decide

Out of that asymmetry comes the change of question that orders the whole evaluation. The wrong question is "did you see what it achieved?" The one that decides is another: can it sustain that result when conditions change, and can we manage it when it does not achieve it?

Notice that the question has two halves and the second one is the less obvious. It does not only matter whether the system sustains performance: it matters what happens the day it does not sustain it. The book is blunt on this point: a resource that produces a convincing answer instead of escalating a doubt is more dangerous than a less brilliant but better governed one, because the error of the first reaches the customer dressed in confidence.

A good occupant of a role is not the one who never runs into an exception. It is the one who knows how to act within limits when the exception arrives.

The questions for the vendor

With that criterion, the sales meeting changes its script. Instead of asking the demo to impress more, interrogate the role the system would have to hold. These questions come directly from the instruments of the book:

On continuity. What happens between executions? Does the system keep identity and history, or does each session start from zero? What does it remember, who can correct what it remembers and how is what must be deleted deleted?

On exceptions. What does it do with a case it does not understand: does it escalate it, improvise it or bury it? To whom does it hand it over and through which channel? Ask to see a demo of failure: what the system looks like when it does not know. The vendor who can only show you successes is showing you the half that matters least.

On authority. What can it approve, modify, commit or spend, and where are those thresholds configured? In numbers or in adjectives? What can it never do, by design, no matter how it is asked?

On observability. Can we reconstruct what it did last week: actions, data consulted, tools invoked, costs? Or do we only have the system's word about itself?

On control. Who can suspend it, in how much time, and what happens to the work in progress when it is suspended? Is there a procedure to give the work back to a person?

Notice what these questions have in common: none of them is answered by looking at the fluency of the conversation. All of them are answered by looking at the governance of the system, which is exactly what the demo does not show. They are, at bottom, a commercial version of the test I develop in the nine properties of an AI Employee.

The correct order: the diagnosis before the purchase

What remains is the deeper defense, the one that keeps you from arriving at the trade show already in a buying state. The book formulates it as an exercise of subtraction: take the initiative you are considering and remove the name of the tool. Describe only the result the organization needs, the frequency with which it must be produced, the consequences of doing it badly and the person who answers for it. Read it without the brand.

If the case loses its meaning when the product disappears, if the only thing left standing was "the thing is, this platform is impressive," you were defending a technology, not solving a problem. The distributor of the case would have discovered in that half hour what cost it three months and a budget: that its problem was not handling inquiries from the website but getting the pricing exceptions out of the head of the commercial manager.

A need that is well understood can admit different solutions. A technology bought before understanding the need usually forces you to deform the problem so that it fits inside it. That is why the selection of the tool must be the last decision, not the first, and that is why the starting point is not a catalog of platforms but the map of your own operation: the invisible workforce first, the design of the role afterward.

The speed of AI did not change the nature of this error, which existed long before conversational demos. It changed its price, because the wrong thing now gets built faster than ever.

Building the wrong thing fast is not progress. It is accelerated waste.

AI Employee

Frequently asked questions

Because of a structural asymmetry: the demo lasts minutes, happens under favorable conditions and shows the path where everything goes well, while real operation lasts every day, receives incomplete information, runs into exceptions and produces consequences. Whoever builds the demo selects clean examples and provides complete context; and if the result is good, the buyer does the rest of the work, projecting onto the system a continuity and a judgment that nobody demonstrated. It is the same trap as the brilliant candidate in the interview who later does not know when to ask for help: nothing was false, but a one-off act was confused with recurring performance.

Questions of governance, not of capability. Does the system keep identity and history between executions, and who governs that memory? What does it do with an exception it does not understand: does it escalate it, improvise it or bury it, and to whom does it hand it over? What can it approve or spend, with thresholds in numbers and not in adjectives, and what can it never do? Can what it did last week be reconstructed, with data, tools and costs? Who can suspend it and how is the work given back to a person? And ask explicitly for the demo of failure: how the system behaves when it does not know. None of these answers is in the fluency of the conversation.

Almost never a visible disaster. In the case of the distributor in the book, the assistant bought at the trade show was left handling inquiries (it contributes something, so nobody cancels it) while the real bottleneck, the quotes that take days, remained intact. The real cost was triple: the transformation budget spent on the part of the problem that did not hurt, the internal credibility burned for the next AI initiative, and a system turned into a demanding guest that has to be fed with data and defended in meetings. Bad decisions rarely explode. Almost always, they stay.

By inverting the order: diagnosis before purchase. First do the exercise of subtraction: describe the initiative without the name of the tool, only with the necessary result, its frequency, the consequences of doing it badly and the person who answers. If the case does not stand without the brand, there is no real problem to solve yet. Second, change the question of the evaluation: not "did you see what it achieved?" but "can it sustain it when conditions change, and can we manage it when it does not?" Third, interrogate the governance of the system (memory, exceptions, authority, observability, suspension) instead of its eloquence. The tool is chosen last; the work is designed first.


Before evaluating any platform, take the inventory of your invisible workforce and learn to distinguish what an AI Employee is from what only looks like one.

Want the full method? Read AI Employee. For executive AI consulting or keynotes and workshops.

Go deeper

Want to bring your team to the next belt?

Book a discovery call or explore the full book.

FAQ

Frequently asked questions

Detailed answer in the article body. See the relevant section.

Detailed answer in the article body. See the relevant section.

Detailed answer in the article body. See the relevant section.

Detailed answer in the article body. See the relevant section.

Keep training