Skip to main content

Distrust is a skill: why demos convince and evals do not lie

The short answer

A demo answers "is it possible?" and nothing more

It shows the vendor's best case on the vendor's input in the vendor's environment, which proves the capability can happen in at least one instance and tells you nothing about how often it happens on your inputs. The persuasion is structural (curated input, a success without a denominator, controlled conditions, a vivid concrete story), so a good demo should raise your urge to evaluate, not lower it.

What you will be able to do

  • Explain why a demonstration is structurally persuasive and structurally weak as evidence: it shows a curated best case chosen by the presenter, on inputs the presenter selected, under conditions the presenter controls.
  • Distinguish what a demo can prove (a capability exists in at least one case) from what only an evaluation can prove (how reliably the capability holds across inputs you did not choose).
  • Decompose any AI demonstration into its separate claims, and for each claim identify the evidence that would confirm it and the evidence that would break it.
  • Detect the leading signals of "AI washing" (dressing up human labor or a simple rule as artificial intelligence), including a hidden automation rate, an inability to run the system on your own inputs, and human-shaped latency.
  • Analyze the difference between an anecdote (a numerator with no denominator) and a rate (a measured outcome over a real sample), and demand the denominator every time.
  • Translate a demo claim into a runnable evaluation: a metric, a representative and adversarial sample, a pass line agreed in advance, and a repeatable procedure.
  • Produce a claims-and-evidence audit for one real system: a list of every claim made about what it can do, each paired with the evidence that exists and the test that would settle it.

The lesson

Between 2018 and 2023, a shopping app called Nate told the market it had solved online checkout using artificial intelligence. The promise was perfectly frictionless. Tap once and an AI agent would navigate any retailer's website, fill in your details, and complete the purchase.

On the strength of a clean, one-tap demonstration, the company raised more than $40 million. That capital did not come from amateurs. It came from serious venture funds managed by sophisticated investors whose entire profession is evaluating technology.

They watched the demo work, and they bought the core claim, a fully autonomous AI agent navigating the web with zero human intervention. In April 2025, federal prosecutors stepped in, charging the company's founder with securities fraud and wire fraud. The Securities and Exchange Commission immediately filed a parallel civil action.

The regulatory fallout was total, leaving the investors with near-complete losses. When investigators pulled apart the mechanics of the app, they found the real engine. The orders were not being placed by a machine.

They were being manually completed by hundreds of human contractors sitting in call centers in the Philippines and Romania, typing fast enough to mimic an algorithm. The Justice Department's core allegation states that the application's actual automation rate was effectively 0%. Every defrauded investor trusted a flawless demonstration.

None of them, at the moment they wired the capital, possessed an evaluation. That gap, the space between watching a demo succeed and measuring how a system actually performs, is the exact place where enterprise AI deployments fail. A demonstration is a highly engineered performance.

It is designed specifically to persuade the relying on concrete visual success. It is not designed to prove reliability. A demo answers one extremely limited question, is this capability possible in at least one specific case? Before you deploy a system to make automated decisions, you have to answer a completely different question, how reliably does this capability hold across inputs I did not choose? Approaching this gap with blanket cynicism is useless.

Cynicism is a mood that rejects everything, which paralyzes the business and blocks innovation. The professional alternative is structured skepticism. This is a trained analytic skill, a repeatable procedure where you isolate specific, professional trust cannot run on a feeling installed by a pitch.

It must be an inspectable artifact built through rigorous methodology. A demo persuades through three flaws. First, inputs are curated in advance, you watch the happy path.

Second, it's a numerator with no denominator, one success, but out of how many attempts? Third, survivorship bias hides failed rehearsals quietly excluded before you entered. A true evaluation earns trust through structure, requiring four components. First, a measured denominator, a statistically significant sample size establishing a genuine rate.

Second, the buyer chooses inputs, deliberately injecting adversarial edge cases that represent your actual environment. Third, the procedure must be repeatable by a skeptic. Finally, a pass line is registered before the result is seen, locking the standard in place so it cannot be lowered.

A demo is structurally engineered to hide failure. An evaluation is structurally engineered to find it. They are answering entirely different questions.

The failure to evaluate leaves organizations vulnerable to AI washing. This is the practice of dressing up human labor or simple rules as autonomous artificial intelligence in order to command a premium craze. The deception hides in the gap between automation and autonomy.

Automation executes a predetermined script, while autonomy requires dynamic decisions. A demo easily simulates autonomy using highly scripted automation. To expose this gap, you need the most diagnostic single number in any AI software purchase, the automation rate.

This pie chart shows what it actually measures, the exact fraction of outcomes completed with zero human intervention. The first red flag of AI washing is evasion. If a vendor claims autonomy but refuses to provide a numerical automation rate, they are refusing to measure the capability they are selling.

The second red flag is human-shaped latency. If response times perfectly track with the daytime working hours of a specific global time zone, the intelligence is likely sitting at a desk. The third red flag is input control.

If the vendor refuses to let you drive the system on your unfiltered data, the product likely cannot survive contact with reality. Supervised AI, a system built to keep humans in the loop, is a highly legitimate and often superior product, provided it is disclosed accurately. Humans in the loop are not the crime.

The fraud lies strictly in concealing them while charging the buyer for autonomy. Even when organizations demand proof, they frequently accept false proxies. The most common is the public benchmark, which creates a dangerous illusion of objective reliability.

A benchmark chosen or especially tuned by the seller carries a heavy conflict of interest. It is essentially a demo with a decimal point. This translates directly into legal risk.

A U.S. state attorney general recently opened an investigation into a healthcare AI vendor's advertised hallucination rate. The regulator didn't care that the demo functioned perfectly. They demanded to see the mathematical denominator substantiating the public marketing claims.

That same risk replicates internally through the successful pilot. Enthusiastic internal teams often launch pilot programs to validate a new AI tool. To win executive approval, the team quietly curates a perfectly clean subset of data.

They filter out the messy operational edge cases, run the tool, and report a massive success. That curated pilot is just an internal demo in disguise. By avoiding the unhappy paths, it shares every structural flaw of a vendor pitch.

Whether a leaderboard or an internal test, any number is an illusion if the evaluator didn't control adversarial inputs and mandate a strict denominator. To operationalize this discipline, we need a method to translate a highly persuasive 30-minute pitch into a defensible governance document. We use a single tool, the four-column claims and evidence audit table.

Column 1 translates vague marketing verbs into precise, testable behavioral claims. Column 2 ruthlessly lists the existing evidence supporting that claim. If the only evidence on hand is a demo, you must label the claim strictly as possible, not yet measured.

Column 3 defines the exact evaluation that would settle the claim. You dictate the denominator, mandate your own adversarial data, and set the passline. Column 4 documents the disclosed automation rate and explicitly records any observed AI washing red flags.

When you complete this exercise, your audit will likely be filled with not-yet-measured rows. That is not a failure. It is an honest map of exactly where your trust currently relies on hope.

Building this table is the mechanism that forces a vendor out of the demo theater. It replaces curated illusions with measurable operational reality. The final operational challenge is maintaining velocity.

Demanding a massive, thousand-case evaluation for every minor AI tool will paralyze your business. To govern efficiently, apply the rule of proportion. You must match the depth of your proof to the specific operational stakes of the deployment.

This matrix plots operational stakes across two axes, the consequence of a system failure and the reversibility of automated decision once made. A low-consequence, highly reversible tool, like drafting an email for human review, carries minimal risk. A light check on a small sample is sufficient proof.

High-stakes, irreversible decisions, like denying a financial benefit, demand exhaustive, adversarial evaluation on your proprietary data. A governed trust decision scales with the risk, ensuring you can completely defend your methodology to a skeptical board of directors or an investigating regulator. Stop buying demos.

Demand the denominator. Ensure your trust is an artifact you can mathematically prove. To operationalize this discipline, we need a method to translate a highly persuasive, 30-minute pitch into a defensible governance document.

We use a single, actionable tool, the four-column claims and evidence audit table. Column one strips away the marketing language. You translate vague verbs, like intelligent, into precise, testable behavioral claims.

Column two ruthlessly lists the evidence that exists today to support that specific claim. If the only evidence on hand is a demo, you must label the claim strictly as possible, not yet measured. Column three defines the exact evaluation that would settle the claim.

You dictate the denominator, mandate your own adversarial data, and set the pass line. Column four documents the disclosed automation rate and explicitly records any observed AI washing red flags. When you complete this exercise, your audit will likely be filled with not-yet-measured rows.

That is not a failure. It is an honest map of exactly where your trust currently relies on hope. Building this table is the mechanism that forces a vendor out of the demo theater.

It replaces curated illusions with measurable operational reality. The final operational challenge is maintaining velocity. Demanding a massive thousand-case evaluation for every minor AI tool will paralyze your business.

To govern efficiently, apply the rule of proportion. You must match the depth of your proof to the specific operational stakes of the deployment. This matrix plots our operational stakes across two axes, the consequence of a system failure and the reversibility of the automated decision once it is made.

A low-consequence, highly reversible tool, like drafting an internal email that a human will review, carries minimal risk. A light check on a small sample is sufficient proof. High-stakes, irreversible automated decisions, like denying a financial benefit or issuing a clinical alert, demand exhaustive, adversarial evaluation on your proprietary data.

A governed trust decision scales with the risk, ensuring you can completely defend your methodology to a skeptical board of directors or an investigating regulator. Stop buying demos. Demand the denominator.

Ensure your trust is an artifact you can mathematically prove.

The ideas, one by one

An eval does not lie because of its structure, not the evaluator's virtue

It has a denominator (a rate over a sample), runs on inputs you chose to represent reality and probe failure, is repeatable by a skeptic, and has a pass line set before the result so the standard cannot drift. Those four properties are what turn a performance into evidence, and they are exactly the properties a demo lacks.

Distrust is a trained skill, not a mood

Cynicism refuses to believe; the skill decides what to believe by testing it, and can always name the evidence that would change its mind. The skill is a short list of questions asked every time: who chose the input and can I choose it instead, what is the denominator, what does failure look like and how often, what is actually doing the work, and where is the eval rather than the demo.

The automation rate is the most diagnostic single number in an AI purchase

When autonomy is the thing being sold, the fraction of outcomes that complete with no human is the thing being measured, and a seller's relationship to that number (states it, hedges it, or hides it) often tells you more than any demonstration. The Nate investors saw the button work and never asked the question; the Department of Justice later put the automation rate at "effectively 0%."

Humans in the loop are not the crime; hiding them while charging for autonomy is

Supervised AI is legitimate and frequently the better product, and you will deliberately place humans at the trust boundary later in this module (see Topic 4.4). Your audit flags the gap between autonomy claimed and autonomy shown, never the honest presence of human labor.

A benchmark can be a demo with a decimal point

A number is only evidence if you can answer whose test it is, who chose or funded it, and whether the scored system is the one you would deploy. Seller-tuned submissions and seller-funded benchmarks are real patterns (see Topic 1.4) (see Topic 12.2); interrogate a benchmark with the same three questions you ask of a demo.

Trust is a rate, and a rate needs a count

One success is a numerator with no denominator. Every "it works" hides a "how often, out of how many, on what inputs?", and only the answer to that question can tell you how much to rely on the system. The vividness of a single success is precisely what makes the missing count so easy to overlook.

Translating a claim into an eval is the action that follows the distrust

Name the claim precisely, choose a metric and a sample from your own inputs including the hard cases, set the pass line before you look, and make it repeatable. That translation is the deliverable, and it is the direct input to the eval suite you build next (see Topic 4.2).

The claims-and-evidence audit makes your trust inspectable

One row per claim, four columns: the precise claim, the evidence that exists today and its type, the test that would settle it, and the automation-and-honesty note. An audit full of honest "not yet measured" rows is not weak; it is a true map of where your trust rests on possibility, and it is the module's work list.

This audit is the first link in the evaluation chain

The tests you write become the eval suite in Topic 4.2 (see Topic 4.2), the human checks feed the trust boundary in Topic 4.4 (see Topic 4.4), and the whole audit becomes evidence in the evaluation report in Topic 4.6 (see Topic 4.6). Build it honestly and every later artifact rests on solid ground; build it on demos mistaken for proof and every later artifact inherits the illusion.

Match the depth of proof to the stakes

Distrust is proportioned, not uniform: a low-consequence reversible tool earns a light check, while a high-consequence irreversible decision earns a full evaluation on your own inputs with the failure cases included. Spending scrutiny where the risk is, and saying so, is more defensible and more welcome than demanding a study for every suggestion box, and it keeps the skill usable rather than obstructive.

The most dangerous demo is often your own team's

You drop your guard with people you trust, so an internal pilot run on hand-picked, cleaned data can smuggle a curated demo into your inventory as an evaluation. Apply the same four questions to your own team's pilot as to a vendor's: who chose the data, did it include the hard cases, what was the rate over a representative sample, and what is really doing the work.

You read it. Now prove it.

Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.

The conversation

The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.

Listen to it as episode 26 of the podcast.

Read the full conversation

Today we are, we're basically turning you into a human lie detector for AI. I like that framing, a lie detector. Yeah, welcome to the deep dive.

Because the core premise we're exploring today is, it's simple but it fundamentally changes how you will look at every single piece of technology you ever buy, build, or invest in. It really does, it rewrites your entire baseline. Right, and the premise is this, distrust is a skill.

It is not a mood, it's not this, vague feeling of paranoia or some cynical posture you adopt in a meeting to look smart. It is a highly structured, trained, analytics skill. And to understand exactly why you need to develop the skill right now, we need to take you back to 2018, to a company called Nate.

Oh, Nate. It really is, it's the ultimate anchor case for this entire topic. It is.

The Nate story is just a complete masterclass in how an illusion is constructed, how it gets funded by incredibly smart people, and how it inevitably shatters when reality finally crashes into the marketing. Yeah, let's really paint the picture for you. Imagine you are a venture capitalist in 2018.

A founder named Albert Sanager walks into your conference room to pitch the shopping app, Nate. And the pitch is essentially the holy grail of e-commerce, right? The absolute holy grail, total frictionless checkout. Exactly.

They told the world they had completely, entirely solved the friction of online checkout using artificial intelligence. And the demo they showed these investors was, I mean, it was breathtaking in its simplicity. It is beautiful.

Right. As a user, you just navigate to a product and you tap a single button once. That is it.

Just one tap. And from there, an autonomous AI agent supposedly just takes the wheel, right? Yes. It navigates the retailer's website, flawlessly fills in your name, your address, your credit card details, selects the shipping method, and completes the entire checkout process.

Without you doing anything else. Right. And they explicitly, aggressively claim this was all happening without human intervention.

So put yourself in that room. You're a VC. You are watching this screen.

You see the finger tap the glass. And moments later, the order confirmation pops up. Pure magic.

It's completely seamless. The latency is low, meaning it happens fast. It feels like you are literally watching the future arrive.

And that demo was so meticulously crafted, so deeply persuasive, that some of the most sophisticated venture capital funds on the planet wired over $40 million to back it. $40 million! Based on a button that worked perfectly in a conference room. Yep.

But here is the rug pull, and it is a massive one. Fast forward to April 2025. The reckoning.

The reckoning. The Department of Justice, specifically the Southern District of New York, along with the SEC, unsealed charges against Sanager for securities fraud and wire fraud. Wow.

And the punchline. The app's actual automation rate, according to federal prosecutors, was effectively zero percent. Zero.

I mean, just let that sink in. When that beautiful button was tapped in the app, there was no sophisticated neural network navigating the web. None.

There was no autonomous agent. That tap routed the order directly to a massive call center operations floor in the Philippines, and later to another operation in Romania. Real people.

Real people. Hundreds of human contractors were just sitting at desks, staring at screens, and frantically manually typing in the orders. That is wild.

They were just operating fast enough that a user on the other end could plausibly believe a highly advanced machine had executed the task. It's just staggering. $40 million incinerated because professional investors, whose entire job description revolves around due diligence.

Right. Literally their only job. Right.

Because they believed the theater of a demonstration. You know, in the field of human-computer interaction, this specific illusion actually has a name. Oh, really? Yeah.

It's known as the Wizard of Oz pattern. Wizard of Oz. Exactly.

It refers to a system where human labor is secretly performing the work that the audience is led to believe a machine is doing autonomously. Pay no attention to the man behind the curtain. Precisely.

Think of the classic movie. The great and powerful Oz is just a guy behind a curtain pulling levers, projecting this giant, imposing mechanical head. Yeah.

Now, if you were in a prototyping lab and you're just trying to test whether a user likes a new interface, the Wizard of Oz method is totally legitimate as long as you disclose it. Oh, sure. Because building the actual AI might take two years, but you want to know today if people will even click the button.

Exactly. You have a human simulate the AI just to see if the button placement makes sense. It saves time.

Right. But when it is concealed, when you take that lab trick out into the real world and you raise tens of millions of dollars by selling that concealed human labor as proprietary, autonomous, artificial intelligence. Yeah.

That crosses a line. That crosses a massive, bright legal line. And the fallout was brutal.

I mean, the investor's losses were near total. The company ran out of money, sold its assets in early 2023, and just left those serious venture funds holding an empty bag. A very expensive empty bag.

Which brings us to the core mission of our conversation today. Because every single one of those investors saw a demo that works. They saw it with their own eyes.

They did. But not one of them, at the precise moment they decided to wire the money, possessed an actual evaluation. That is the key difference.

So our goal today is to give you the exact framework to separate what a demo can prove from what only an evaluation can prove. We are going to ensure that you never fund, buy, or deploy the next NAIT. We have to close the gap between perception and reality.

Because the demonstration is structurally engineered to bypass your critical faculties. Structurally engineered. Yes.

We need to deconstruct how that engineering works and then replace that engineered feeling of trust with an actual architecture of truth. I love that. The architecture of truth.

So if a demo is an illusion, let's break down how the magic trick actually works. Let's do it. Because a demo really only answers one specific question, right? Is it possible? Just that.

Nothing more. But it seems like people immediately extrapolate that to mean, does it work all the time? How do sophisticated people get fooled so completely by that jump? We just fall into a trap of assumptions. You know, we tend to think that if someone shows us a tool working, they have provided evidence of its quality.

But a demo is not a test. A demo is a curated performance. It's designed from the ground up to persuade, not to inform.

Okay, let's talk about the first ingredient in that persuasion. It's something you call happy path bias. Yes.

The single most important fact you must internalize about any demonstration is that the presenter chose the input. They chose the input. They decided exactly what to type, what to click, or what to ask the AI.

And they decided it weeks in advance, knowing exactly what their system does brilliantly. Wow. This is the happy path bias.

In software engineering, the happy path is a specific route through the code where absolutely no errors occur. The perfect scenario. Right.

The presenter maps out this route, carefully stepping around all the messy, unpredictable edge case variables of the real world. You are watching the system's absolute best case scenario, selected by the person who has the highest financial motivation for you to trust it. You know, it makes me think of a movie trailer.

Oh, it's a great analogy. Right. When you watch a trailer, you are seeing the absolute best 90 seconds of a film.

Yep. It has been meticulously cut together by an editor whose only job is to get you to buy a ticket on Friday night. The trailer tells you that the movie has, you know, explosions.

It has a funny joke. It has a romantic kiss. It looks amazing.

But it tells you absolutely nothing about how disjointed, boring, or just terrible the actual 90 minute film might be. The trailer proves the studio had the budget to shoot one cool car chase. It does not prove they made a good movie.

Exactly. And in the software world, the equivalent of that bad movie is a system that just crashes the moment a user does something slightly unexpected. Right.

The demo tells you that the system can succeed on at least one input. It tells you that a capability exists. A capability.

Yes. But it provides zero information about how often it fails on the inputs you would actually give it in the wild. And this brings us to the most crucial distinction you have to make as an evaluator.

The difference between capability and reliability. Capability versus reliability. Okay.

I want to make sure I have this crystal clear because it sounds like conflating these two is the root of the whole problem. It is the root of almost every bad AI purchase. Wow.

Capability is simply doing a thing correctly in at least one curated case. Which is the demo. Exactly.

That is what a demo proves. It proves the thing is possible. Reliability is entirely different.

Reliability is doing that thing right consistently on inputs that you did not choose across a wide variety of unscripted chaotic situations. Real life. Real life.

And a demo structurally cannot prove reliability. And that all comes down to a missing piece of math, right? The denominator. The denominator is everything.

Everything. When you watch a demo, you see a success. The button is tapped.

The order is placed. The AI summarizes the 100-page legal document perfectly. You saw it work.

That single success is a numerator. It is the top half of a fraction. But a numerator floating in space by itself means nothing.

You saw the one time it worked. Was that one success out of one attempt? Right. Or was it? Or was it one success out of 500 attempts and the presenter is just hiding the 499 times the system completely hallucinated? You have absolutely no idea.

You have no idea because a demo has no denominator. But, you know, why doesn't our brain demand the denominator in the moment? Because, I mean, if someone tells me they made a basketball shot from half court, I usually want to know how many tries it took them, right? Of course. But in a tech demo, we just kind of nod and hand over the check.

Why? It is a massive cognitive blind spot. It's called the availability heuristic. The availability heuristic.

Yeah. Our brains are practically hardwired to favor vivid concrete stories over dry statistical aggregates. OK.

You see the button work in a slick interface and it's a story. It is emotionally vivid. It is cognitively cheap to process.

Right. It's easy. Reading a 50-page evaluation report with a thousand rows of testing data, that takes intense mental effort.

So what happens? Our brains silently, unconsciously invent a denominator that isn't there. Oh, wow. We see it work once and our brain automatically assumes, well, it must work like this 100% of the time.

The successes are highly available to your memory and the failures are invisible due to survivorship bias. Survivorship bias. You only see the version of the demo that survives rehearsal.

You do not see the 20 attempts that crashed the system at 2 a.m. the night before. That is so true. The demo is theater.

The presenter controls the lighting. They control the network. They control the data.

Yep. And if it glitches, they just blame the HDMI cable. Always the HDMI cable.

Always. But my actual business, the environment where this tool actually has to survive is the chaotic street outside that theater. So if a demo only proves a capability exists, how do we actually prove reliability? We have to strip away the performance.

What does the architecture of truth actually look like? This is where we shift from observing to evaluating. An evaluation or an eval does not lie. Okay.

But it is vital to understand that it doesn't lie because the person running it is some highly virtuous, deeply moral saint. It's not about them being a good person. No, it is not about the character of the vendor.

It is about structure. The architecture of a true test forces reality to the surface. There are four structural properties of true evidence that you must demand.

All right. Let's break these down because this is the real toolkit. Property number one has to be the denominator, right? We have to fix the math.

Exactly. A true evaluation measures an outcome over a defined sample. Okay.

It does not say it worked. It says it produced a correct result on 847 out of 1,000 cases. Ah, I see.

The denominator converts a vivid anecdote into a mathematically sound rate. Trust in an automated system is fundamentally a question about frequency. Frequency.

Yes. When I deploy this to 1,000 employees, how often will it be wrong? A demo has no frequency. Only an eval with a denominator gives you the frequency you need to make a governed trust decision.

Okay. So property one is the denominator, but, I mean, if the vendor gets to choose the 1,000 cases, they could just feed it 1,000 easy layups, right? Precisely. Which brings us to property two.

Self-chosen representative and adversarial inputs. Right. If the presenter isn't allowed to choose the inputs to make it look good, who does and what kind of inputs are we actually looking for? You choose the inputs, or an independent third party does, and you do not choose them to flatter the system.

You choose the inputs to probe its breaking points. You want to break it. Yes.

A real evil runs on cases that represent your actual daily grind. If you are a logistics company, you don't test the AI on perfectly formatted digital invoices. Right.

You test it on the crumpled, coffee-stained, handwritten shipping manifests that your worst supplier faxes to you in three different currencies. The absolute nightmare scenario. Exactly.

Crucially, the evil must include explicitly adversarial cases that the demo deliberately avoided. What do you mean by adversarial in this context? Like, are we trying to hack it? Not necessarily hack it in a cybersecurity sense, but stress test its logic. Okay.

You feed it conflicting instructions. You give it a document where the core premise is a logical paradox. You use unusual layouts, faint document scans, extreme edge cases.

Push it to its limits? The evil is built specifically to find the exact failures the demo was built to hide. That is why an evil and a demo so often disagree. And when they disagree, the evil is the only thing you believe.

I love that. Okay. Property three, repeatability by a skeptic.

I love the phrasing of this. A claim that cannot be checked by a skeptic is worth practically nothing. Right.

Think of it like the scientific method. A true evaluation is fully documented. The sample data, the labels of what constitutes a correct answer, the exact procedure used.

It's all laid out. They're all recorded so that anyone, even someone who actively wants the product to fail, can run the exact same test on the exact same system and get the exact same mathematical measurement. It's bulletproof.

It does not rely on the presenter's charisma, their specific laptop, or a rehearsed sequence of clicks. It is objective and repeatable. And that leads to the fourth property, which dives deep into human psychology again, a preset pass line.

Very important. The threshold for success must be set before you look at the results. I want to push on this.

Why is setting the pass line in advance so critical? Like if I run the test and it gets an 80%, can't I just look at that 80% and decide in the moment if I'm happy with it? No, because if you run the test first and set the pass line after, you fall victim to motivated reasoning and the sunk cost fallacy. Motivated reasoning. Right.

Imagine you are the executive buyer. You really want this tool to work. Yeah.

You've spent political capital advocating for it. You've hyped it up to the board. You want to be the innovative leader who brought AI to the department.

I'm invested. You are deeply invested. So you run the evaluation and it comes back at 72% accuracy.

If you haven't set a hard pass line in writing before the test, your brain will immediately start rationalizing to protect your ego and your investment. Right. I would probably say something like, well, 72% is actually pretty good for a first pass, right? The AI will learn over time.

And besides, we can just put a human on the other 28% to double check it. Exactly. You mentally move the goalposts to justify the purchase.

I'm negotiating against myself. You redefine success to fit the disappointing reality you just bought. Setting the bar before you see the result protects the measurement from yourself.

That is so smart. You decide what good enough means while you are calm, uninvested, and objective. Right.

You might say, we require 95% accuracy because anything less causes unacceptable downstream errors in our supply chain. Right. If it hits 94%, it fails.

The standard cannot drift. It's the same discipline as setting a pre-registered rollback trigger when deploying code. The bar you set before the result is the only bar you can trust.

That makes total sense. We really have to protect ourselves from our own desire to believe the magic. Exactly.

So we know demos are engineered for persuasion and evils are engineered for truth. But let's get practical here. If I am sitting in a conference room with a vendor or even with my own internal engineering team, how do I actually force them to provide a real evil? I can't just cross my arms and say, I think you're lying to me.

No, you can't, because that is destructive cynicism. Right. We have to clearly separate cynicism from structured skepticism.

OK, what's the difference? Cynicism assumes everything is fake and everyone is lying. If you are a cynic, you refuse to believe anything, which paralyzes the business. You cannot innovate if you reject every new tool out of hand.

Yeah, you'd never get anything done. Structured skepticism, on the other hand, is a repeatable procedure. It is a specific set of interrogative questions you ask of any AI claim to settle it.

Ah, OK. You aren't aggressively rejecting the claim. You are politely but firmly demanding the exact evidence that would change your mind.

OK, let's weave these into a real scenario. I'm the buyer. The vendor is running a beautiful demo on the screen.

It looks amazing. I want to buy it. But I'm practicing structured skepticism.

What is my very first move? Your first move is to attack the happy path bias. You ask question one, who chose the input and can I choose it instead? This is where I say, can I drive? Yes. The vendor shows you how flawlessly their AI tracks data from an invoice.

You say, that looks incredibly powerful. I'd love to see how it handles our actual workflow. Here is a PDF of a messy handwritten invoice from our most difficult supplier.

Can we run this one through right now? And what happens when they say no? Because they always have an excuse, right? Oh, the system isn't tuned for that specific format yet, or we don't have that module loaded on this demo environment. If they welcome the messy data, they are offering evidence. OK.

But if they hesitate, deflect, or give you those exact excuses you just mentioned, you have your answer. You are watching theater, not a product. OK, so I've disrupted the happy path.

Now they start throwing stats at me to recover. Well, our model is highly accurate. It works incredibly well across our client base.

What is question two? Question two is, what is the denominator? Ah, the denominator again? Every single time a presenter uses a vague, qualitative phrase like, it works or it's highly accurate, your reflex must be to force them into a rate. Pin them down. You ask, how often out of how many total attempts and on what specific inputs, if they claim 95% accuracy, you do not accept the number at face value? Right.

You ask, 95% of what? How large was the sample size? Were those easy cases or hard ones? The quality of their answer or their inability to produce one will instantly tell you if a real evaluation even exists. And what if it's not perfect, which leads to question three? I assume no AI is 100% perfect, so I need to ask, what does failure look like? Yes. You are forcing the vendor to describe their system's failure modes.

If you ask a vendor, show me exactly what happens when the model hallucinates or fails, and they cannot give you a specific decaled answer, you are in dangerous territory. Why? Because one of two things is true. Either they truly do not know how their own system fails, which means they haven't rigorously evaluated themselves, or they do know exactly how it fails and they are making a deliberate choice to hide it from you.

Oh, wow. Both scenarios should immediately disqualify them from your trust. That is so powerful.

Now we arrive at question four, and this is the big one. This is the question that, if the NAIT investors had asked it, would have saved them $40 million. Question four is, what is actually doing the work? The automation rate question.

This is the automation rate question. When you see a result on the screen, is it being produced by a massive large language model? Is it being produced by a human sitting in a call center? Is it being produced by a simple deterministic if-then rule? Or is it some blend of all three? And in what exact proportion? And the final question, question five seems to kind of wrap all of this up. Where is the evil? Stop looking at the screen with the flashy user interface and ask for the documentation.

Show me the receipts. Tell them, give me the measured result on a representative sample, define the metric you used, define the sample size, and tell me who chose those inputs. If a vendor can only offer you the demonstration, they are showing you a possibility and asking you to pay for a certainty.

I can hear executives listening to this right now and thinking, if I go into a vendor meeting and aggressively interrogate them with these five questions, they are going to get incredibly offended. It is going to ruin the partnership before it even starts. That is a very common fear.

But it completely misunderstands the market dynamics of honest technology. How so? Polite, rigorous due diligence is welcomed by honest vendors. Welcomed? Yes.

If a vendor has actually spent the time and capital to build a reliable, robust product, they want you to ask these questions. Oh, because it makes them look good. Because their answers will absolutely crush their competitors who are selling vaporware.

If a vendor takes offense to you asking for a denominator or asking to run your own data, that offense is highly diagnostic information in itself. It's a massive tell. It is a massive tell.

Politeness is not a reason to skip the questions. It is just a reason to ask them warmly. Ask them warmly, but ask them.

Let's zoom in on question four. What is actually doing the work? Because this brings us to the concept of AI washing. Yes.

If autonomy is the core feature being sold, then the automation rate is the ultimate test. It is the most diagnostic single number in an AI purchase. If a vendor markets their system using words like hands-free, autonomous, or requiring no human intervention, and then you ask for the automation rate, they hedge, deflect, or refuse to state as a hard mathematical percentage what fraction of outcomes complete without a human touching them.

Alarm bells. You should hear deafening alarm bells. Autonomy is not a vibe.

It is a highly measurable quantity. A seller of autonomy who refuses to measure it is making a deliberate choice to conceal reality. So this is what we call AI washing.

It's a term thrown around in headlines constantly, but let's define the actual mechanics of it. AI washing is the practice of presenting human labor. A simple handwritten rule, or bought-in commodity software, as if it were proprietary autonomous artificial intelligence.

Okay. And it is done specifically to raise venture capital, win an enterprise contract, or charge a massive premium over what the actual underlying technology is worth. It's a spectrum though, right? It sits on a spectrum.

On the extreme criminal end, you have the Nate case. Right. Claiming zero human intervention, while actually using hundreds of offshore workers to type in credit card numbers.

Total fraud. In the middle, you have misleading marketing. A product that is 30% machine learning and 70% human reviewers.

But it's sold aggressively as AI powered. Which technically isn't a lie, but it's highly deceptive. Right.

And on the mild end, you have technically true but empty claims. Like a vendor taking a basic 15-year-old spreadsheet macro and rebranding it as an AI workflow optimizer. I want to push back on something here because I think it confuses a lot of buyers.

I am buying software to lower my headcount or increase efficiency. If you are telling me that finding humans in the loop is a red flag, aren't we just saying that human supervision is bad? Shouldn't the AI be doing everything? This is a critical nuance. Humans in the loop are not the crime.

They're not. Supervised AI, where a human reviews the output, catches edge cases, or guides the model, is highly legitimate. In many high-stakes environments like medical diagnostics or legal contract review, supervised AI is significantly better and safer than full autonomy.

You want humans in the loop there. You want them. The defect of AI watching is not the presence of human labor.

The core defect is the concealment of that labor while charging the buyer for autonomy. It's the lie, not the labor. Exactly.

If a vendor tells you, look, our system automatically resolves 60% of tier one customer support tickets, and the remaining 40% of complex issues are automatically routed to our human specialists for review, that is a fantastic product. Right, because it's transparent. They are selling a supervised system, honestly.

You can price it correctly, you can staff your own teams for it, and you can govern it. Yeah. But if they claim it handles 100% autonomously and quietly use humans for the 40% without telling you they are AI washing, your audit should never flag this uses humans as a problem.

It should only flag this claims a level of autonomy it has not demonstrated, and it is hiding the gap. But AI washing isn't just hiding offshore labor in a call center, right? Sometimes the deception is purely technical. Oh, very often.

Let's look at example five from our source material. Simple deterministic threshold rules, basic if statements being sold internally as complex AI risk models. This happens frequently, and often it comes from internal enterprise teams trying to secure budget or prestige.

Okay, how does that work? You will be told that an internal tool is a sophisticated AI risk model for flagging fraudulent transactions. Sounds impressive. But when you apply structured skepticism and look under the hood, you find it is just a set of hard-coded handwritten rules.

If the transaction amount is over $10,000, and the IP address is in this specific list of five countries, then flag it. That's just an if statement. That is a deterministic rule.

It is basic logic. It is not machine learning. It is not AI.

But hold on. If it works, why does the label matter? If the if statement catches the bad transactions and stops the fraud, who cares if we call it AI or a neural network or a deterministic rule? It's just a word. It is much more than a word.

It matters immensely for governance, compliance, and liability. The underlying mechanism dictates how you manage the risk. Okay, unpack that for me.

Let's look at the technical difference. A deterministic rule is fixed logic. It does not need massive beta sets to train, and it does not change its behavior unless a human rewrites the code.

A machine learning model is probabilistic. It learns patterns from historical data, which means it carries entirely different obligations. If you mislabel a basic rule as an AI model, you will subject your engineering team to complex, expensive, totally unnecessary valuations that don't apply.

You're wasting time treating a calculator like a brain. Right, but the inverse is even more dangerous. How so? If you think a system is just a simple, stable rule, but it is actually a learned, probabilistic model, you might fail to monitor it for drift, bias, or hallucinations.

The mislabeling corrupts your entire organizational AI systems inventory. Oh, wow. The AI label dictates your defensive posture.

AI-washing with if statements absolutely breaks your governance framework because you don't know what kind of machine you are actually operating. Okay, so we know to look out for flashy stage demos, we know to hunt for hidden human labor, and we know to check if the AI is just a bunch of if statements in a trench code. Exactly.

But what happens when the demo wears the clothes of a rigorous evaluation? This introduces a totally different beast, disguised demos. This is where the persuasion becomes incredibly tricky, because the illusion is disguised as mathematics. Let's talk about public benchmarks and leaderboards.

A vendor comes into your office and says, we understand you want real data. Look, we don't just have a demo. We scored 94% on this massive industry-standard public benchmark leaderboard.

We are ranked third in the world out of 50 models. I mean, that sounds incredibly convincing. It sounds like a denominator.

It sounds exactly like the evil you've been telling us to demand. It does. But public benchmarks are very often just demos with decimal points.

Demos with decimal points. A number on a website is not automatically evidence of reliability for your specific use case. You have to ask the same structured questions.

Whose test is it? Who funded the creation of that benchmark? Where I follow money. Was this independent test quietly sponsored by the very vendor who miraculously ranks number one on it? You have to look for the conflict of interest. And you have to look for Goodhart's law, which states that when a measure becomes a target, it ceases to be a good measure.

Vendors frequently create specially tuned, highly over-optimized submissions that excel only at answering the specific questions on that public benchmark. They essentially teach the model the answers to the test. Right, they game the test.

But when you take that model off the benchmark and put it in the real world, it fails. That's terrifying. Even more insidious is the model they submitted to the leaderboard, the exact same version of the model they are deploying in your enterprise environment.

Let me guess. It's not. Often it's totally different.

The leaderboard rank reflects a highly optimized, expensive system you cannot actually buy. It is a curated performance wearing the clothes of an independent evaluation. But the danger of disguised demos doesn't just come from external vendors trying to trick us, right? Let's bring in example eight.

Cleaned internal data pilots. This one hurts because it's friendly fire. Yes, your own enthusiastic internal innovation team comes to you.

They ran a successful pilot of a new AI tool. They report it worked great. 98% accuracy.

Massive projected time savings. High fives all around. You trust your colleagues.

They work for the same company. So you drop your guard and approve the million dollar rollout. But three months later, it fails catastrophically in production.

The system is hallucinating. Users are angry. Why? What happened? Because that internal pilot ran on a handpicked, heavily cleaned, totally sanitized subset of data.

The team manually curated the pilot data to ensure the project would look good and get approved by leadership. They took out all the messy, adversarial, real world edge cases. Right.

They built a happy path for themselves. They curated a happy path to secure their budget. That pilot was not an evaluation.

It was an internal demo in disguise. Wow. You must apply the exact same structured skepticism to your own internal teams as you do to a hungry external vendor.

Who chose the pilot data? Did it include the unhappy paths? If not, it proves nothing about production reliability. And the stakes for getting this wrong are escalating rapidly. It's no longer just about internal embarrassment or a wasted software budget.

Regulators are officially entering the chat. They absolutely are. Let's look at examples two and seven from the source material.

The regulator subpoena. We are seeing a major shift in the legal landscape. U.S. state enforcement agencies and attorneys general are now issuing subpoenas to tech vendors.

Specifically, health tech vendors demanding the raw math and the actual denominator behind their public marketing claims about critical hallucination rates. They want the receipts. A vendor publicly claimed their medical AI system practically never hallucinated.

The regulator stepped in and said, prove it. Show us the evaluation. Show us the sample size.

Show us the representative inputs. And it demoed well in the boardroom is not a legal defense. It is absolutely not a defense.

Regulators are increasingly demanding the mathematics behind the marketing. If you deploy a high stakes AI system and it fails, causing harm to a consumer or a patient, the regulators will ask for your claims and evidence audit. Right.

If you could only point to a slick marketing demo or a seller funded leaderboard benchmark, you are completely legally exposed. You have no defensible basis for your trust decision. I want to provide a breath of fresh air here, though.

Not everyone is hiding the ball, right? Example six from our sources shows what honest reliability reporting actually looks like in practice. The published autonomous agent experiment. There was a fascinating case where researchers built an AI shopping bot, an autonomous agent designed to negotiate prices.

And they released it into the wild to see how it would handle real world unscripted negotiations with humans. Sounds risky. It was.

They published the results completely transparently. And in those results, they admitted that the bot was successfully talked into, giving away $1,000 in discount losses to clever adversarial users. They published their own failure.

They documented exactly how they got beat. And from a governance perspective, that transparency is gold. It is the exact opposite of AI washing.

They didn't just show a tightly controlled demo of the bot successfully closing a perfect sale. They showed the reality. They measured its reliability over many real interactions.

They exposed the unhappy paths. They found the denominator. And they documented the exact failure modes.

That is what transparent reliability reporting looks like. An organization that openly publishes how its agent fails is modeling the exact honesty your internal audits should demand from any vendor. So we have this powerful framework.

We know how to interrogate a vendor. We know the difference between a demo and an eval. We know how to spot AI washing.

And we know not to trust public benchmarks blindly. Right. But in the real world, under immense pressure to ship product, executives still make critical procedural mistakes.

Let's hit some of the most common traps so our listeners can avoid them. Let's do it. The first trap is conflating demoed well with working.

We've hammered this, but it bears repeating because the psychological pull is so strong. A demo proves possibility in one curated case. It never proves reliability.

Do not let a good feeling in a conference room replace a denominator. Trap two, mistaking a SOC 2 report for an evaluation. I see this constantly.

A vendor hands over a massive 100-page SOC 2 compliance document. It has stamps and signatures from auditors. And the buyer says, great, they're certified.

The model is accurate. I love the restaurant analogy for this. SOC 2 report is like a health inspector's certificate framed on the wall of a restaurant.

OK. It proves the kitchen is clean. The refrigerators are kept at the right temperature.

And the staff washes their hands. It says absolutely nothing about whether the food actually tastes good or if the chef followed the right recipe. That is so clear.

SOC 2 attests to operational security, data protection, and access controls. It does not attest to model accuracy, hallucination rates, or automation percentages. It is a real, vital, and useful artifact.

But it is not a model performance evaluation. Trap three. Believing that if a system were mostly human, we would naturally notice.

We think we are too smart to be fooled by the Wizard of Oz. We think I'd spot a guy typing in the Philippines. The investors in Nate were highly sophisticated venture capitalists.

They could not tell. Well-built AI washing is designed specifically to produce results at human plausible latency. If a contractor in a highly organized call center uses hotkeys, macros, and autofill software, they can complete complex tasks fast enough to feel like an automated system to the end user.

You cannot detect the hidden human simply by watching the output speed. You only detect it by forcing the automation rate question and running the system on unstructured, messy inputs the vendor couldn't possibly prepare a macro for. Trap four.

Refusing to evaluate due to speed pressures. You're in a meeting and someone says, we don't have time for a full evil. We need to deploy this quarter to hit our OKRs.

The demo is convincing enough. Let's just ship it. This is a psychological trap.

There is a golden rule for this. The strength of your urge to skip the evaluation is directly proportional to how persuasive and engineered the demo was. Oh, that's good.

The moment you feel the least need to check the system, because the demo was just so magical, is the precise moment the check matters the most. Time spent evaluating before deployment is incredibly cheap compared to the staggering cost of defending a system that fails in production because it was never what the vendor claimed. Trap five.

Treating an evaluation as a permanent license instead of an expiring certificate. If I run an evil on January 1st and it passes, why can't I just trust it forever? Microsoft Word doesn't suddenly forget how to bold text after six months. Software doesn't expire.

Software logic doesn't expire. But machine learning models suffer from model drift. A model is a probabilistic representation of the data it was trained on.

As the real world changes, as consumer behavior shifts, as language evolves, as global supply chains alter it, the world moves away from the model's original training window. OK, give me an example. Think about a fraud detection model trained in 2019.

When the pandemic hit in 2020, consumer purchasing behavior changed overnight. Right, everyone's buying online. People were buying different things from different locations at different times.

The model, which was highly reliable in 2019, suddenly started flagging legitimate purchases as fraud because the data distribution shifted. The world changed, but the model didn't. A pass today does not guarantee a pass in six months.

The repeatable evil you design must be rerun on a schedule and immediately after any vendor update. A pass is a license with an expiry date. OK, I have to push back here on behalf of every executive listening.

If we demand a 1,000-case, fully repeatable, adversarial-owned data evaluation for every single AI tool we touch, business will grind to a complete halt. We use dozens of AI tools. We will never deploy anything.

How is this actually scalable? It isn't scalable to do it for everything, and you shouldn't do it for every tool. That is miscalibration, not rigor. OK.

If you demand a massive, expensive evaluation for every minor feature, the business units will simply route around your governance entirely. The skill of distrust is proportion suspicion. We call it the proportioning rule.

You match the depth of proof to the stakes of the decision. How does the formula for the proportioning rule work? It is consequence times reversibility times run frequency, right? Exactly. Let's look at the axis.

Consequence. How much harm does a wrong output cause? Reversibility. Can the harm be easily undone once acted upon? OK, give me a low-stakes example.

If you are deploying an AI tool that simply suggests a draft reply to a customer service email, and a human agent will read, edit, and click send on that email, the consequence of a bad draft is low, and the action is entirely reversible before it goes out. Right. That earns a very light check.

A solid demo plus a small spot check on your own data is a perfectly defensible basis for trusting it. But if we flip the stakes... Right. If you are deploying an AI tool that generates clinical alerts in a hospital, or a system that automatically declines credit card applications for fraud without human review, that is high consequence and irreversible.

A mistake causes immediate harm. You can't just spot check that. No.

That requires a full, rigorous, own-data evaluation with adversarial cases. Nothing less is legally or ethically defensible. And run frequency.

And you multiply this by run frequency. Even a low-consequence tool, if it runs 10 million times a day, aggregates massive risk over time, requiring deeper proof. Rigorous governance doesn't mean stopping the business.

It means spending your deep scrutiny exactly where the catastrophic risk is, and letting the low-stakes tools move faster. Which perfectly sets up our close. We need a practical, scalable way to execute everything we've talked about when we walk into the office on Monday morning.

And that is the claims and evidence audit. This is the single most valuable action you can take to protect your organization. Yeah.

It is the artifact that forces truth out into the open. Okay. You build a structured four-column table, and you rank the rows by trust risk, putting those high-consequence, irreversible claims at the very top of the list.

Okay, I have my spreadsheet open. Walk me through how I fill out these four columns, so it doesn't just turn into a useless checklist. Column one is the precise claim.

Do not write down the marketing fluff. You don't write, it is highly intelligent, or it optimizes workflows. You write the specific, checkable behavior the vendor is promising.

It correctly extracts the invoice total and vendor name from scanned, handwritten supplier PDFs. Very specific. Okay, what's next? Column two is evidence today.

What actually backs this precise claim right now? The only thing you've seen is a demo. You write demo only. Ouch.

If it's a public benchmark, you know who funded it. You must be brutally honest with yourself if your evidence is weak. Okay, column three.

Column three is the test that would settle it. This is your blueprint for the evil. You define the metric.

You define the representative sample from your own messy internal data. You state the hard pass line you require before you buy. And the last one.

Column four is the automation and honesty note. This is where you record the answer to the automation rate question. What is actually doing the work? Are there hidden humans? Is it just an if statement? Are there any AI washing tells? It is brilliant because it makes the gaps impossible to ignore.

It forces the uncomfortable conversation. Just look at E&R's operations review at Cedarline Freight in Manchester. Using this exact audit structure, force the vendor to admit they routed a third of complex assignments to a human queue, instantly transforming a risky, autonomous illusion into a highly governed, supervised reality.

Ian didn't have to accuse anyone of fraud. He didn't have to be cynical. He just methodically applied the four columns, asked for the automation rate as a hard number, and the vendor came clean.

Yeah. Ian ended up buying a supervised system. He staffed his teams for it correctly, and he avoided deploying a dangerous illusion.

That is what a governed trust decision looks like on one page. A skeptical board member could read that audit and find absolutely no place where a feeling was recorded as a fact. We have covered incredible ground today.

We started with the Nate investors losing $40 million on a flawless fake demo, and we ended with a simple four-column audit that would have saved them every single penny. We explored the psychology of the happy path bias, the mathematical power of a denominator, the anatomy of true evaluations, and the broad spectrum of AI washing. And we learned that distrust, when applied as a structured proportion skill, isn't about being a cynic who blocks progress.

It is about demanding the specific evidence required to make innovation actually safe. So I want to leave you with a final provocative thought. Think about the AI system that your organization relies on the most right this second, the one driving your critical workflows.

If I walked into your office today and I asked you to show me the denominator behind its reliability claim, would it exist? Could you show me the math? Or are you just running your business on a really good demo? Thank you for joining us for this deep dive. Keep asking for the denominator.

Real cases

These examples show the demo-versus-eval gap and the AI-washing tells in real cases, with the analytic reasoning stated. The deep anchor is the Nate case; the others are referenced to sharpen a specific point and are owned, for deep treatment, by their own topics.

Example 1 (the anchor): the Nate shopping app. Nate, founded in 2018 by Albert Saniger, marketed an app that let a user buy from any online store with a single tap, performed by an AI agent "without human intervention" except in edge cases. The demonstration was exactly the kind that answers Question 1 (is it possible?) vividly: a tap, then an order confirmation. What no investor evaluation established was Question 2 (how reliably, across real stores and real carts?) and Question 3 (what is actually placing the orders?). The answer to Question 3, per the Department of Justice, was that hundreds of human contractors in a call center in the Philippines, and later in Romania, manually completed the purchases, and the "actual automation rate was effectively 0%" (US DOJ SDNY, "Tech CEO Charged In Artificial Intelligence Investment Fraud Scheme," April 2025). Nate raised more than forty million dollars on the strength of the demo, ran out of money, and sold its assets in January 2023, leaving investors with near-total losses; in April 2025 prosecutors charged Saniger with securities fraud and wire fraud, and the SEC filed a parallel civil action. Read through this topic's lens, every AI-washing tell from 3E was present and unpressed: no disclosed automation rate, orders completed at human-plausible speed, and autonomy claimed far beyond anything an evaluation had shown. The lasting lesson is not that fraud is possible; it is that a demo can be fully persuasive and carry effectively zero of the reliability and honesty it seems to promise, and only a question no one asked ("what fraction runs with no human?") would have exposed the gap.

Example 2 (a demo whose reliability was measured and found wanting): a hospital's own hallucination-rate check. A generative-AI vendor selling to hospitals is a setting where the demo-versus-eval gap is a matter of patient safety, and where a US state enforcement action established that a vendor's own reliability claims about its "critical hallucination rate" were the thing that mattered and had to be substantiated rather than asserted (the deep treatment of that enforcement action, and of building the eval suite that tests such a rate, belongs to Topic 4.2 (see Topic 4.2)). The analytic point to carry here is narrow and general: the claim "our system rarely hallucinates" is a claim with a denominator hiding inside it (rarely, out of how many, on what inputs?), and the only responsible response to it is to demand that denominator and, where you can, measure it yourself on your own cases. A demo of a correct summary proves the correct summary was possible; it says nothing about the rate of the dangerous error, and in a clinical setting the rate is the entire question.

Example 3 (autonomy claimed, humans doing the work): the offshore-labor pattern. A recurring shape of AI washing is a product marketed as an autonomous AI agent while a large fraction of its work is performed by remote human workers, a pattern serious enough to have drawn US securities enforcement in its own right (that specific vendor-interrogation case is owned by Topic 3.3 (see Topic 3.3)). The reason to name the pattern here, without re-telling that case, is that it isolates the single most useful probe in this topic: the automation rate. Two products can look identical in a demo, one genuinely autonomous and one 70 percent human, and the demo cannot tell them apart because both show a tapped button and a good result. The only thing that separates them is the number the honest one will state and the dishonest one will not. When autonomy is the thing being sold, the automation rate is the thing being measured, and a seller's relationship to that number (states it plainly, hedges it, or hides it) is often more diagnostic than any demonstration.

Example 4 (the benchmark is not automatically the eval): a seller-tuned submission. A vendor's impressive score on a public leaderboard feels like the evaluation this topic is asking for, but a number can be a demo in disguise. A model can be specially tuned for a benchmark and submitted in a form that differs from the version a customer would actually receive, so that the leaderboard rank reflects a system you cannot buy (see Topic 1.4); and a benchmark can be quietly funded by the very organization being measured, a conflict that changes how much the score is worth (see Topic 12.2). The analytic move is to treat a benchmark with the same three questions as a demo: whose test is it, who chose or funded it, and is the scored system the one I would deploy? A benchmark that survives those questions is real evidence; one that does not is a number wearing an eval's clothes.

Example 5 (a rule dressed as intelligence, applied to your own estate): the "AI" that is an if-statement. Not all AI washing is fraud or offshore labor; a common and quieter form is a simple deterministic rule sold as machine learning. Consider an internal tool your organization was told is an "AI risk model" that, on inspection, applies a handful of hand-written thresholds (if the amount is over X and the country is on a list, flag it). This may even be a perfectly good tool; the defect is the mislabel, and the mislabel matters for governance because it changes what you must evaluate, disclose, and defend. The tell is that the "model" cannot be shown any training data, any evaluation on held-out cases, or any behavior that a rule could not produce. The probe is to ask what it learned and from what, and to notice when the answer describes an author writing rules rather than a system learning patterns. Your AI systems inventory (see Topic 0.2) is where this correction gets recorded, because an accurate inventory distinguishes a learned model from a relabeled rule, and the two carry different obligations.

Example 6 (an autonomous agent that was genuinely tested, and revealed): running a real shop. The opposite of AI washing is a system whose autonomy is measured honestly, including its failures, and a published experiment in which an AI agent was left to run a small real-world shop and was talked into discounts and giveaways until it finished its run well below where it started is a clean illustration of what an honest reliability test looks like (that case is owned, for its economics, by Topic 8.6 (see Topic 8.6)). The point for this topic is the contrast: the useful knowledge was not that the agent could take an order (the possible question, easily demoed) but the measured, unflattering record of how it behaved over many real interactions (the reliable question), reported rather than hidden. An organization that publishes how its agent fails is doing the opposite of AI washing, and it is modeling exactly the honesty your claims-and-evidence audit demands of every vendor and of yourself.

Example 7 (the demo that a regulator later measured): the state that put the vendor's own number on trial. Enforcement has begun to do, with subpoena power, exactly what this topic asks a buyer to do voluntarily: demand the denominator behind a vendor's reliability claim. When a US state attorney general took interest in a healthcare AI vendor's public statements about how rarely its system produced dangerous errors, the pressure was not on whether a demo could show a correct output (it plainly could) but on whether the reliability rate the vendor advertised was substantiated by evidence rather than asserted (deep treatment of that action, and of building the eval suite that tests such a rate, is owned by Topic 4.2 (see Topic 4.2)). The lesson to carry, without re-telling that case, is that "how well does it work?" is increasingly a question you will have to answer to an outside authority with a measurement, not a demonstration, so the claims-and-evidence audit you build here is not only good practice, it is a rehearsal for the day someone with legal power asks you to show your denominator. A trust decision you cannot defend to a skeptical board is one you probably cannot defend to a regulator either.

Example 8 (the internal demo that was never your own data): the pilot that proved nothing about you. A frequent and self-inflicted version of the demo-versus-eval gap happens entirely inside an organization, with no vendor to blame. A team runs a "successful pilot" of an AI tool, reports that it "worked great," and the tool is rolled out, and only later does someone notice that the pilot ran on a hand-picked, cleaned subset of data that a team member curated to get the project approved. The pilot was an internal demo wearing an evaluation's clothes: the inputs were chosen, there was no representative denominator, and the hard cases were quietly excluded. The tell is the same as for a vendor demo, applied to yourself: ask who chose the pilot data, whether it included the messy and adversarial real cases, and what the rate was over a representative sample rather than the curated one. The distrust skill is not reserved for outsiders; the most dangerous demo is often the one your own enthusiastic team showed you, because you were not on guard against people you trust. Your claims-and-evidence audit applies the same four columns to internal claims as to vendor ones, precisely so an internal pilot cannot smuggle a demo into your inventory as an evaluation (see Topic 0.2).

Where people go wrong

  • "It demoed really well, so it works." A demo shows the vendor's best case on the vendor's chosen input in the vendor's environment. It proves the capability is possible in at least one case; it says nothing about how often it succeeds on your inputs. "Demoed well" is an answer to "is it possible?", never to "can I rely on it?" Treat a good demo as a reason to build the eval, not a reason to skip it.
  • "Distrust means assuming everything is fake." That is cynicism, and it is as useless as credulity. The skill is structured skepticism: you decide what to believe by testing it, and you can always say exactly what evidence would change your mind. A cynic and a skilled analyst both doubt the demo; only the analyst knows what would resolve the doubt and asks for it.
  • "They showed us a number, so we have an evaluation." A number is only as good as who chose the test, who chose the sample, and whether the scored system is the one you would deploy. A seller-selected, seller-funded, or seller-tuned benchmark is a demo with a decimal point. Ask the same three questions of a benchmark as of a demo before you treat it as evidence.
  • "Humans in the loop means the vendor is lying." No. Supervised AI is legitimate and often the better product, and Module 4 will have you deliberately place humans at the trust boundary (see Topic 4.4). The defect AI washing names is not the presence of humans; it is the concealment of them while claiming and charging for autonomy. Flag the hidden gap, not the honest human.
  • "If it were mostly humans, we would obviously be able to tell." The Nate investors could not tell, and they were sophisticated. Human labor completing tasks fast enough to feel automated is precisely what a well-built AI-washing product delivers. You cannot detect the hidden human by watching the output; you detect it by asking for the automation rate and by running the system on inputs the vendor did not prepare.
  • "One success is evidence of reliability." One success is a numerator with no denominator. You saw the time it worked; you did not see how many times it was tried. Reliability is a rate, and a rate needs a count. Always ask "out of how many, on what inputs?" The vividness of the single success is exactly what makes this error so easy to fall into.
  • "We do not have time to evaluate; the demo was convincing enough." The strength of your urge to skip the evaluation is proportional to how convincing the demo was, which is proportional to how carefully it was constructed to convince you. The moment you feel least need to check is the moment the check matters most. Time spent evaluating before you deploy is cheap against the cost of defending a system that was never what you said it was.
  • "An audit full of 'not yet measured' rows is a failure." It is the opposite: it is an honest map of where your trust currently rests on possibility rather than proof, and it is the exact to-do list of evals to build next (see Topic 4.2). A polished audit with every row marked "proven" and no denominators behind the proof is the one to distrust.
  • "The vendor benchmark and my use case are close enough." Close enough is a phrase that hides a denominator. A benchmark on general documents tells you little about your invoices; a fraud model tuned on someone else's transactions tells you little about your fraud. The eval that settles a claim runs on your inputs, because the gap between "their sample" and "your reality" is where the trust breaks.
  • "Once it passes an eval, we are done evaluating." An eval is a measurement at a moment, and systems drift as the world moves away from their training window (see Topic 4.6). A pass is a license with an expiry, not a permanent certificate. The repeatable eval you designed here is designed to be re-run precisely because today's rate is not forever.
  • "Asking these questions will offend a good vendor." A good vendor expects them and answers them, because a buyer who does due diligence is a buyer who will not churn in three months when reality diverges from the pitch. The vendor who is offended by "let me run it on my own data" is giving you information. Politeness is not a reason to skip the questions; it is a reason to ask them warmly.
  • "Distrust means demanding a full evaluation of everything." No; that is miscalibration, not rigor, and it teaches the organization to route around governance. The skill is proportioned: match the depth of proof to the stakes, so a low-consequence reversible tool earns a light check and a high-consequence irreversible decision earns a real evaluation on your own inputs. Spending your scrutiny where the risk is, and saying so, is more defensible and more welcome than demanding a thousand-case study for a suggestion box.
  • "An internal pilot our own team ran is trustworthy because they are not trying to sell us anything." The most dangerous demo is often the one your own enthusiastic team showed you, because you drop your guard with people you trust. A pilot on a hand-picked, cleaned subset chosen to win approval is an internal demo wearing an evaluation's clothes. Apply the same four questions to your own team's pilot as to a vendor's demo: who chose the data, did it include the hard cases, and what was the rate over a representative sample.
  • "They have a SOC 2 report, so it is covered." A SOC 2 report is a real and useful artifact, and it is answering a different question than the one your audit asks. It typically attests to security and operational controls (how data is protected, who can access it, how changes are managed), not to whether the model performs the claimed task reliably or how much of the work is human. A vendor's SOC 2 says little or nothing about the automation rate or the accuracy of a claim; ask directly whether the report, or any addendum to it, covers model performance at all, and do not let a real certification stand in for the evaluation it was never designed to be.
  • "We ran the evaluation, so the number is settled." An evaluation can itself be wrong: a sample that quietly excluded the hard cases, a metric that measures something adjacent to the decision, contaminated data, or a scoring rule applied inconsistently can all produce a confident-looking rate that does not mean what it appears to mean. A bad evaluation is more dangerous than no evaluation, because it supplies the appearance of rigor without the substance, and a false confidence is harder to question than an open "not yet measured." Before you rely on a passing rate, apply the same scrutiny to the test that you applied to the demo: who chose the sample, what exactly counted as a pass, and could the number be checked by someone else.
  • "If a regulator ever asks, we will explain that it demoed well." "It demoed well" is not a defense to a board and it is not a defense to a regulator, who increasingly ask for the denominator behind a reliability claim with legal power to compel it. The claims-and-evidence audit is your rehearsal for that day: if you cannot show a measurement on your own inputs now, you will not be able to show one when it is demanded, and the gap between the autonomy you claimed and the autonomy you measured becomes the finding.

Questions people ask

What is claims-and-evidence audit?
The one-page artifact this topic produces for an AI system. It is a table with one row per claim the system makes, and four columns: the claim stated precisely, the evidence that exists today and its type, the test that would settle it, and the automation-and-honesty note. It is the first artifact of the Evaluation and Trust module and the input to the eval suite built in Topic 4.2.
What is demo (demonstration)?
A shown performance of a system on inputs the presenter chose, in conditions the presenter controls. It can prove that a capability is possible in at least one case; it cannot prove how reliably the capability holds on inputs the presenter did not choose, because it has no denominator.
What is evaluation (eval)?
A repeatable, measured test of how a system behaves across a defined sample of inputs, reported as a rate against a metric. Its trustworthiness comes from four properties: a denominator, inputs chosen to represent reality and probe failure, repeatability by a skeptic, and a pass line set before the result.
What is denominator?
The count of cases a rate is measured over ("correct on 847 of 1,000"). A claim without a denominator ("it works," "80 percent") is an anecdote until the count and the sample behind it are stated; demanding the denominator is the core move of turning a story into a rate. More on Denominator
What is automated versus autonomous?
Automation means the system executes a predetermined set of instructions; autonomy means the system decides, on its own, which instructions apply to a situation it was not specifically scripted for. A demo can show automation (the task completed) without ever showing autonomy, because the input was chosen to be one the system, or a hidden person, was ready for. The distinction matters because "autonomous" is the word most often oversold, and the gap between the two is where AI washing hides.

Keep going