The vendor interrogation: questions that expose what the sales deck hides
The short answer
The deck is edited truth, not lies
A sales deck is optimized to close a deal, so it minimizes, redefines, or omits everything that would slow the deal down. Most vendors are not committing fraud; they are selecting. The interrogation exists because the selected-out information (who does the work, whose model it is, what the metric counts) is exactly what becomes your incident later.
What you will be able to do
- Explain why a sales deck systematically hides certain classes of information, and name the incentive structure that produces "AI washing" (overstating how much artificial intelligence actually does the work).
- Apply the seven-family vendor interrogation framework (demo-to-deployment gap, metric definition, human-in-the-loop reality, model provenance and ownership, failure and incident history, data handling, and evidence in writing) to any AI vendor you are considering.
- Analyze a vendor's answer and classify it as a confirm (verifiable commitment), a dodge (a true-sounding non-answer), or a refuse (a decline that is itself information).
- Distinguish a headline metric a vendor reports from the definition underneath it, and construct the follow-up question that exposes the definition.
- Produce a written vendor interrogation record: the questions asked, the answers given, what was confirmed, what was dodged, and what the vendor would and would not put in the contract.
- Connect the interrogation to your accumulating dossier: the AI systems inventory that told you this system exists, the scoping decision that said you would buy or wrap rather than build, and the eval suite and conformity file that will later have to test whatever the vendor actually delivered.
- Judge when a vendor's answers are good enough to proceed, when they demand a narrower deployment, and when they are a reason to walk away.
- Trace a single transaction end to end through a vendor's described workflow to expose the humans and third parties an "it is automated" claim conceals.
- Convert every material verbal claim into a specific written-commitment ask (warranty, service level, audit right, or indemnity) and record what the vendor will and will not sign.
The lesson
A customer pulls up to a drive-thru. They speak a messy order over a hissing soda machine, and an AI reads it back flawlessly. The sales engineer hits pause in the conference room.
The deal moves forward. Hold that flawless clip up to the light. In January 2025, the U.S. Securities and Exchange Commission charged a company that sold exactly this kind of product.
They told investors their voice AI eliminated human order-taking, boasting a 95% automated completion rate. The reality was decidedly manual. Offshore workers in the Philippines and India were entering roughly 70% of those orders by hand.
The sales deck is not a forged document. It is a carefully edited truth. Every number is technically defensible, if you already know the private definition the vendor is using.
You are stepping into an environment of extreme information asymmetry. The vendor knows exactly what the system is, and exactly what they removed to make the presentation shine. Standard software procurement tactics won't save you here.
Checking uptime or SOC 2 compliance does absolutely nothing to catch AI washing, the practice of overstating how much artificial intelligence actually does the work. To close that knowledge gap, you need a targeted operational upgrade. This is the 7-Family Vendor Interrogation Matrix.
Its sole objective is to force hitting variables into the light, so you can stop being a passive audience for a pitch, and start doing the aggressive diligence required before you sign. We start with Family 1, the demo-to-deployment gap. Every demo you see is a best-case scenario masquerading as a typical day.
You must force the vendor to run your organization's hardest edge cases on the exact production model. Babylon Health marketed an AI symptom checker with flawless promotional claims, but a wide demo-to-deployment gap meant real-world performance fell drastically short. You have to test reality before you buy it.
That brings us to Family 2, Metric Definition. Every metric rides on a definition. You have to get the definition first.
When a vendor claims 95% accuracy, you need to know the denominator, the exact rule for counting a success, and precisely who or what is excluded from the count. Requesting the specific denominator and counting rule prevents the vendor from shifting the goal posts. A high percentage looks impressive until you realize it only measures the easy cases, while a massive excluded population is quietly ignored.
This is exactly how Presto's 95% automated metric functioned. Their counting rule defined automated as an order completed without the restaurant's own staff intervening, which quietly excluded all the offshore humans doing the actual work. Accepting a headline percentage without forcing the vendor to name exactly who and what is left out guarantees you will purchase a misleading metric.
Family 3 is the human-in-the-loop reality. The rule here is simple. Trace the humans and the third parties.
Take a single transaction and force the vendor to map it end-to-end. Name every human who touches it, where they sit, and who employs them. Major autonomous systems regularly hide manual labor.
Amazon's JustWalkOut technology relied heavily on remote video reviewers, and the Department of Justice charged the founder of the shopping app Nate over an AI checkout system allegedly run by human workers offshore. Family 4 examines model provenance. You are about to depend on a model, and you need to find out whose it is.
Buying a wrapper where the vendor does not actually own the underlying core model means your reliability is tied to a supply chain you cannot see. If that undisclosed upstream provider changes their pricing, deprecates the model, or cuts off access, your downstream service fails instantly. An undisclosed human-in-the-loop or third-party dependency is not a feature.
It is a hidden, unpriced risk that completely invalidates your capacity planning and cost models. Family 5 forces out the failure history. No system is error-free.
You must reject process-based answers like, we test extensively, and demand to see actual field outcomes. Make them show you their absolute worst production output to date. The gap between what they hope the system does and what it actually misses is where your incident lives.
You are specifically hunting for silent failures. If a support bot invents a non-existent company policy and states it as fact, the system doesn't throw an error code. The customer sees the fabrication before your monitoring system catches it.
Family 6 establishes data handling. You have to trace exactly what happens to your corporate and customer data the moment it enters the vendor's ecosystem. Is your data being used to train their models? Finding a do-not-train toggle hidden in a UI settings page is not enough.
That exclusion must be a legally binding boundary written into the contract. By forcing the vendor to admit exact failure modes and set rigid data boundaries, you dictate the specific alerts and monitoring framework you will need to build upon deployment. That leads to Family 7, evidence in writing.
A vendor will say almost anything in a meeting, but the contract is the product, not the conversation. To separate talk from commitment, run every answer through a strict analytical filter. Classify every response as a confirm, a dodge, or a refuse.
A confirm is a specific, verifiable commitment that the vendor is willing to warrant in writing. These are the load-bearing beams of your decision. A dodge is a true-sounding non-answer, a redefinition of your question, or an appeal to social proof rather than evidence.
Dodges require written follow-ups. And a refuse is a decline to answer or commit. A refusal is information, not a wall.
It pinpoints the exact location where the vendor believes the truth will cost them the deal. The gap between a vendor's verbal assurance in a meeting and what they will actually warrant on paper is the true product you are buying. Let's dismantle a few common buyer mistakes.
Mistake 1 is conflating the polish of a confident sales presentation with the mechanical reliability of the system. Confidence tells you about the salesperson. It tells you nothing about the software.
Mistake 2 is assuming that hidden humans in the loop act merely as a safety net. In reality, undisclosed human labor is a hidden bottleneck, a single point of failure that limits scaling when that workforce becomes unavailable. Mistake 3 is the fatal error of deferring hard technical questions until the contract negotiation phase, assuming you'll figure it out later.
The pre-signature interrogation is your last cheap moment. Your leverage as a buyer peaks when the vendor still wants your business. The moment you sign, that leverage collapses entirely, and your switching costs begin to climb.
Every material question you fail to force before signing becomes a question you will be forced to litigate after a live incident. Your Monday morning action is to compile all of these answers into a written interrogation record. This strips away the sales narrative, leaving only what the vendor will actually commit to.
This record forces one of three definitive outcomes, proceed, walk, or the most common good outcome, proceed narrowed. A credible vendor is often safe, but in a much smaller scope than they market. By narrowing the deployment to only the specific menu items or use cases the evidence actually supports, you isolate verified capabilities from the vendor's hype.
The flawless demo is real, and the 70% of orders entered by hand are real. Your fundamental duty as a buyer is to figure out which world you are acquiring before the contract is signed, not after the system fails.
The ideas, one by one
Every metric rides on a definition; get the definition first
A number without its denominator, counting rule, and excluded population is not information. The Presto "ninety-five percent automated" figure was technically true and structurally misleading because "automated" excluded the offshore humans doing the work. Never record a number until you have asked what it counts and who it excludes.
Trace the humans and the third parties
Automation language hides people and sub-vendors. For any "it's automated" claim, trace a single transaction end to end and name every person and third party who touches it. A system sold as autonomous that secretly depends on people has a hidden cost, a hidden failure point, and a risk model built on a fiction.
A refusal is information, not a wall
When a vendor declines to answer or to commit, they have marked exactly where the truth would cost them. Weigh it: walk on load-bearing refusals (whose model, who does the work, will you warrant the number) that will not close; proceed narrowed on peripheral ones. Either way, the refusal goes in the record.
The contract is the product, not the conversation
A vendor will say almost anything in a meeting and warrant far less on paper. The single most important line in your record is the gap between what they claimed and what they will put in writing as a warranty, service level, audit right, or indemnity.
Classify every answer: confirm, dodge, or refuse
The pattern across answers matters as much as any single one. Confirms are load-bearing beams. Dodges become written follow-ups and, if repeated, findings. Refuses are located weak points. The classification is what turns a conversation into an analysis.
The interrogation resolves to proceed, proceed-narrowed, or walk
The most common good outcome is proceed-narrowed: the vendor is credible but safe only in a smaller scope than you planned, so you deploy to the subset the answers support and write the boundary into the contract and your monitoring.
The record is a link in a chain, written to be used later
The confirmed claims become your eval suite's test cases, the dodged claims become your monitoring alerts, and the refusals become the risks named in your board memo. Write the record so Topics 3.4, 4.2, 5.6, and 8.6 can pull from it. An interrogation you do not write down is a conversation you will misremember under pressure.
Match interrogation depth to stakes, on purpose
A low-risk internal tool needs a lighter pass than a system that decides things about your customers. But make the depth a conscious, recorded decision, not an accident of how much time you had that day.
The interrogation is your last cheap moment
Your leverage peaks before you sign and collapses after. Every material question you fail to force now becomes a question you have to litigate after an incident, when the vendor has your money and you have the switching cost. Force the definitions and the writing while asking is still free.
Refusing to move on is the whole skill
The clever question is worth little without the willingness to ask it a second and third time. Material claims break open on the follow-up, not the first ask, which is exactly how the Presto figure would have fallen. Stay on the denominator, the transaction trace, and the written warranty through the vendor's reframes; the answer that arrives only under pressure is the true one.
You own the deployment, so you must own the knowledge
When the bought system fails, it is your incident, your customers, and your regulator, not the vendor's. A contractual remedy pays you after the harm; it does not un-harm anyone. That is why the interrogation is worth real effort even for a reputable, fairly priced vendor: reputation and price do not transfer the deployment risk back, so the knowledge has to come with it.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 19 of the podcast.
Read the full conversation
So, picture this. You are sitting in a really well-appointed conference room. Or, you know, maybe just staring at a Zoom screen these days.
Right, exactly. And you are about to sign off on a massive AI vendor contract. It is a huge moment.
Oh, absolutely. The pressure is on. Yeah, and your operations director is sitting there practically vibrating with excitement because they just love this product.
You want it deployed yesterday. Exactly. The internal sponsor who championed the project has been pushing for months.
You've seen the highly polished sales deck. You've watched the pristine demo. And honestly, it looks totally flawed.
It always does in the demo. Right. It seems like the exact technological silver bullet your organization has been waiting for.
It's going to cut operational costs, streamline workflows, boost efficiency, all the buzzwords. The dream scenario. And the slide on the screen right now literally says 95% automated.
So, all you have to do is pick up the pen and sign. But before you do that. Yes, before you do, I want to pull you out of that conference room.
We're going to drop you straight into a fast food drive-thru lane in January of 2025. Which is, frankly, the perfect place to start because that drive-thru lane represents exactly what happens when you buy the slick slide deck instead of the actual operational reality. It really does.
So, let's look at January 2025. The U.S. Securities and Exchange Commission, the SEC, brings a really highly publicized case against a company called Presto Automation. Yeah, this is a massive deal in the governance world.
Right. So, let's look at what Presto was actually selling to investors in these massive restaurant chains. Picture a sales engineer at the front of a room.
They hit play on a video clip. A very carefully selected video clip, I'm sure. Oh, completely.
It's a customer pulling up to a drive-thru. They lean out the window and they speak this messy, complicated, highly specific order. Like changing their mind halfway through.
Exactly. They want no pickles, extra sauce, a completely different sized drink. And the environment is chaotic.
There's a soda machine hissing loudly, an ambulance siren in the distance. The worst case audio scenario. Right.
And the voice AI handles it perfectly. It reads the order back flawlessly. The slide in that sales pitch claims the system is 95% automated.
And the executives in the room, they just nod. Because why wouldn't they? The deal moves forward. Yeah.
I mean, the promise of 95% automation in a high turnover, low margin business like fast food, that is economically irresistible. It fundamentally changes the profitability of the entire restaurant. It does.
But let's look behind the curtain for a second. Because behind the scenes, from June to December of 2023, there were offshore human workers. Sitting in the Philippines and India.
Exactly. And they were manually entering roughly 70% of those orders by hand on the company's advanced pilot version. 70%.
Yeah. And if you go back to the original version of the system, human intervention was required in every single instance. 100% human in the loop.
This is why this is the defining case study for this entire field of governance and vendor management. We have to do a deep dive here. Right.
Because the stakes are huge. They are. Because here is the absolute essential point about that 95% figure.
And this is what trips up even the smartest executives. The SEC did not say that the 95% figure was a forged document. Wait, really? Really.
The vendor didn't just invent a number out of thin air and print it on a slide. The number was, according to a very specific internal definition, technically true. Hold on.
Let me get this straight. They didn't lie to the SEC or the buyers. They were manually typing in 70% of the orders, but the 95% number was technically true.
Yep. How does that math possibly work? Well, it works through the magic of definitions. They simply excluded all of those offshore workers from their internal definition of the word intervention.
Oh, wow. Yeah. When they said 95% automated, what they actually meant was that 95% of the time, the restaurant's own local staff inside the physical building didn't have to walk over to the screen and intervene.
So the offshore humans who were desperately typing the orders in real time, half a world away, were just quietly left out of the math. Exactly. They essentially defined automation as not your employees.
That is wild. I mean, if a publicly traded company can obscure 70% of its human labor behind a 95% automated metric and actually successfully pitch that to sophisticated investors and enterprise buyers, we have a massive problem. A huge problem.
It means the standard ways we evaluate enterprise software are entirely broken when it comes to AI. They are completely broken. Yeah.
Because the stakes are just, well, they're fundamentally different when you are buying an AI system compared to standard deterministic software. Like a database or a CRM. Right.
With a CRM, you know what it does. But when you make the decision to buy or wrap an AI system instead of building it from scratch internally, you are taking ultimate responsibility for a complex probabilistic piece of technology that you cannot fully see inside. It's a black box.
It really is. You didn't select the training data. You didn't break it in testing.
You literally don't know where its boundaries are. Which means we need a totally new way to interrogate these vendors before we sign. And that is exactly our mission for this deep dive.
Yes. We are going to arm you with a rigorous framework. Right.
We're taking an executive education level guide on vendor interrogation. Specifically, we're going to cover the seven families of questions that surface what an AI sales deck actively leaves out. Because right now, sitting at that desk before you sign, you are in a very specific window of time.
Explain that window. Well, you are in what we call your last cheap moment. I love that phrase.
It's crucial. Before you sign the contract, the vendor desperately wants your business. Their sales team is highly incentivized to answer your questions.
Your leverage as a buyer is at its absolute maximum. After you sign. The moment the money changes hands and you begin integration, your switching costs start climbing exponentially.
The exact same questions that we're completely free to ask today in the conference room. They will sound like hostile legal accusations next month when the system is failing in production. So this framework is about utilizing that last cheap moment before knowledge gap turns into a catastrophic organizational incident.
Exactly. So let's start with the foundation of this entire problem. Right.
Let's call this first spine concept by name. The deck is edited truth, not lies. This is so important to internalize if you want to effectively question a vendor.
It really is because it is rarely about outright fraud. Right. It's about curation.
A sales deck is a precision tool optimized to do exactly one thing. Close the deal. By design, it minimizes, softens or entirely redefines anything that introduces friction into the sales cycle.
Right. And there's a term for this, right? Yeah. The industry term for this specific failure mode is AI washing.
It's the practice of describing a product as more automated, more intelligent or more AI driven than it actually is an operational reality. And I imagine that exists on a pretty wide spectrum. A very wide spectrum.
So let's walk through that. On the harmless end of the spectrum, what does AI washing look like? On the totally honest, completely legal end, you have standard market speak. So a vendor takes an old school deterministic rules engine, a piece of software that just follows basic if then logic.
Nothing fancy. Right. Nothing fancy.
But they put a modern conversational interface on it and they call it agentic or AI driven, simply because that is the vocabulary the market currently expects and more importantly, funds. So it's just marketing. It might be slightly annoying to a purist, but it isn't malicious.
Exactly. But then on the far end of the spectrum, you have literal prosecuted deception. Like the Nate app.
Yes. Nate shopping app. Look at the U.S. Department of Justice and the SEC charging the founder.
They were pitching this seamless, autonomous AI checkout system that could navigate literally any e-commerce site on the web. It sounded like magic. It did.
But federal prosecutors alleged that the checkout process was not an advanced neural network at all. It was actually being run by human workers sitting in the Philippines and Romania. Just manually typing in credit card details.
Typing in credit card details and shipping addresses while the user just stared at a loading screen on their phone. Wow. I think it's easy for a buyer to hear stories like Nate or Presto and think, OK, the market is just full of scammers.
I just need to spot the liars. But that's a dangerous oversimplification. Right.
I want to look at the middle of that spectrum because that's where most enterprise buyers actually get hurt. It's not malice. It's more like looking at a nutrition label on a box of cookies.
I like this analogy. Yeah. Think about it.
You pick up the box and it says in massive bold letters, 90% fat free. Which sounds incredibly healthy. Right.
But if that claim is based on the physical weight of the food and the food has a super high water or sugar content, most of the actual calories you're consuming could still be coming straight from fat. So they aren't lying. No.
The cookie company isn't committing fraud. They're just picking the absolute most flattering denominator to show you on the box. That is the perfect analogy.
Yeah. They're picking a flattering denominator to highlight a specific feature while hiding the actual systemic impact on your health or, in our case, your business. And the danger is that even the most reputable, honest vendors do this.
Structurally, it is just how software is sold. But when you are buying AI, that flattering definition leaves you, the buyer, holding massive amounts of unpriced risk. Because at the end of the day, when you deploy that system to your customers, you own the deployment.
Precisely. If I buy an AI chatbot to handle customer service for, say, my airline, and that chatbot hallucinates a fake refund policy. Or worse, hurls a completely unprompted insult at a high-tier loyalty customer.
Right. That customer does not care who my third-party sauce vendor was. They don't know the vendor's name.
They only care that my airline just insulted them. The liability shift is absolute. It is.
If the system fails, it is your incident. It is your brand equity that evaporates overnight. It is your customer trust that is breached.
And it is your industry regulator knocking on your door, asking for an explanation. Now, a vendor might say, well, we have a contractual remedy. We'll pay you back some of your licensing fees if we mess up.
But a financial remedy pays you after the harm has occurred. It does not unharm your customer. It does not unpublish the viral screenshot on social media.
Exactly. Which means we can't rely on the deck. And we honestly can't just rely on the vendor's brand reputation.
We have to actively look for the specific things a sales deck will reliably hide. And there are four of them. Right.
Four specific things. The conditions of the demo, the definitions sitting underneath the headline metric, the people and third parties quietly operating in the loop, and the failures. Because no deck leads with its worst outputs.
No sales engineer boots up a PowerPoint to show you the time their model confidently fabricated a racist historical fact. Never going to happen. So if the danger lies in these flattering measurements and curated conditions, we need to look exactly at how those measurements are constructed.
If the problem is the nutrition label, we need to learn how to read the fine print on the back of the box. Which brings us to our second core spine concept. Every metric rides on a definition.
Get the definition first. A headline number in a sales presentation like 95% automated or 99% accurate or 85% autonomous resolution is entirely meaningless as a piece of information until you know the mechanics of how it was calculated. So what are the mechanics? What are we looking for? You need to identify three specific elements for every single metric.
First, you need the denominator. Meaning what is the total pool of data the percentage is measured against? Exactly. Second, you need the counting rule, which is the exact rigid boundary of what qualifies as a success versus a failure.
And third, you need the excluded population. Who or what was deliberately left out of the count entirely? Right. So let's apply this practically.
Our source material gives us seven families of interrogation questions, and the first two families tackle these metrics head on. Let's look at family one, the demo to deployment gap. This is the massive chasm between the highly curated, brightly lit environment where the demo is recorded and the messy, chaotic, completely unstructured reality of your actual real world production environment.
Because a demo is just a best case scenario dressed up in the clothes of a typical case. That's a great way to put it. Let me give you an example of how that gap destroys value.
Take the case of Babylon Health. The telehealth firm in the UK. Yes.
At one point, they were valued near $4 billion. They marketed an AI symptom checker heavily, and their demos were incredibly impressive. I remember seeing those.
They looked fantastic. But the reporting and medical scrutiny that came later indicated that its real world performance never seemed to match the massive promotional claims made in those controlled demonstrations. Why not? Because in a demo, the inputs are clean.
The user describes their symptoms in clear medical adjacent terminology. But in the real world, a panicked patient is typing messy slang, misspelling words, combining totally unrelated symptoms. Right.
The gap between the demo world and the real world was catastrophic. And the company eventually collapsed into administration. And it isn't just about the data being messy in the real world, is it? Sometimes the model they demo for you is literally structurally different from the one they sell you.
Yes. And this is a crucial mechanism for executives to understand today. There was a highly publicized case recently involving a public AI benchmark leaderboard.
Okay, what happened? A company submitted a specially tuned, hyper-optimized version of their model to the leaderboard. This version was explicitly trained to excel at the conversational patterns tested by that specific benchmark. And it scored vastly higher than the base model.
But the base model was the one you could actually buy and run in your own enterprise environment. Exactly. Same model name, totally different tuning and context window restrictions.
So it results in materially different answers. Right. If you buy the model based on the leaderboard score or the highly tuned demo, you are buying a ghost.
So what's the specific interrogation question for this family? You must ask, can we run our own hardest, messiest, most unstructured historical inputs through the exact production model we will be buying before we sign this contract? Not the demo model, not the leaderboard model. The production model running on the actual compute infrastructure we are paying for. That is a phenomenal question.
Our own hardest inputs. If they say no to that, that is a massive red flag. It's a deal breaker for me.
So that covers the demo gap. Let's look at the actual math of the metric itself. This is family two, the metric definition.
This is an area where we are seeing regulators actively stepping in because the definitions are so manipulated. Look at the 2024 Texas Attorney General settlement. This was over a healthcare AI vendor.
Yes. They settled an investigation over exactly this issue. The vendor was making representations to hospitals about a critical hallucination rate in a clinical safety critical setting.
Which sounds like an important metric. It is. But that rate means absolutely nothing until you know exactly how the vendor defined the word critical and how they actually counted it.
Right. If the AI tells a doctor a patient has a common cold instead of the flu, is that a hallucination? Yes. Is it a critical hallucination? Well, who decides that? The vendor's data labelers, usually.
And if they define critical as only hallucinations that result in, say, immediate fatality, their critical hallucination rate will look incredibly low on paper. But the model might still be incredibly dangerous for everyday diagnostics. Precisely.
Let's do a deep dive right now to show how easily a metric can be dismantled. Let's do it. We're back in our conference room.
The vendor is sitting across from us and they put up a slide that proudly declares 85% autonomous ticket resolution. As an operations director, I love that number. That sounds phenomenal.
Sounds like I can rethink my entire support staffing model, cut my tier one agents, and save millions. It is entirely designed to make you think exactly that. But let's crack the metric open using those three elements we discussed.
Denominator, counting rule, and exclusions. Watch how quickly that 85% dissolves. OK, let's start with the denominator.
You stop the presentation and you ask the vendor, 85% of what? And they would smoothly say, of your customer support tickets, of course. But you have to push further. Is it 85% of all incoming support tickets across all channels? Or is it 85% of only the specific tickets that the AI chose to attempt because it had a high confidence score? Oh, that is a massive distinction.
Right. If the AI looks at a complex billing issue, decides its confidence is low, and refuses to touch it, does that count against the 85%? Often, no. They exclude it from the denominator entirely.
So let's do the math on that. Let's say you have 100 tickets. And the bot only attempts 40% of your total incoming tickets because it only handles simple password resets.
OK, so it only touches 40 tickets. And it resolves 85% of those 40 tickets. It is actually only resolving 34% of your total ticket volume.
From 85% down to 34% just by forcing them to define the denominator. Yes. That alone completely changes the return on investment calculation.
I can't fire my tier 1 support staff if the bot is only handling 34% of the volume. Not without causing a massive backlog. But let's keep going.
What about the counting rule? How do they inflate the successes? Now you ask what exactly constitutes a resolution in your system? Good question. If a customer gets frustrated by the bot going in circles, they just abandon the chat, and the bot auto closes the ticket after 24 hours of inactivity, does your system's dashboard score that as a resolved ticket? Oh, wow. If they do, they are literally counting customer rage as a successful AI interaction.
Exactly. Or even more common, if the bot realizes halfway through the conversation that it can help the customer and it successfully routes the chat to a human agent, do you score that handoff as a success in the 85% metric? Wait, because technically the bot successfully performed the routing action? Yes. Wow.
If a handoff to a human counts as an autonomous resolution, the metric is actively measuring the exact opposite of what it implies. It's inflating the automation number with human labor. You thought you were buying an autonomous agent, but you are actually just buying an incredibly expensive routing switch.
Exactly. And what's fascinating is the headline number 85% did not change. The vendor didn't lie about the math.
But its meaning completely collapsed under the weight of two structural questions about the denominator and accounting rule. Which naturally leads us to the next massive problem. If the metric is artificially inflated because the AI is handing the hard tickets off to a human, or if a human is quietly stepping in to fix the transcription, like in the Presto drive-thru case.
We desperately need to know who those people are, where they are sitting, and how much they cost. Which takes us to our third core concept. Trace the humans and the third parties.
This covers the next two families of interrogation questions. Family three is the human-in-the-loop reality. Right.
You have to understand that enterprise automation, language words like seamless, agentic, autonomous, is specifically engineered to hide people. Think about the Amazon just walk out technology that was heavily deployed in their physical grocery store. Oh, I remember this.
You scan your app at the turnstile, you walk in, grab your groceries, put them in your bag, and you just walk out. No checkout lane. It felt like pure magic.
It felt like magic because the technology narrative told you it was magic. It was framed as an incredible triumph of computer vision and AI. But extensive reporting later indicated that the system relied very heavily on a massive workforce of humans.
Often sitting in data centers in other countries. Exactly. Manually reviewing the video feeds to verify exactly what people were picking up and putting in their carts.
So the AI is doing it, plus a room full of people thousands of miles away watching me buy cereal. Yes. And I want to be very clear here.
Human-in-the-loop is not automatically a bad architectural design. Right. A system with disciplined, well-trained human review can be vastly safer, more accurate, and more reliable than a fully autonomous neural network.
But the problem arises when a system is sold as autonomous while secretly depending on people. Because your capacity planning, your cost models, and your risk models as a business are all built on the fiction of infinite software scale. Okay, let me play devil's advocate for a second.
Go for it. Let's say I'm that operations director and I really want this deal to go through. Why do we actually care if there are hidden humans in the loop as long as the output is accurate and the price is what we agreed to? It's like if I buy a subscription to a self-driving delivery service.
I think it's all AI. But I find out later that a remote human driver takes over the steering wheel on the really hard, complicated routes. Okay.
The package still arrived on time. The price was fixed. Every enterprise system has fallback.
AWS goes down. We have backups. Why is an AI fallback suddenly a deal breaker? Why should I care how the sausage is made if it tastes good? You must care because undisclosed humans are unpriced operational risks.
Walk me through that. Let's look at a fictional scenario from our source material to illustrate exactly how this breaks your business. Yeah.
Imagine a company called Northwind Kitchens. Okay. They are buying a voice AI for their drive thrusts, very similar to the Presto situation.
The governance lead at Northwind is smart. So they ask the vendor, what happens if the vendor's remote human support team, the people cleaning up the messy orders, experiences an internet outage and goes offline during the Friday night dinner rush? A completely plausible scenario. Yeah, absolutely.
The vendor admits that if the remote humans disconnect, the system would fall back to the restaurant's own staff at the physical drive-through window. Okay. So the local restaurant staff suddenly has to put the headset back on and take the orders manually, just like they used to.
Yes. But now, look at the business case that justified the purchase in the first place. The operations director wanted to buy this AI specifically so they could cut the staffing headcount at the drive-through window.
That was the ROI. But if the vendor's offshore team goes down and the volume falls back to the local window, Northwind Kitchens cannot actually cut those staff members? Exactly. They have to keep them on payroll, standing around, just in case the vendor's secret human workforce drops offline.
The entire cost-saving premise of the deal was built on a fiction. That makes total sense. If it's pure software, it scales infinitely.
If it's humans disguised as software, it has a latency limit and a capacity limit. Precisely. A disclosed human is a design choice you can plan around, negotiate SLAs for, and build redundancy for.
A hidden human breaks your operation the moment they get overwhelmed. So what's the interrogation question we must ask for Family 3? You ask. Walk me through a single complex transaction end-to-end.
List every human who touches it, where they sit, what their latency limits are, and who legally employs them. And tracing the humans leads us directly to tracing the third-party technology. This is Family 4. Model provenance and ownership.
Right. You are about to heavily depend on a model to run a core part of your business. You need to find out whose model it actually is.
Because in the current AI ecosystem, we deal heavily with something called a wrapper. Exactly. A wrapper is a product where the core capability, the actual brain of the product, is provided by an upstream foundation model that the vendor does not own.
They just wrap their user interface in some prompt engineering around an API from OpenAI, Anthropic, or Google. Which means the vendor I'm signing a contract with doesn't actually control the core technology they're selling me. Precisely.
And this is completely different from traditional software supply chains. How so? Well, if a CRM vendor uses an open source database library, that library is static. It does its job.
It doesn't change on its own. But if an AI vendor is just a wrapper, their entire cognition engine is out of their control. Oh, right.
If that upstream provider decides to deprecate the specific model version your vendor relies on, your product changes behavior overnight. It might suddenly refuse to answer questions it answered perfectly yesterday. Or if the upstream provider drastically changes their API pricing, your vendor's entire business model might collapse and they go bankrupt.
Exactly. Going back to the Presto case, the SEC specifically noted that for a period of time, the core speech recognition technology deployed in those drive-thru units was actually owned and operated by a third party, not by Presto itself. Yes.
If you don't ask, is the core model yours or a third party's, and will you explicitly name every subprocessor in the chain, you are buying a black box supply chain. And supply chains carry enormous fragility. Okay, so we've cracked the metric, we've found the hidden offshore workers, and we've mapped the third party API dependencies.
We're doing good. Once we find all these hidden dependencies, we have to start asking what happens when those dependencies inevitably fail. Because they will fail.
This brings us to the final two foundational question families before we get into the psychology of analyzing the vendor's answers. Family five is failure and incident history. This is where we look for the silent failures.
Every system fails. There's no such thing as 100% uptime or 100% accuracy in software. And especially not in probabilistic AI.
Right. If a vendor implies their system is failure-free, they are telling you one of two things. Either their monitoring is so immature that they don't even know when they are failing, or they are actively unwilling to disclose what they know.
And in AI, the most dangerous type of failure is a silent failure. Let's break down the mechanics of a silent failure. Why is it unique to AI? It goes back to how large language models function.
They are next-token prediction engines. They do not have an internal database of truth and falsehood. They probabilistically determine the most likely next word.
Because of this, a silent failure occurs when the system produces a highly confident, completely fabricated output, rather than throwing an error message. Right. A traditional software system that throws an error 404 is manageable.
It stops, your IT team sees the error in the logs, and you handle it. But a system that fabricates a confident answer fails in a way that your customers see before you do. Give me a real-world example of this silent failure.
Let's look at the Cursor AI incident in early 2025. Cursor, the developer tool company. Yes, a highly respected company.
They deployed an AI customer support agent. And this agent confidently invented a completely non-existing company policy regarding refunds or usage limits. It just made it up.
Completely made it up, and it stated to real, paying users as absolute, unbending fact. The AI didn't stutter. It didn't throw an error code.
It just lied with perfect syntax. Customers got angry. Some canceled their subscriptions based on this hallucination.
And the vendor had to do massive public damage control. Or look at the Federal Trade Commission settlement with DonotPay in 2024. Oh, the robot lawyer.
Yeah. DonotPay aggressively marketed an AI service as the world's first robot lawyer, claiming it could replace the entire legal industry for basic tasks. But the FTC alleged they hadn't actually tested whether the AI's output matched a human lawyer's quality.
Exactly. They were making massive capability claims without an established testing history behind them. So for family five, you cannot simply ask the vendor, do you test your system extensively? Why not? Because that is a process question.
Every vendor on earth will say yes to a process question. Of course we test. You have to ask for outcomes.
What's the phrasing? You say, show me your absolute worst production outputs from the last 90 days. Walk me through this specific class of unstructured input that reliably causes your system to fail silently. And if they can't? If they cannot name their own failure modes, you end the meeting.
Wow. Okay. And that brings us to family six, data handling.
This is where things get legally perilous. Data questions are where a polished, smooth talking vendor will often get visibly, physically uncomfortable in the room. Because the honest answers to data questions often lose deals.
Exactly. You must ask, what exactly happens to the proprietary data that we and our customers send you? Are you using our data to train your future models? And worse, are you passing our data to your upstream provider so they can train their models? Right. And the risk here isn't just that our competitors get our data.
What happens if the data the vendor's model was originally trained on turns out to be a legal nightmare? That is the structural risk. We saw this play out in South Korea in 2021 with a highly publicized privacy case. What happened there? A company built a conversational AI chatbot and they trained it using deeply personal private messages that users had shared on a completely different service owned by the same company.
Without telling them. Right. The national privacy regulator stepped in and found that this reuse of data lacked valid, informed consent.
They had to destroy the data set. So if you buy a product built on poison data, you might inherit that legal exposure or your product might get shut down overnight by a regulator. So what is the specific interrogation tactic here? You have to ask if the data training off switch is legally bound in the text of the contract.
As opposed to what? As opposed to just being a toggle hidden in the user settings page that they can silently update in their terms of service next month. The discomfort you see in the room when you ask these data questions, that discomfort clusters exactly where their automation story is the thinnest. OK, we have built the foundation.
We have a first six families of questions. We know what to ask about the metrics, the humans, the models, the failures and the data. But having the questions is only half the battle.
Exactly. The real executive skill is in how we interpret the incredibly slick, highly rehearsed answers we are going to receive from a professional sales engineer who does this every single day. Which brings us to our next core concept.
Classify every answer, confirm, dodge or refuse. Because when you ask these hard structural questions, you are rarely going to get a clean, simple yes or no. Sales teams are trained to answer in a register that sounds highly responsive, enthusiastic and authoritative without actually committing to anything legally binding.
Therefore, as the interrogator, you must mentally classify every material answer into one of three buckets. Let's break those buckets down so you know exactly what to listen for. First bucket, confirm.
A confirm is exactly what it sounds like, is a specific, verifiable and ideally writable commitment. Give me an example. The production model we are deploying for you is version 4.2 and we will commit to giving you 30 days written notice before any upstream API change.
OK, that's clear. Confirms are checkable. They survive the journey from the verbal conference room conversation into the written legal contract.
These confirms are the load-bearing beams that will actually support your deployment decision. Got it. Second bucket, the dodge.
A dodge is a true-sounding non-answer. It is an absolute art form in enterprise sales. How so? The salesperson might answer a much narrower question than the one you actually asked.
Right. For example, you ask, does your system use offshore human labelers in real time? And they answer. They answer, our core AI model is 100% proprietary and developed in-house.
Ah, they answered a technology question when you asked a labor question. Exactly. Or they might substitute social proof for actual technical evidence.
You ask about their hallucination rate and they respond by naming three Fortune 500 banks that use their product. We work with JP Morgan or whatever. Which tells you nothing about the hallucination rate.
Or they might use a phrase with no verifiable referent whatsoever, like our model is state-of-the-art or we use military-grade encryption. What about legal dodges? I feel like we comply with all applicable laws is a phrase you hear constantly when you asked about data privacy. It is incredibly common and it is a classic dangerous dodge.
It sounds wonderfully reassuring, but it doesn't specify which jurisdictions laws they consider applicable or how they are complying. And more importantly, AI regulations like the EU AI Act or various state laws often place heavy duties on the deployer of the system. That is you.
Exactly. Their internal compliance does not automatically shield you from liability if the system discriminates against your customers. Now, I want to clarify something.
A dodge isn't necessarily proof of malice or a scam, right? No, not at all. It's like watching a skilled politician in a televised debate. The moderator asks a difficult question about the economy and the politician smoothly pivots to the talking point they want to discuss, like education.
They didn't lie. They just refused to engage with the danger zone. But in a vendor interrogation, recognizing a dodge means you convert that verbal evasion into a mandatory written follow-up.
That's exactly right. A dodge just means the question remains unanswered. It is a placeholder.
Are there red flags we should watch out for here? Yes. The ultimate tell across all of these interactions is if a vendor simply cannot name a single weakness of their own system. If they sit there and tell you it is flawless, that it handles everything perfectly and stales infinitely, you have a massive red flag.
A seasoned operator who has actually deployed a system in the messy real world can name its weak spots in one sentence. If they can. They're either wildly inexperienced and haven't seen a break yet, or they are entirely unwilling to be honest with you.
Both are major findings that should pause the deal. Which brings us to the third bucket and our next core concept. A refusal is information, not a wall.
Because sometimes if you press hard enough, they won't dodge. They will just flat out say, we can't share that information with you. And this is where beginners in governance often fail.
They hit a refusal. They feel awkward. They feel like the meeting failed, or they get discouraged and move on.
But an expert interrogator writes it down as a located weak point. Exactly. A refusal to answer a question or to make a commitment is extremely high value information.
The vendor has just put a giant X on the map, marking exactly where the operational truth is too expensive, too fragile, or too damaging for them to share. But surely not all refusals mean we flip the table, rip up the contract and walk out of the room, right? They have proprietary secrets to protect. Of course.
The key executive skill here is distinguishing between a load-bearing question and a peripheral question. What's a peripheral question? A peripheral question might be about their internal compensation structure for their developers. If they refuse that, fine.
And a load-bearing question. Something like, whose upstream foundation model is powering this? Or who exactly is doing the human review of our data? Or will you legally warrant that 95% accuracy number you put on the slide? If they refuse those questions, the fundamental foundation of the deal is cracked. Let me challenge that on the model provenance specifically.
Shouldn't a refusal to name an upstream model mean we automatically walk away? Well, if I'm buying standard enterprise software, like a CRM or an HR platform, that vendor doesn't tell me every single open-source code library they use to build their database. They hide their libraries as proprietary architecture. Why is AI different? Why is an AI refusal load-bearing? That is a great pushback, and it comes down to the fundamental nature of the dependency.
An invisible software library and a CRM usually performs a discrete static function, like sorting a list or compressing an image. If it breaks, they patch it. But in an AI product, particularly a wrapper, the entire capability, reasoning, and reliability of the product might rest on one single upstream foundation model.
The AI is the product. If OpenAI deprecates the specific model version your vendor relies on, your product changes behavior overnight. If they refuse on a load-bearing question, you must treat it as a confirmed active risk in your board memo until proven otherwise.
You don't necessarily walk out of the room immediately, but you weigh it incredibly heavily in the risk column. That makes total sense. So we have categorized our notes.
We have our confirms, our dodges, and our refusals. But the ultimate test of all three of those tuckets comes down to what happens when you slide a piece of paper across the table and hand them a pen. This is our next core concept and the seventh and final question family.
The contract is the product, not the conversation. This is about demanding evidence in writing. There is a massive structural difference between what a vendor will cheerfully claim while drinking your coffee in a conference room and what their general counsel will legally warrant on paper when their own revenue is on the line.
We call this the warranty gap. Let's define the specific contractual mechanisms here. We need to talk about warranty, service level, audit rate, and indemnity.
How do these apply to AI? Okay, a warranty is a binding contractual commitment that a specific claim is true backed by a legal remedy if it is proven false. And a service level. A service level or SLA is a measurable performance, standard-like guaranteed uptime, or a specific accuracy rate.
Audit rate. An audit right is your explicit contractual permission to test their system on your own inputs on a regular schedule to verify performance. And finally, indemnity.
That's their promise to cover your financial losses if their system causes a harm, like a copyright infringement claim, from their training data. I want to pause on indemnity for a second. If a small AI startup offers me full indemnity, am I actually protected? Often, no.
It is worth noting that AI insurance coverage in the broader market is rapidly narrowing. Insurers are excluding hallucination risks and copyright risks, right? Exactly. So an indemnity from a startup with six months of runway might be legally binding, but practically worthless if their insurer excludes the very AI output that caused the lawsuit.
You can't squeeze blood from a stone. This is where the interrogation gets incredibly sharp. Asking, will you put that exact phrasing in the contract is the ultimate truth serum.
It really is. Remember that 85% autonomous ticket resolution claim we deconstructed earlier? The one that shrank to 34% when we looked at the denominator and the counting rule? That original 85% claim evaporates into thin air the moment you ask the vendor to warrant a 34% resolution rate with a contractual financial penalty if they miss it. Exactly.
If we connect this to the bigger picture, this goes back to the concept of your last cheat moment. Your leverage is at its absolute peak before you sign. If the vendor dodges the contract question by saying, oh, we can sort out the specific service level metrics later during the implementation phase, you have to recognize what that really means.
In enterprise software, sorted out later reliably means concede it later. Once they have your money and the integration has begun, their incentive to agree to strict financial penalties drops to zero. Okay, so we've run the interrogation.
We've cracked the metrics, traced the humans, mapped the API dependencies, categorized everything into confirms, dodges, and refusals, and pushed for the written contract. What do we actually do with this written record on Monday morning? We need to make a decision. The interrogation must produce a written record.
An interrogation you do not write down is just a conversation you will misremember under pressure from your internal sponsor. That written record resolves into one of three operational decisions. Proceed, walk, or proceed narrowed.
Perceive means everything looks good, the confirms are strong, the risks are manageable, and they sign the warranties. Walk means the load-bearing questions were refused, the risk is unpriced, and we're out. But proceed narrowed.
You say this is the most common good outcome for an enterprise AI deployment. Yes. Proceed narrowed means the vendor is credible, they are acting in good faith, but your rigorous interrogation revealed that their system is only genuinely safe and reliable in a much narrower scope than their marketing slides suggested.
So you don't kill the deal? No, you don't kill it, but you deploy it only to the specific subset of use cases that the evidence actually supports. And you write that precise boundary into the contract and into your operational monitoring. Give me an example of proceed narrowed.
If the interrogation reveals that the voice AI handles the top 15 menu items flawlessly, but hallucinates terribly on custom dietary modifications, you proceed narrowed. You only let the AI take orders for those 15 standard items, and you build a routing rule that sends any utterance containing the word allergy or custom directly to a human immediately. You scope the deployment to the evidence.
To get to that outcome, the person running the interrogation, you listening right now, needs one specific psychological skill above all others, the willingness to refuse to move on. This is the hardest habit for beginners to build. Moving on feels polite.
Keeping the meeting flowing smoothly feels productive. You don't want to seem difficult. But the clever question is entirely useless without the willingness to ask it a second, a third, and a fourth time.
You must stay on the denominator question through all of the vendor's reframes, pivots, and smiles. The answer that arrives only under pressure after three dodges is the real operational truth. And this interrogation record is not just a one-time gate that you pass through and throw in a drawer.
It is a reusable chain of logic for your entire governance program. Exactly. The claims you tagged as confirms during the interrogation become the exact test cases for your engineering evaluation suite.
You test what they promised. The dodges become the alerts you program into your production monitoring system because those are the exact areas the vendor wouldn't reassure you about. And the refusals become the named documented risks in the memo you hand to your board of directors.
So if we want to give you the single most valuable move you can make this Monday morning, what is it? What is the immediate action item? On Monday morning, I want you to take the single most impressive, highest impact metric your current or prospective AI vendor is marketing to you. The one that made your operations director so excited. Do not accept it.
Email the vendor or bring it up in your very next meeting and ask this exact sequence of questions. What is the precise denominator for this number? Who or what is excluded from the count? And will you put that specific definition along with a service level remedy in our contract? And do not move on to the next agenda item until they answer it. Doing that is how you prevent your organization from becoming the next SEC headline for AI washing.
You are stopping the sales pitch and demanding the operational reality. Let's do a quick recap of the journey we just took. We learned that sales decks are carefully edited truths, not outright lies.
We learned that you have to crack the metric definition by finding the denominator and the counting rule. You have to trace the hidden humans and the third parties hiding behind the word autonomous. You must classify every answer into confirm, dodge, or refuse.
You have to use refusals as high value maps of weak points. And finally, you have to realize that the legally binding contract, not the friendly conversation, is the actual product you are buying. Before we wrap up, I want to leave you with a new angle to ponder, something we haven't touched on yet, but which builds directly on everything we've discussed today.
What's that? We've spent this entire deep dive treating the vendor as an outside company, an external third party trying to sell you something. Right. But what happens when the vendor is your own internal IT or data science team pitching an AI build to the C-suite? Oh, wow.
That changes the dynamic completely. Internal teams use polished slide decks, flattering metrics, and highly curated best case demos too. Their careers, their promotions, and their department budgets depend on the AI project getting approved by leadership.
So ask yourself, would your own company's internal AI pitch survive this exact same seven family interrogation? Or are they hiding human-in-the-loop fallback and API dependencies from the board? That is a phenomenal, terrifying question to end on. If your internal team can't name their own denominator, you are holding the exact same unpriced risk as if you bought it from a stranger. Thank you so much for joining us on this deep dive.
Stay curious, protect your organization, and always, always ask for the denominator.
Real cases
These examples show the interrogation lens applied to documented cases. Each is real and cited; the point in each is which hidden thing a good interrogation would have surfaced.
Example 1: Presto Automation, the metric that redefined "automated" (United States, 2025). The SEC's core finding is a case study in Family 2 and Family 3. Presto told investors Presto Voice "eliminat[ed] human order taking" and reported roughly ninety-five percent "automated order completion." The interrogation questions that break this: "ninety-five percent of what, and who is excluded from the count?" and "for one order, name every human who touches it and who employs them." The answers, per the SEC, would have been that "automated" excluded offshore agents in the Philippines and India who entered orders by hand (about seventy percent of orders on the more advanced pilot version from June to December 2023, and every order on the original version), and that the underlying speech-recognition model was, for a period, a third party's, not Presto's (Family 4). The number was defensible only under a definition no buyer would assume. (SEC, Presto Automation administrative proceeding, 2025.)
Example 2: Babylon Health, the demo that outran the deployment (United Kingdom, 2023). Babylon, a telehealth firm once valued near four billion US dollars, marketed an AI symptom checker whose real-world performance never matched its promotional claims; the company collapsed into administration and was sold for parts. The Family 1 question this case rewards: "can we run our own hardest cases through the system before signing, and what is the input most likely to make it perform worse than the demo?" A demo-to-deployment gap this wide is exactly what pre-signature testing on your own inputs is designed to catch. (TechCrunch, "The fall of Babylon," 2023.) (see Topic 3.1)
Example 3: The DoNotPay "robot lawyer," the capability claim with no testing behind it (United States, 2024). The FTC alleged that DoNotPay marketed an AI service as "the world's first robot lawyer" that could "replace the $200-billion-dollar legal industry," yet the company had not tested whether its output matched a human lawyer's and had not retained attorneys to check. The settlement required consumer notice and barred unsupported substitution claims. The Family 5 and Family 7 questions this case rewards: "what testing validates this capability claim, and will you warrant it in the contract?" A capability a vendor will not test and will not warrant is a capability that does not exist for your purposes. (FTC, "Operation AI Comply," 2024.) (see Topic 5.7)
Example 4: Amazon "Just Walk Out," the automation that ran on people (United States, 2024). Reporting indicated Amazon's cashierless checkout relied substantially on people reviewing video to verify purchases, and Amazon wound the system down from its stores. The Family 3 question this case rewards: "what percentage of transactions require a human, and what happens if that human workforce is unavailable?" The lesson is not that human review is wrong; it is that a system's true cost and risk model depend on knowing the humans are there. (Bloomberg, 2024.) (see Topic 0.2)
Example 5: The healthcare hallucination-rate metric (United States, 2024). A US state attorney general settled with a clinical AI vendor over how it had represented a safety-critical accuracy metric to hospitals. Without naming the vendor's specific numbers here (that treatment is owned by another topic), the transferable point is Family 2: a metric offered to a buyer in a high-stakes setting must be interrogated for its definition and measurement, because "critical hallucination rate" means nothing until you know how a critical hallucination was defined and counted. (Texas Attorney General, 2024.) (see Topic 4.2)
Example 6: The confident-wrong support bot (United States, 2025). An AI customer-support agent for a developer-tools company invented a nonexistent company policy and stated it to users as fact; customers canceled, and the vendor had to explain publicly how the system had failed. The Family 5 question this case rewards: "what class of input makes your system fail silently with a confident wrong answer rather than an error?" A vendor who cannot answer that has not characterized their own silent-failure surface, which becomes your reputational incident. (The Register, 2025.) (see Topic 3.4)
Example 7: The AI-washed checkout prosecuted as fraud (United States, 2025). Federal prosecutors and the SEC charged the founder of a shopping app whose "AI" checkout was, they alleged, run almost entirely by human workers offshore. It is the same shape as Presto: an autonomy claim with people underneath it. The transferable point is that Families 3 and 4, pressed hard and put in writing, are not paranoia; in these cases they were the exact facts a regulator later proved. (US DOJ, 2025.) (see Topic 4.1)
Example 8: The benchmark that flattered a model (2025). It was documented that a company submitted a specially tuned, conversation-optimized version of a model to a public leaderboard, where it scored well above the ordinary public weights of the same base model. No fraud is alleged; the point is narrower and directly relevant to Family 1. A demo or a benchmark score reflects a specific model version and configuration, and the version you can buy may not be the version that produced the number. The interrogation move is to require the exact production model and version in writing and to test on it, rather than trusting a figure produced under conditions you cannot reproduce. (TechCrunch, 2025.) (see Topic 1.4)
Example 9: The reused chat data behind a consumer AI (South Korea, 2021). A company built a chatbot on personal messages that users had shared for a different service, and the national privacy regulator found the reuse lacked valid consent. Without retelling the case that another topic owns, the transferable point is Family 6: a vendor's answer to "where did your training data come from and on what consent" can expose a legal problem you would inherit as the deployer. Ask it, and ask whether the vendor will indemnify you if the answer turns out to be wrong. (PIPC, the Personal Information Protection Commission, South Korea, 2021.) (see Topic 2.2)
Cross-example pattern. Read the nine examples together and one shape repeats: the claim is real, the definition or configuration behind it is favorable, and the unfavorable part (the humans, the third-party model, the demo conditions, the field misses) is the part left off the slide. This is why the interrogation is organized by families rather than by a list of clever questions. The families map onto the categories of hidden information, so working them in order guarantees you look everywhere the shape tends to hide. The alternative is asking three sharp questions about the one thing that happened to catch your eye, and missing the family that will actually produce your incident.
Where people go wrong
- "A confident, polished vendor is a trustworthy vendor." Polish is a sales skill, not evidence. The Presto and Nate cases involved companies confident enough to make their claims to investors and regulators. Confidence tells you about the salesperson, not the system. Interrogate the claim, not the delivery.
- "The headline metric is the fact I need." The headline metric is the fact the vendor chose to give you. A number without its denominator, counting rule, and excluded population is not information. "Ninety-five percent automated" collapsed the moment the SEC asked what "automated" excluded. Always get the definition before you record the number.
- "Human-in-the-loop means the vendor is being cautious and safe." Sometimes. But a system sold as autonomous that secretly depends on people is a hidden cost and a hidden single point of failure, not a safety feature. The problem is not the humans; it is the humans being undisclosed, so your cost model, capacity plan, and risk model are all built on a fiction.
- "If I ask hard questions, I will offend the vendor and lose the deal." A vendor worth signing expects a real buyer to interrogate the product; the good ones respect it. If hard, fair questions genuinely end the relationship, the vendor was protecting something, and you just learned what. The interrogation filters for exactly the vendors you want.
- "A refusal to answer means I should walk away." Not necessarily. A refusal is high-value information about where the truth is expensive to the vendor, but the right response is to weigh it: is it a load-bearing question (whose model, who does the work, will you warrant the number) or a peripheral one? Walk on load-bearing refusals that will not close; proceed narrowed on peripheral ones.
- "The verbal assurance in the meeting is what I'm buying." You are buying the contract. A vendor will say almost anything in a room and warrant far less on paper. The gap between the claim and the warranty is the real product. Convert every material claim into a "will you put that in writing" and record the answer.
- "Once I've interrogated the vendor, governance of this system is handled." The interrogation is the first link, not the whole chain. The confirmed claims still have to be tested by your eval suite (see Topic 4.2), the dodged claims still have to be watched by your monitoring (see Topic 3.4), and the refusals still have to be managed as named risks in your board memo (see Topic 8.6). The record exists so the chain can use it.
- "AI washing only happens at scam companies." It runs on a spectrum. Outright fraud sits at one end, but honest vendors also report genuine metrics under flattering definitions and describe systems in the most autonomous language the truth will bear. The interrogation is needed for the honest vendors too, because their softening is legal and still leaves you holding risk you did not price.
- "I should interrogate every vendor with the same intensity." No. Match the depth to the stakes. A low-risk internal productivity tool needs a lighter pass than a system that makes decisions about your customers or feeds a regulated process. But make the depth a conscious, recorded decision, not an accident of how much time you had.
- "The model in the demo is the model I will get." Not necessarily. A demo can run on a different model, a different version, or a more favorable configuration than the one your contract delivers, and the deck will not flag the difference. The same model name can produce materially different answers under different tuning and context. Require the exact production model, version, and configuration in writing, and run your pre-signature test on that, not on whatever the demo used.
- "A big-name upstream provider means the dependency is safe." A reputable upstream model is more reliable than an obscure one, but it does not remove the supply-chain risk. The upstream provider can still change the model, deprecate a version, change pricing, alter usage terms, or restrict access, and any of those flows straight through to your service because your product is a wrapper around theirs. Reputation reduces some failure probabilities; it does not transfer the dependency risk off your books. Name and warrant the upstream model regardless of the provider's size.
- "The vendor told me they comply with all applicable laws, so that box is checked." "We comply with applicable law" is a dodge, not a finding, because it does not say which laws, in which of your jurisdictions, or how. It also answers the wrong question. AI-specific rules increasingly place duties on the deployer, not just the provider, so a vendor's own legal compliance does not automatically cover the obligations that land on your organization once you put the system into production. Ask what the vendor's compliance claim actually covers, and separately confirm what it leaves for you to handle.
- "An indemnity clause means I am covered if the system causes harm." An indemnity is only as good as the vendor's ability and insurance to pay it, and AI-specific insurance coverage has been narrowing, with some insurers moving to exclude AI-output liabilities. A generous indemnity from a thinly capitalized vendor whose own policy excludes the exact harm is a comfort on paper and nothing behind it. Ask what the vendor's insurance actually covers, not just what the contract promises. (see Topic 8.3)
Questions people ask
- What is vendor interrogation?
- The structured process of questioning an AI vendor, before buying, to surface the information a sales deck omits: the conditions of the demo, the definition under each metric, the humans and third parties in the loop, the failure history, the data handling, and what the vendor will commit to in writing.
- What is sales deck?
- A vendor's presentation optimized to close a deal. It is not usually false, but it is selective: it shows the best-case demo, the most favorable metric, and the most autonomous-sounding description of the system, minimizing or omitting whatever would slow the deal down.
- What is AI washing?
- Describing a product as more automated, more intelligent, or more AI-driven than it actually is. It runs on a spectrum from calling a rules engine "AI" to meet market expectations, through reporting a genuine metric under a flattering definition, to outright deception that regulators prosecute (as in the Presto Automation case, SEC, 2025).
- What is metric definition?
- The denominator, counting rule, and excluded population that give a headline number its actual meaning. "Ninety-five percent automated" is meaningless until you know ninety-five percent of what, what counts as a success, and who is left out of the count.
- What is denominator?
- The base a percentage is measured against. Changing the denominator changes what a metric means without changing the number. A central target of any metric-definition question. More on Denominator
Keep going
This lesson builds AI vendor due diligence and third-party risk, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.