Build, buy, or wrap: the decision framework with your organization's real budget
The short answer
Three options, one accountability
Build means you own the machine, buy means you rent a finished product, wrap means you own a thin app around a rented general-purpose model. The option changes what you build and what you rent. It never changes who answers for the outcome: you do.
What you will be able to do
- Distinguish the three ways to get an AI capability, build, buy, and wrap, and state in one sentence what your organization owns and what it rents in each.
- Analyze a real capability your organization needs and recommend build, buy, or wrap using explicit decision axes, not gut feel or vendor charisma.
- Estimate the total cost of ownership (TCO) of each option, the full lifetime cost including the run cost, the maintenance cost, and the supervision cost, not just the sticker price.
- Name the supervision tax (the recurring human-oversight cost) as a required line item, and explain why removing it, as Michigan did, is what turns a cheap build into an expensive disaster.
- Weigh control against speed, differentiation against commodity, and lock-in against convenience, and defend the tradeoff you chose.
- Locate where liability lands in each option, so you know before you sign whether a failure is yours to answer for, the vendor's, or split. (see Topic 8.4)
- Classify the decision as reversible or hard to reverse, and favor the reversible path when the evidence is thin. (see Topic 0.4)
- Choose the shape of human oversight the decision needs, human-in-the-loop or human-on-the-loop, and price it, so the supervision cost matches the stakes rather than the option you preferred.
- Audit a build, buy, or wrap proposal that someone else hands you, finding the missing cost buckets, the over-weighted axis, the assumed volume, and the absent exit before you approve it.
- Produce a one to two page decision memo that a skeptical finance officer and a skeptical regulator could both read and find defensible.
The lesson
When your organization needs a new AI capability, you have to decide exactly how you're going to get it. That choice comes down to three distinct architectural paths. Build, buy, or wrap.
If you choose to build, you are choosing to own the machine end-to-end. You develop the model, write the code, and manage the data pipeline entirely in-house. If you buy, you are renting a finished product from a vendor.
You get a working solution immediately, but you own none of the underlying machinery. It operates entirely as a black box. And if you wrap, you build a thin application layer around a general-purpose AI model trained by someone else.
You own the prompt layer, but it sits on top of a rented brain. Most leaders make this architectural decision by looking at a single number, the upfront sticker price to get the system deployed. That is an illusion.
To evaluate these options, you have to price the entire life of the system, not just the sticker. You calculate the compounding costs over a multi-year horizon, including every dollar spent to keep it running and accurate. No matter which of these three paths you take, there is one constant.
The organization deploying the AI under its own name is the entity held accountable for the outcomes. Prioritizing initial speed and a cheap invoice over that long-term accountability bakes a catastrophic flaw into your architecture before you even flip the switch to go live. In 2011, the state of Michigan signed a $47 million contract for an automated unemployment system called MIDUS.
The goal was to process claims faster and cut operational costs. The state configured the new system to flag suspected fraud without a human adjudicator reviewing it first. Every human removed from that loop represented a salary the state could save immediately.
Between 2013 and 2015, the automated system went to work and falsely accused roughly 40,000 citizens of unemployment fraud. When the Michigan Auditor General eventually reviewed 22,000 of those fraud determinations, the findings were stark. About 93% of the flags did not actually involve any fraud.
The state aggressively pursued those false flags, seizing tax refunds and garnishing wages. People suffered extreme financial distress. Some filed for bankruptcy and others lost their homes entirely.
Michigan faced years of class action litigation, ultimately settling the Bouserman lawsuit in 2022 for $20 million. Every AI system has a supervision tax. You either pay it up front in reviewer salaries to ensure the system is safe, or you refuse to pay it and face the much larger bill in massive liability and ruined lives later.
To evaluate the lifespan cost of an AI capability, we use the total cost of ownership matrix. This matrix breaks the lifetime cost into four mandatory buckets, acquire, run, maintain, and supervise. Let's look at the buy option.
The initial acquire cost looks tiny, just a subscription fee and some integration. But that subscription compounds year over year, creating a massive run cost for a black box you cannot control. With a wrap, the engineering effort to acquire it is extremely low.
However, you pay a per-unit API bill that scales aggressively with usage volume, plus a heavy maintenance cost when the underlying provider inevitably updates or deprecates their model. The fourth bucket, supervise, is the recurring human cost of watching the system and auditing its decisions. Notice this row on the matrix.
For any high-stakes capability, the supervise bucket is identical and massive across all three paths. The supervision tax is strictly dictated by the real-world stakes of the decision the AI makes. Buying a vendor's product or wrapping a rented model does not give you permission to stop watching it.
Any proposal that attempts to show a cheap AI deployment by zeroing out the human reviewers isn't actually cheap. It is unpriced risk masquerading as a savings plan. To strip emotion from the acquisition debate and break deadlocks between engineering teams and finance, you run your capability through the six-access scorecard.
Access 1 is differentiation. You only justify the extreme cost and responsibility of a build if this specific capability creates a unique competitive moat that your customers actively choose you for. Access 2 is control.
Wrapping a model gives you total control over the interface, but zero control over the core behavior of the foundation model sitting underneath it. Access 3 is speed. A looming deadline is real business pressure, but speed should only be used to break ties between safe options.
It should never override fundamental safety checks. Access 4 is lock-in. An auto-renewing contract or a tightly coupled model version becomes a one-way door, making it incredibly painful to leave if the AI's performance begins to drift.
Access 5 is data sensitivity. If you are handling highly protected personal records, strict privacy regulations will often legally force you to build or self-host your wrap, regardless of what your budget prefers. And Access 6 is liability.
Renting a vendor's model outsources the compute power, but it never outsources the blame for failures. The burden lands entirely on the deploying organization. Often, these axes conflict.
If you need a cheap commodity capability, but it carries high stakes consequences for your users, you have a tension in the scorecard. You must explicitly acknowledge that conflict and fund the supervision required. Running these six axes in strict order forces your team to confront the architectural trade-offs you are making, long before anyone signs a contract.
Now we map your organization's specific risk profile against the three acquisition paths using the completed scorecard. In regulated industries, where AI decisions directly impact lives or finances, you must choose to build or run a self-hosted wrap. This requires the highest upfront total cost of ownership, and strictly implementing a human-in-the-loop to review every consequential output.
If the task is a commodity capability, like summarizing internal documents, you need it quickly and the risk is low. You branch toward a buy or a wrap. For these profiles, skip reinventing the wheel.
You can safely accept vendor lock-in and utilize lighter human-on-the-loop monitoring, randomly sampling outputs instead of reviewing every single one. When your organization wants to try a new feature, but lacks the hard evidence to justify a massive build, you need a reversible path. Using a wrap serves as a two-way door.
It allows your teams to test features rapidly in production, with high reversibility if the pilot doesn't deliver the value you expected. Mapping these profiles proves that AI acquisition is never just a one-size-fits-all IT purchase. It is a highly tailored architectural commitment to a specific level of risk.
The central trade-off that defines this entire process is simple. Renting the technology does not mean you rent the liability. You have to accurately price the entire lifespan of the AI by running the axes, and you can never justify a deployment by zeroing out the humans who ensure its safety.
Ultimately, deciding how to acquire an AI capability defines exactly who will be held responsible when it inevitably makes a mistake. Make sure you are prepared to answer for the path you choose.
The ideas, one by one
Price the life, not the sticker
Total cost of ownership has four buckets: Acquire, Run, Maintain, and Supervise. The sticker is only Acquire, usually the smallest over a system's life. The decision lives in the other three, and most disasters live in the last one.
The supervision tax is set by stakes, not by acquisition route
A consequential decision about a person needs human oversight whether you built, bought, or wrapped. Choosing to buy or wrap changes who supervises, never whether supervision is required. Any low TCO that leaves Supervise near zero for a high-stakes system is not cheap, it is unpriced risk.
Michigan proves the tax comes due either way
MiDAS was a 47 million dollar build whose entire saving was deleting the human adjudicators; about 93 percent of the fraud determinations the auditor checked were wrong, roughly 40,000 people were falsely accused, and the state settled for 20 million dollars (Michigan sources, 2016 to 2022). The oversight you refuse to pay for as salary returns as harm and liability, larger.
Run the six axes in order
Differentiation, control, speed, lock-in, data sensitivity, liability. Differentiation, control, and liability push high-stakes crown-jewel work toward build; speed and commodity push everyday work toward buy or wrap. When the axes conflict, say so and defend the tradeoff.
Speed is one axis, not the decision
"We need it now" is the sentence most often used to skip the other five axes. Speed breaks ties among safe options; it rarely wins the argument alone.
Price the exit before you enter
Every option locks you into something: a vendor, a model version, or your own maintenance burden. Ask what it costs to leave in two years and whether leaving is possible. An option you cannot exit is a decision you cannot reverse. (see Topic 0.4)
Renting the model does not rent the blame
In a wrap, you are the party the public and the regulator hold responsible, because you deployed it under your name. The comfortable myth that the model provider absorbs liability is exactly the myth that ends in a settlement. (see Topic 8.4)
The labels blur; price what you actually get
A "custom model" can be a wrap in disguise, a cheap buy can hide an expensive integration, and a wrap's model can change under you. Ask what you own, what you are locked into, and what happens when the underlying model changes.
The memo is the artifact, and it will be attacked
Your build, buy, or wrap decision memo feeds the vendor interrogation and the shipping plan, and a hostile board pulls it out in Module 11. Write it today as a first-class record a CFO and a regulator could both read, not a throwaway.
Demo accuracy is not deployment accuracy
The performance the vendor showed was measured on the vendor's chosen data. On your real population it will almost always be lower, and the gap is the exact size of the supervision you must fund to catch the errors the demo hid. Price the gap, do not admire the demo. (see Topic 4.6)
Audit the proposal, do not just approve it
When someone hands you a build, buy, or wrap recommendation, find the missing buckets, the over-weighted axis, the assumed volume and stakes, and the absent exit. The governance skill is taking a proposal apart, not signing the flattering slice its author happened to see.
Price over the same horizon, or the comparison lies
A build front-loads cost; a buy or wrap compounds it. Over one year the build looks dearest; over five the subscriptions overtake it. State the horizon you will actually run the capability and price all three options on it, or you will reach whichever answer you started with.
The answer can be all three
A capability is often several components, and the defensible call is frequently to build the one part that is your differentiator, buy the commodities around it, and wrap the flexible layer. Decompose the capability and run the framework per component, so build budget goes only where it is earned.
Match the oversight shape to the harm
Human-in-the-loop review on every output is right when a single mistake is severe and irreversible, and wasteful when it is not. Spend the expensive review where the harm is; use lighter on-the-loop monitoring where errors are visible and recoverable.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 18 of the podcast.
Read the full conversation
Imagine checking your mail on, like, a random Tuesday. Right. You're expecting nothing more than, you know, utility bills or flyers.
Right. But instead, you open a letter from your state government stating that you are guilty of unemployment fraud. Just out of nowhere.
Completely out of nowhere. And before you can even pick up the phone to figure out what kind of clerical error this is, you check your bank account. And you realize the state has already seized your tax refund.
Oh, wow. Yeah. And then the next week, your wages are garnished.
So you try to appeal, you try to call a caseworker, but there is literally no human to talk to. Just an automated system. Exactly.
A machine flagged your file, a machine ruled against you, it issued the penalty, and a machine took your money. And, you know, this isn't some dystopian science fiction scenario. No, it's not.
This is exactly what happened to roughly 40,000 people in the state of Michigan between 2013 and 2015. Yeah. And it is honestly one of the most chilling, well-documented failures of automated decision-making in modern history.
It really is. Because the state of Michigan, they signed a $47 million contract with a technology vendor, Hess Enterprises, back in 2011. Right.
And the whole objective was to replace their aging, you know, mainframe-based unemployment system with a modernized platform. They called it MIDUS, which stands for the Michigan Integrated Data Automated System. And on paper, I mean, this was pitched as a massive leap forward, right? Oh, absolutely.
Yeah. It was all about faster processing, modernized databases, and, you know, this shiny automated system to catch fraudulent claims. But the architecture of that implementation is really where the nightmare started.
Because they didn't just automate the filing process, did they? No, they didn't. They automated the actual judgment. Right.
They configured MIDUS to ingest data, flag suspected fraud, and, well, here's the really critical part issue, a binding legal determination without a single human adjudicator ever reviewing the case. Which is just wild to think about. And that decision, I mean, that was entirely driven by budget, wasn't it? 100%.
The state was looking at the massive cost of building and deploying the software, and they needed to justify the return on investment to the taxpayers. So how do you do that? You eliminate the ongoing operational cost. Exactly.
Every human adjudicator left in the loop was a salary that the new automated system was supposed to save. They basically ran the system completely blind. Under the assumption that the machine's, you know, state-matching logic and statistical thresholds were just infallible.
Right. They trusted the math over human oversight. I want to look at the mechanical failure here, because this wasn't just a slight miscalculation.
No, not at all. When the Michigan Auditor General finally went in and pulled a sample, I think it was 22,000 of these automated fraud determinations, just to see what the machine was actually doing, do you know what the error rate was? It was staggering. 93%.
Yeah. 93% of the people penalized had committed absolutely no fraud whatsoever. It's hard to even wrap your head around that number.
So how does a $47 million system fail that spectacularly? Like, what was it actually misreading? Well, it was a catastrophic failure of data pipelines and really rigid logic. You see, the system was pulling historical data, cross-referencing it with employer records, and flagging literally any discrepancy as intentional fraud. Wait, any discrepancy? Any discrepancy.
So if an employer reported a slightly different date of termination than the employee, or if there was, like, a minor typo in a wage report, the system didn't flag it for review to ask, you know, is this a clerical error? It just assumed the worst. Exactly. It instantly defaulted to a rigid classification of intentional malicious fraud.
Because in human systems, a caseworker uses semantic reasoning, right? Precisely. A human looks at a one-day date discrepancy and dismisses it as a typo. Like, oh, they put the fourth instead of the fifth.
But the machine lacked semantic reasoning entirely. It just performed brute force matching. Right.
It found a mismatch, and it just executed the penalty algorithm without a second thought. And the state did not pull its punches on those penalties either. No, they went after people hard.
They imposed quadruple financial penalties on these citizens. People lost their homes. People were pushed into bankruptcy over this.
It's heartbreaking, and it took years of really brutal litigation to fix. Finally, Bowserman v. Unemployment Insurance Agency became this massive class-action lawsuit. Right.
And in October 2022, the state agreed to a $20 million settlement, which was actually only just finalized recently, in January 2024. So that's a full decade of financial ruin for tens of thousands of people. Which brings us to the absolute core of what we need to dive into today for our listeners.
Yeah, let's pivot to the practical application of this. Because when we talk about artificial intelligence and automated logic in a corporate or government context, we are talking about software-making, or at least recommending, decisions that we use to hand to trained human beings. Whether that's deciding if a citizen committed fraud or deciding if a job candidate gets an interview.
Or even deciding if a patient is flagged for medical risk. Exactly. The stakes are incredibly high.
So if you are an executive or an operations manager or maybe a board member staring down a proposal to deploy an AI capability this week, this deep dive is your master class. We are going to deconstruct exactly how to acquire an AI capability without walking your own organization into a multi-million dollar disaster like Michigan did. And this entire discussion really hinges on a single fundamental paradigm shift regarding how you calculate the cost of these systems, right? That's right.
The core thesis here is this. The decision is never just about how you get the AI. It is about how you price its entire life.
Price its entire life. I like that. Yeah.
Michigan's failure wasn't simply that they chose to build a bespoke system instead of buying an off-the-shelf one. Their real failure was pricing the wrong number entirely. They priced the cost to build the software minus the salaries of the humans they fired.
Exactly. What they completely failed to add to that equation was the astronomical cost of the harm caused when the system was inevitably wrong and no human was there to catch it. So to avoid a Michigan-level catastrophe, we need a really precise decision framework.
And that starts with understanding what is actually on the table when someone says, you know, we need AI for this. Right. Let's lay out the options.
Let's break down the procurement avenues. Yeah. From my reading of the source material, there are really only three ways to acquire an AI capability.
Build, buy, or wrap. That's it. Just those three.
But I want to get extremely specific about the actual mechanics of each one. Okay. So we define these three options by asking one strict question.
What do you own and what do you rent? Okay. Let's take them in order. Option one is build.
Right. With a build, you develop the system yourself or you pay a contractor to build it to your exact specifications. Michigan's Midas was a build.
Got it. In a modern AI context, a build means you own the source code, you own the data pipelines, and crucially, you own the model weights. You are literally the manufacturer.
Hold on. Let's stop right there. Because I sit in vendor meetings all the time where people just throw around this phrase, we will own the model weights.
It's a big buzzword right now. It is. But what are they actually holding in their hands? Is it just standard code? Because they make it sound like this mystical asset.
That is a great question because we really need to demystify it. Model weights are not standard human readable logic code. It's not like a script that says, if X, then Y. Okay.
So what is it? When you train an AI, especially like a neural network, you feed it massive amounts of data. The system learns the relationships between different pieces of data by adjusting numerical values. And those values act as multipliers.
And those multipliers are the weights. Exactly. Physically, if you own the model weights, you own a massive file, often gigabytes or even terabytes in size, containing millions or billions of these highly specific decimal numbers.
They represent the learned experience of the model. If you own them, you have total mathematical control over how the system processes information. Got it.
So building gives you the ultimate control. Every single brick in the foundation is yours. Yes.
But I imagine that also means every single crack in the foundation is yours to fix. Precisely. Build gives you maximum control and maximum competitive differentiation, but it hands you the highest upfront cost, a massive ongoing maintenance burden, and total unmitigated responsibility for the outcomes.
Okay. So option two is buy. This is your standard software as a service or SaaS, right? Right.
You purchase a finished product from a vendor. They host it on their servers. You pay a monthly or annual subscription, and your team just logs in and uses it.
You are essentially a tenant. And buy is almost always the fastest path to a working capability, I'd assume. It is, but it requires a massive concession, which is you understand the least about what you just deployed.
Because it's a black box. Exactly. You do not own the model weights.
You usually cannot even see them. You are living entirely under the vendor's rules. So if they change the underlying model, update the interface, or just double the subscription price, you just have to absorb that disruption.
Yeah. You have no choice. Okay.
And the third option is wrap. And this has just exploded in popularity recently. Everywhere.
It's the most common approach right now. So you build a very thin application layer, like your custom prompts, some safety guard rails, your branded user interface, and you wrap that around a general purpose AI model, a GPAI, that some massive tech company has already trained. Right.
And you connect your thin layer to their massive model via an API. And an API, an application programming interface, is essentially just a paid doorway over the internet. Okay.
Walk me through how that actually looks in practice. Sure. Your software sends a packet of data, say a customer's email, through the doorway.
The giant model processes it on their servers and sends the summarized answer back through the doorway to your app. And you pay a fraction of a cent for that transaction. Exactly.
So in a wrap, you own your thin application layer, but you are entirely running the intelligence underneath. You are essentially building on rented land. Let's use a real estate analogy to really lock this in for everyone.
I like that. So build is like designing and constructing your own custom home from the dirt up. It takes years, it costs an absolute fortune, but you know exactly where the plumbing is, and you can add a new room whenever you want.
Perfect analogy. Buy is like signing a lease on a luxury apartment in a high rise. It's fully furnished.
It's ready today. But if the landlord paints the lobby a color you absolutely hate, or the pipes in the wall start leaking, you are totally at their mercy. Right.
You can't go breaking down the walls. Exactly. And wrap is like building a really beautiful bespoke pop-up shop on a piece of land owned by a massive corporation.
You own the shop itself, but the landlord might suddenly quadruple the ground rent or change the soil composition under your foundation, or just decide they don't want pop-up shops on their property anymore. The real estate analogy works perfectly to visualize the operational control, but it actually falls short when we talk about the legal reality, which brings us to the first major principle of our framework. Okay, hit me.
Three options. One, accountability. Now, I want to push back on this idea of accountability right away.
Go for it. Let's say I choose to buy that sauce apartment. I sign a massive enterprise contract with a major vendor.
And that contract includes a strict service level agreement and SLA. I'm paying them millions of dollars a year precisely to transfer the risk to them. Right.
So if their software miscalculates a mortgage rate and discriminates against an applicant, how does that liability not sit entirely with the vendor who built the flawed math? Because of the legal and regulatory concept of the deployer. And listen closely, because this is the single most dangerous assumption an executive can milk right now. Okay.
You read an SLA and you think you bought an indemnification shield. You didn't. If you read the fine print of that contract, vendors work incredibly hard to guarantee uptime, not outcomes.
Uptime, not outcomes. Exactly. They guarantee the software will turn on and be available.
They do not guarantee it won't break the law when applied to your specific customers in your specific context. So they're basically saying the hammer works perfectly. If you hit someone in the head with it, well, that's your fault.
That's exactly it. When you deploy a system to make decisions about your customers, your employees or your citizens, you are the deployer. Renting the model does not rent the blame.
Oh, wow. Think back to Michigan. They could not turn around to those 40,000 people they falsely accused and say, hey, don't look at us.
Fast Enterprises wrote the bad code. Right. Because the state of Michigan deployed it.
Yes. And the state of Michigan answered for it in federal court. If your bought fraud detector freezes a customer's account wrongly, that customer does not care who your sauce vendor is.
They have absolutely no relationship with your vendor. They are suing you. Three options.
One, accountability. Exactly. That is a very sobering reality check.
Yeah. You cannot outsource the liability. So if we accept that the accountability stays on our balance sheet no matter what, we have to accurately forecast what it will cost to maintain that accountability over time.
We need to talk about the real budget. Right. And the framework calls this price the life, not the sticker.
Yes. Most organizations fail spectacularly at this because their procurement processes are designed to buy static assets. You know, like office chairs or server racks.
Right. You look at the sticker price to acquire the asset. You get your budget approval and you move on.
But an AI system is not a static asset. The total cost of ownership, the TCO, is what it costs to keep the system working safely and accurately for its entire lifespan. So let's break down this TCO framework.
There are four distinct buckets of cost. Let's cover the first three starting with bucket one, acquire. Okay.
This is the sticker price we just talked about. If you are doing a build, this is the cost of your engineering team. But from reading the source material, it seems the massive hidden cost here is actually data acquisition and labeling.
Oh, huge. Why is that so expensive? Because models don't learn by magic, right? They learn by example. So if you want to build a model that reads medical scans and identifies tumors, you can't just give it raw images.
It doesn't know what it's looking at. Exactly. You need highly paid human experts, actual radiologists, to sit down, look at thousands of images, draw precise bounding boxes around the tumors, and label them.
Wow, that sounds incredibly tedious and expensive. It is. The cost of acquiring clean data and paying domain experts to label it so the machine can actually learn from it routinely eclipses the cost of writing the code itself.
Okay, so that's build. If you buy, the acquire bucket is the license fee plus the custom integration cost to make their software talk to your archaic legacy databases. Right.
And if you wrap, the acquire cost is the engineering time to build your application layer and wire up those API connections. Spot on. And here is the trap for executives.
Acquire is usually the smallest cost over the system's total life, even though it is almost always the only number prominently presented on the vendor's proposal. Right, it's the tip of the iceberg. Exactly.
Okay, bucket two is run. This is your daily operational cost. For a build, you're paying for your own compute power, the cloud hosting, the server electricity.
For a buy, it's the ongoing SaaS subscription. And for a wrap, it is the API bill. And I want to dig into this API bill because the source material flags this as a massive blind spot for companies when they transition from a pilot program to full production.
It's a huge pitfall. How exactly does an API bill catch you off guard? Because you pay per unit of computation, which is usually measured in tokens, which roughly equate to parts of words. Okay.
So when you run a pilot, you might have, say, 50 employees testing the wrap internally. They send a few queries a day. The API bill at the end of the month is like $40.
Right. And management looks at it and says, wow, this is incredibly cheap. Let's roll it out to our entire customer base.
Exactly. And your customer base is what? Millions of people? Right. Millions.
Suddenly, instead of 50 employees asking simple questions, you have millions of customers feeding massive multi-page documents into the API 24 hours a day. Your $40 bill doesn't just scale linearly. It explodes exponentially based on the length of the prompts and the complexity of the outputs.
A wrap that looks like a literal rounding error in a controlled pilot can completely bankrupt your IT budget at production volume if you haven't modeled the math correctly. That is terrifying. Okay, bucket three is maintain.
Yes. Now, this is where I really want to push back hard. I understand that traditional software needs maintenance, you know, security patches, bug fixes.
Sure. But if I spend $5 million to build a perfectly trained, highly accurate AI model today, why do I need a massive maintenance budget tomorrow? The math doesn't change. Like, a calculator from 1990 still knows that 2 plus 2 is 4. Can't we just set the AI and forget it? You absolutely cannot set it and forget it.
And your calculator analogy perfectly highlights exactly why. Okay, tell me. A calculator operates on absolute unchanging rules.
Math is static. AI operates on statistical distributions of the world at the exact moment it was trained. And the world moves.
The specific term for this in data science is drift. An unmanaged drift is a documented cause of silent, catastrophic degradation in AI systems. Walk me through the mechanics of drift.
How does a model silently degrade? Think of a deployed model like a highly detailed map of a city. If you print a map today, it is perfectly accurate. It reflects reality.
But over the next year, reality changes. Right. They close a bridge.
Exactly. The city closes a major bridge for repairs. A new subdivision gets built.
Traffic patterns reverse on downtown streets. The ink on your printed map hasn't changed at all. It still looks incredibly authoritative and precise.
But it is now silently wrong because it's describing a world that no longer exists. Ah, I see. So if I apply that to data, say I build an AI to detect fraudulent credit card transactions.
Perfect example. I train it on all the fraud data from 2023. It learns exactly what 2023 fraud looks like.
But in 2024, the criminals change their tactics. They start routing small transactions through different countries before making a big purchase. My model is still perfectly looking for the 2023 patterns.
Exactly. The mathematical relationships, those model weights we talked about earlier, are locked in. When the real-world feature distributions shift, the model's accuracy plummets.
This is concept drift. So what does maintenance look like for a build then? Maintenance for a build means paying your data scientists to constantly monitor the incoming real-world data, manually label the new fraud tactics, and expensively retrain the model to update the map. Okay, but what if I buy or wrap? Doesn't the vendor handle the maintenance? That seems like a massive point in favor of renting.
It is a trade-off. If you buy a SaaS product, yes, the vendors' engineers are the ones fighting drift. They are updating the core model.
That is a genuine cost saving. But you pay a different kind of maintenance tax. Which is? You live under their update schedule.
If their update unexpectedly breaks your internal workflow or changes the interface your employees rely on every day, managing that chaos internally is your maintenance cost. That makes sense. And for a wrap? In a wrap, the model provider is constantly tweaking, updating, or deprecating the underlying general purpose AI.
When they release a new version of the model, your carefully engineered prompts, the ones that worked perfectly yesterday, might suddenly start producing bizarre or incorrectly formatted outputs today. So maintenance in a wrap means constantly scrambling to update your application layer every single time the landlord changes the soil underneath your pop-up shop. Precisely.
It sounds like when you're analyzing these three buckets, acquire, run, and maintain, the most critical mistake is comparing them on different timelines. Oh, that is the ultimate procurement sin. A build's cost is heavily front-loaded in the acquire bucket.
You spend millions in year one. A buy or a wrap looks incredibly cheap in year one because the cost is spread across a subscription or an API build. Right, it looks like a steal.
But if you compare a build's five-year TCO against a buy's first-year TCO, you aren't doing rigorous financial analysis. You are just manipulating a spreadsheet to justify the answer you already wanted. You're comparing apples to a single slice of an orange.
Exactly. You must price the options over the exact same time horizon, say, a strict three-year or five-year window, so you can clearly see the exact month when the compounding subscription of a buy overtakes the front-loaded capital expenditure of a build. Okay, that gives us a really solid grip on acquire, run, and maintain.
But there is a fourth bucket in this total cost of ownership framework. Yeah. And according to the source material, this is the bucket that makes or breaks an executive's career.
Yes. It's the bucket Michigan deliberately ignored. The fourth bucket is the supervised bucket.
We call it the supervision tax. The supervision tax. This is the recurring human cost of watching the AI.
It is the salary of the reviewers checking the outputs, the auditors sampling the machine's decisions, and the IT staff kept on call when the system inevitably hallucinates or fails at two in the morning. Right. And you call it a tax because it is absolutely unavoidable for consequential systems.
And just like a real tax, every corporation spends an immense amount of energy looking for a loophole to avoid paying it. Which they never find. The central principle here is, the supervision tax is set by stakes, not by acquisition route.
I think of this like a bank vault. Okay, let's hear it. The number of armed guards you need to hire to stand outside a bank vault is dictated entirely by how many millions of dollars are sitting inside the vault.
Those are the stakes. Right. The number of guards is not dictated by whether you forged the steel door yourself in a custom build, bought the door from a security catalog in a buy, or just rented a secure building in a wrap.
If there are millions of dollars at risk inside, you need the guards. Period. That visualization hits the nail on the head.
The supervision tax does not shrink just because you chose to rent the software instead of building it. It doesn't magically go away. No.
If you buy an off-the-shelf AI fraud detector, you still absolutely need your own human employees reviewing its accusations before you freeze a customer's bank account. Renting the model does not rent you an excuse to stop watching. So how do we actually calculate this tax? Because an executive can't just throw a random budget number at the wall.
There has to be a formula. There is a strict mathematical formula, and you must run it before you sign any contract. Okay.
What is it? You take your expected volume of outputs per month. You multiply that by the percentage of outputs a human reviewer must check to maintain safety. You multiply that by the number of minutes a competent review takes.
And finally, you price that time at your fully loaded staff cost. Let's run this math live so we can really feel the weight of it. Let's do it.
Let's say we deploy an AI to review customer service claims, and it flags 10,000 claims a month for potential denial. Okay. 10,000.
We decide we can't trust it blindly, so a human reviewer needs to sample just 10% of those flags to keep the system honest. 10% of 10,000 is 1,000 claims a month that a human must look at. Keep going.
How long does a review take? Let's say the human has to pull up the customer's history, read the AI's logic, and make a call. Let's be aggressive and say it takes three minutes per review. Three minutes is fast, but okay.
1,000 reviews times three minutes is 3,000 minutes. That is 50 hours of human labor a month. But wait.
If this is a high-stakes decision, like denying a medical claim, we can't just do a 10% sample. No, absolutely not. We have to review 100% of the denials.
Run that math. 10,000 claims times three minutes each. 30,000 minutes.
That is 500 reviewer hours every single month forever. Yeah. Assuming a standard 40-hour work week, that is more than three full-time employees whose entire job is just babysitting the software we bought to replace human labor.
Exactly. And let me assure you, a vendor's spreadsheet-only proposal rarely highlights those three full-time salaries. No, they want you to think it runs itself.
Now, notice how changing the sample rate from 10% to 100% radically changed the cost. That introduces the two distinct operational shapes of human oversight. Right.
Human-in-the-loop and human-on-the-loop. Let's define the mechanical difference between those two, because they dictate the multiplier in our math. Human-in-the-loop means a person reviews and explicitly approves every single consequential output before it is allowed to affect the real world.
The AI might recommend freezing an account, but the system is hard-coded to wait until the human clicks approve before the lock engages. And human-in-the-loop is incredibly expensive, as your math just proved, but you are forced to use it when an error is severe and irreversible. Contrast that with human-on-the-loop.
This means the system is authorized to act on its own in real time. It routes the email. It approves the minor refund.
Right. But a human is monitoring the dashboard, sampling a percentage of the decisions after the fact, watching for statistical patterns of error, and remaining ready to intervene or roll back the system if things go off the rails. And on-the-loop is vastly cheaper, because you are only sampling maybe 5 or 10%, but you can only ethically and legally use on-the-loop when errors are immediately visible, and more importantly, when they are easily recoverable.
Right. So if an AI chatbot routes a customer's question about baggage fees to the wrong department, that is a recoverable error. The customer is annoyed, but they get transferred.
On-the-loop is fine. Exactly. But if an AI denies someone a home mortgage or, as we saw in Michigan, accuses them of a felony, that is a severe, life-altering error.
You must have in-the-loop. Which brings us full circle to the principle, Michigan proves the tax comes due either way. Let's connect this TCO framework directly back to the Midas disaster.
Michigan didn't just miscalculate the supervised bucket. They actively, intentionally zeroed it out. They completely ignored it.
The entire ROI justification for spending $47 million to build the system was firing the human adjudicators. Yeah. They didn't have humans in-the-loop to approve the fraud flags.
They didn't even have humans on-the-loop to sample the data and realize the machine was hallucinating a 93% error rate. No. They just plugged it in and let the machine execute the penalties.
They chose out-of-the-loop for a highly consequential, irreversible decision. And here is the brutal lesson of the supervision tax. The tax that Michigan refused to pay up front, which would have been the salaries of a few dozen adjudicators to review the flags, came due anyway.
It always does. You either pay the tax as salary up front, or you pay it as harm, brand destruction, and legal liability later. And the later bill is always orders of magnitude larger.
They tried to save a few million in salaries, and it cost them a $20 million settlement, a decade of horrific headlines, the complete destruction of public trust, and tens of thousands of ruined lives. This is the definitive takeaway for the listener right now. If a vendor or an internal engineering team puts a proposal on your desk showing a miraculously low total cost of ownership, and you notice the supervised bucket is left near zero for a system that makes high-stakes decisions.
A run. Yeah. You are not looking at a magically cheap system.
You are looking at massive unpriced risk. You are looking at a Michigan-style disaster waiting for an audit. Okay, so we have thoroughly mapped out the true costs.
We know you can't outsource accountability. We know the four TCO buckets, acquire, run, maintain, and supervise. We know we have to pay the tax.
Yes. But when you are sitting in the boardroom, how do you actually make the final strategic decision between build, buy, and wrap? You can't just go by gut feeling, and you definitely shouldn't let a charismatic vendor dictate the architecture. No.
You use a strict sequential framework, we call it, run the six axes in order. You evaluate the proposed capability against six specific business dimensions. Let's take them one by one.
The first axis is differentiation. Are we asking if this AI is our competitive moat or if it is just a commodity? Precisely. A moat is the fundamental reason a customer chooses to give your business money instead of giving it to your competitor.
If the AI capability you are discussing is your core moat, you are pulled heavily toward build. Because you want to own it. Right.
Building is how you create and own intellectual property that no competitor can simply rent from a vendor tomorrow morning. But if the capability is a commodity, meaning it is necessary for business, everyone have it, but nobody wins market share because of it, you are pulled strongly toward buy or wrap. I'll give a real world example.
Customer service called transcription. Perfect. I do not choose which airline to fly based on which airline has the most technologically advanced AI transcribing my angry phone calls.
No one does. It is purely a back office commodity. For an airline to spend millions of dollars and two years building a bespoke speech-to-text model from scratch is absolute madness.
It is burning capital to reinvent a wheel they could rent via an API for fractions of a penny per minute. The golden rule here is never build commodities. Excellent rule.
And the second axis is control. If differentiation dictates what you should build, control dictates what you are legally required to understand. Meaning how much do you need to see inside the machine? Exactly.
How much do you need to see inside and how much do you need to be able to change its internal logic? Right. High-stakes decisions about human beings, credit approvals, hiring filters, health benefits demand a concept called explainability. If the machine says no, you must be able to audit exactly why it said no.
And this need for explainability pulls you firmly toward a build or perhaps a self-hosted wrap where you control the environment and the data flow. Because a biased SaaS product is almost universally a black box. Let me push back on that.
Because I see SaaS vendors all the time marketing their AI as transparent or explainable AI. Are you saying they are lying? They are providing a dashboard that gives you a high-level summary of the decision. But you cannot inspect the actual mathematical model weights.
You cannot see the raw training data. Right. It's a proprietary trade secret.
Exactly. If an applicant asks your HR department why your bought software rejected their resume and you call the vendor, the vendor is not going to open up their proprietary algorithm for you to inspect. If you cannot explain the decision to a regulator because the vendor won't let you look under the hood, you have a massive control failure.
That makes perfect sense. The third axis is speed. How fast do you need this capability working in the real world? And buy is obviously the fastest, right? You just turn on a subscription.
Yeah. Wrap is quite fast, too. A good engineering team can wire up an API and a UI in a few wints.
Build is painfully slow, often operating on a timeline of 18 to 24 months before you see a return. We will come back to the dangers of speed in a moment. But let's move to the fourth axis.
Lock-in. This is huge. How hard is it to leave once you commit? The hard truth that executives need to swallow is that every single option locks you into something.
Let's map the traps. Buy locks you into a specific vendor. They can auto-renew your contract at a 50% markup.
Or worse, they can get acquired and discontinue the product entirely, leaving you stranded. Very common. Wrap locks you into a model provider.
You are dependent on their deprecation cycle. If they turn off the model version your problems rely on, your system breaks overnight. Also very common.
And build locks you into your own bespoke maintenance burden. You become completely dependent on the tribal knowledge of the one specific engineering team that built the spaghetti code. If your lead data scientist quits, your maintenance costs skyrocket.
This is why you must rigorously define reversibility before you sign anything. Is this project a two-way door, meaning it is technically and financially cheap to walk back out if the project fails? Or is it a one-way door, where untangling the system from your core operations would be so devastatingly expensive that you are trapped? You must price the exit costs before you walk through the door. The fifth axis is data sensitivity.
Where does the data physically go, and under what legal framework? Right. If your capability requires processing highly sensitive protected data like hypo-protected health records, or financial data sending that data out of your secure servers across the internet to an external API for a wrap, or to an external vendor's cloud for a buy, might lack a legal basis entirely. Right.
It doesn't matter if an external wrap is 90% cheaper than building it internally. If data privacy laws explicitly forbid you from sending unencrypted customer health records to a public generative AI model, external wraps are instantly excluded from the conversation. Exactly.
And finally, the sixth axis is liability. We have already established this deeply. The deployer is accountable.
You answer when it fails. So, to recap the six axes, differentiation, control, speed, lock-in, data sensitivity, and liability. But here is the reality of the boardroom.
What happens when these axes point in completely opposite directions? And often do. Let's say I am analyzing a task that is a total back-office commodity. That screams buy.
But the task handles highly sensitive customer financial data, which screams keep it in-house and build. And because it deals with finances, I need extreme control and explainability over the decisions. What do I do when the framework contradicts itself? This is where executives earn their paychecks.
The conflict between the axes is the analysis. You do not try to average them out to make the conflict disappear on a PowerPoint slide. You explicitly name the trade-off in your decision memo.
You write it down. This capability is a commodity, which means we should buy it. But it carries high-stakes consequences and sensitive data, which means we carry massive liability.
Exactly. So you price the tension. You might choose to buy the software to get the commodity economics, but because you recognize the massive liability and control deficit, you intentionally overfund the supervised bucket.
You hire an entire team of human reviewers to monitor the black-box sauce tool. Ah, I see. A recommendation that pretends there was no trade-off is a recommendation that is hiding unpriced risk from the leadership team.
That's a great point. We mentioned speed as the third axis. Let's dig into that because the source material has a specific warning.
Speed is one axis, not the decision. In the corporate world, speed usually shouts louder than all the other five axes combined. The phrase, we need this deployed by next quarter, is frequently used to stampede a terrible decision through procurement.
It happens every day. Look, speed is a completely valid business value. First-mover advantage matters.
But its proper structural place in this framework is to break ties among options that have already been deemed safe and compliant. It's a tiebreaker. Yes.
Speed never justifies skipping the differentiation analysis. It never justifies ignoring data privacy laws. And it certainly never justifies zeroing out the supervision tax.
Taking the fastest route without checking the lock-in exit cost is exactly how companies trap themselves in irreversible multi-million dollar mistakes. Okay, the framework sounds clean on paper. Four TCO buckets.
Six decision axes. Run the math. Write the memo.
But the real world is messy. Very messy. Vendors are incredibly savvy and they know exactly how to disguise these three options to make them look like whatever your procurement department is mandated to buy this year.
Let's talk about edge cases and disguises. Let's do it. Disguise number one is the build that is secretly a wrap.
Oh, this happens all the time. A vendor comes into your office and pitches a custom bespoke AI model tailored specifically for your industry's unique needs. It sounds amazing.
It hits all the right buzzwords. You are paying premium, front-loaded build prices. Right.
You think you're getting a custom house. But under the hood, they aren't training a custom model at all. All they have done is write a thin application wrapper and a clever 100-word system prompt over the exact same general-purpose base model that everyone else has access to.
It is a massive value extraction. You are paying a bespoke build sticker price. But you are inheriting all the lock-in of the vendor and all the vulnerability of the underlying base model they are quietly renting.
So how do you catch them? The tell here is simple. Ask the vendor what specific base model architecture they are using and ask them exactly what happens to your custom system when that base model is deprecated by the original provider next year. If they dodge the question or say, oh, we handled that seamlessly, you are buying a wrap wearing a built price tag.
Great advice. Okay, disguise number two is the exact opposite. It is the buy with a build hidden inside it.
Oh, yes. You find a sauce product. The sticker price in the acquire bucket is phenomenally low.
It looks like a massive win for the budget. Until you try to install it. Exactly.
To make this cheap software actually function with your company's archaic identity systems, your messy data lakes, and your highly specific workflows, it requires an army of external consultants and internal engineers writing thousands of lines of custom integration code. And that integration cost becomes so massive it rivals building a system from scratch. And worse, you now have to maintain that highly brittle integration glue forever.
Every time the vendor updates their software, your custom glue breaks. So you have to aggressively price that integration inside the acquire and maintain buckets rather than treating it as a free IT afterthought. Exactly.
Many supposedly cheap buys are actually hyper expensive builds hiding inside a monthly subscription. But there are strategic ways to play these options against each other too. Let's talk about the reversible first move.
Sometimes wrapping a model can be a brilliant strategic two-way door for a brand new feature. Say you aren't sure if your customers actually want an AI chatbot. Spending two years building one is a massive gamble.
Instead, you use a wrap to stand up a prototype in three weeks. You test the waters quickly and cheaply. It is an excellent strategy for gathering production evidence before committing massive capital.
But the source material warns that it is only genuinely reversible if you deliberately engineer the switching costs to remain low. Yes, you have to be disciplined. You have to pin and test the specific model version you are using and ruthlessly avoid building deep custom dependencies on that one single API provider.
Because if you build your entire company workflow around the quirks of one rented model, you can't leave. Right. A wrap without a meticulously planned exit architecture is just a build's lock-in but without any of a build's control.
And finally, there's the mixed estate. The reality is, for a complex enterprise capability, the right answer isn't build, buy, or wrap. It's all three.
It usually is. You decompose the architecture. You build the specific core algorithm that is your unique competitive differentiator.
The piece that actually makes you money. Right. You buy the commodity tools surrounding it, like the cloud database, the security layer, and the identity management.
And you wrap a flexible, off-the-shelf language model to handle the conversational user interface layer. Exactly. You run the six-axis framework on each individual component of the system.
This ensures your massive, expensive build budget goes only toward the components that actually earn their keep as a moat. Okay, let's bring this entirely down to earth. Okay.
Because as an executive, you won't always be making this decision from a blank whiteboard. More often than not, an eager manager, an ambitious engineer, or an aggressive vendor is going to slide a slick, finished 80-page proposal across your desk and ask for your budget signature. It happens every day.
So how do you audit a handed-down proposal? Let's use an immersive scenario to walk through the exact steps of auditing a proposal. We will look at a hypothetical operations manager named Felix. Perfect.
Felix is an operations manager at a mid-sized credit union called Rivermark Mutual. They have about 90,000 members. Felix's board wants an AI capability to scan and rank member accounts for potential fraud every night so the human fraud team knows exactly who to investigate first when they log in the next morning.
Okay, so Felix is smart. He doesn't start by looking at vendors. He starts by writing the job to be done in one single, precise sentence.
Which is? Rank member accounts by likelihood of fraud each night so human reviewers work the riskiest first. And he deliberately underlines the words human reviewers. Right.
He knows the Michigan story. He knows what happens when someone tries to delete the humans to save a quick buck. He then runs the six axes on this exact job.
Differentiation. Low. Every credit union on Earth does fraud triage.
It's a defensive commodity which points to buy. Control. High.
This touches people's money. If a member complains, Rivermark must be able to audit why an account was flagged. Data sensitivity.
Extremely high. It's protected financial data. Liability.
100% Rivermarks. So Felix looks at his access analysis and sees the inherent conflict immediately. The commodity nature of the task screams buy.
But the high data sensitivity and the need for control scream build. Or at least keeping it entirely in-house. Exactly.
Furthermore, because of the data privacy laws governing banking, he excludes external public wraps entirely right at the start. Sending unencrypted banking data to a public API lacks a clean legal basis. So he's narrowing his options.
He eventually chooses buy because the board needs it fast and Rivermark doesn't have the internal engineering talent to credibly build and secure a massive fraud model in a single quarter. Right. But because he understands the framework and because the stakes are so high-freezing, a member's account wrongly is a catastrophic, irreversible failure of institutional trust, he attaches a hard condition to the procurement.
What's the condition? He forces a permanently funded, human-in-the-loop review team into the budget. The AI ranks the accounts overnight, but a human must review the logic and clear every single flag before any account freeze is executed. No automated Michigan-style actions.
Felix audited his options and made a highly defensible, mature choice. He priced attention, but when you are the executive sitting across the desk and someone hands you a proposal, you need a checklist to ensure the person who wrote it was as thorough as Felix. You do.
There are four specific checks you must run on any proposal before you sign it. Let's lock these in. Check number one, find the missing buckets.
Okay, so almost every vendor proposal will clearly highlight the acquire bucket, the sticker price, it's on page one. Many will show the run bucket, the monthly subscription or estimated API cost, but very few will honestly price the maintain bucket, detailing how much internal IT time will be spent fighting drift and broken integrations. And almost none will accurately price the supervise bucket.
Exactly. Because adding a team of fully loaded human reviewer salaries makes the whole proposal look uncompetitively expensive. If the maintain and supervise lines on the spreadsheet are near zero for a system that makes decisions about people, the proposal is a fantasy.
Send it back. Check number two, find the over-weighted axis. Every biased proposal leans heavily on one single axis to carry the entire argument.
An internal engineer who really wants to build a cool new system for their resume will lean entirely on control and differentiation, completely ignoring speed and long-term maintenance. Or a finance officer will lean entirely on the lowest sticker price, ignoring liability and data sensitivity. Right.
You have to find the axis the author leaned on and then ask them explicitly what the other five axes are telling them. The axes they skipped are usually the ones that prove the recommendation is dangerous. Check number three, find the assumed volume and stakes.
This is where the math gets manipulated. Was the run cost calculated based on a small controlled pilot volume or the massive volume you'll actually see in full production? And crucially, was the supervision cost calculated based on the vendor's demo accuracy? Let's hammer on this because it is a massive trap. What is demo accuracy versus real-world accuracy? Demo accuracy is the success rate measured on the vendor's perfectly clean, highly curated, beautifully formatted test data.
It always looks like 99%. Of course it does. Real-world deployment accuracy, running on your company's messy, incomplete, typo-ridden legacy data is almost always significantly lower, sometimes drastically lower.
So if a vendor says our AI is 99% accurate, you might calculate your supervision tax thinking your humans only have to review 1% of the decisions. You budget for 10 hours a week. But in the real world, the AI is only 80% accurate on your messy data, which means your humans have to review a massive chunk of the outputs and your supervision tax explodes to 200 hours a week.
Exactly. If you price your human reviewer hours based on the flawless demo, you are hiding the true, massive cost of supervising a system that will inevitably make real-world mistakes. Treat a successful, low-volume pilot as evidence that a concept is technologically feasible.
Never treat a pilot as proof of deployment cost or proof of deployment accuracy. And the final check, check number four, find the exit. Flip to the end of the proposal.
Does the document explicitly state what it will practically and financially cost to leave this vendor or deprecate this model in two years? If the author cannot tell you the step-by-step process and cost to leave, they are asking you to approve a one-way door and calling it a simple purchase. Right. If a proposal fails any of these four checks, you do not sign it.
You send it back. You force the team to add the missing TCO buckets. You force them to balance the six axes.
You demand real-world production volume estimates. And you make them price the exit. You force the organization to make a decision based on the whole picture, not just the slattering, optimized slice the author wanted you to see.
And that fundamentally is the essence of executive governance. We have covered massive ground today, and we promised a master class. Let's summarize the absolute spine of this framework for the listener to take back to their desk.
Let's do it. First and foremost, three options, one accountability. Build, buy, or wrap it does not matter.
Renting the model does not rent the blame. When it fails, the deployer answers for it. Second, price the life, not the sticker.
You must run the math on the total cost of ownership across four buckets. Acquire, run, maintain, and supervise. Third, the supervision tax is set by stakes, not by acquisition route.
As Michigan painfully learned, removing human supervision to save a few salaries on a high-stakes system will cost you astronomically more in harm and liability later. Michigan proves the tax comes due either way. Fourth, run the six axes in order.
Differentiation, control, speed, lock-in, data sensitivity, and liability. And remember, speed is one axis, not the decision. Do not let the urgency of a deadline stampede you into ignoring data privacy or liability.
And finally, always audit a handed-down proposal. Find the missing buckets. Find the over-weighted axis.
Check the assumed volume against real-world messiness. And find the exit door before you enter. I want to leave the listener with one final concrete move, the single most valuable action you can take this coming Monday morning.
Before you approve any AI acquisition, force the team to write a specific, mathematically testable proof-of-failure line directly into the decision memo. Give me the actual phrasing. What does a proof-of-failure line look like in practice? It is a hard, pre-agreed metric.
For example, Felix at Rivermark might write, If, after 90 days in full production, our human review team finds that more than one in five of the accounts flagged by the AI have no genuine fraud signal, we treat this tool as mathematically failing, we freeze its use, and we initiate the exit plan. I think that is the most powerful tool we've discussed today. A decision you cannot test is a decision you can only admire.
By setting a proof-of-failure line, you remove the emotion and the sunk-cost fallacy from the evaluation. You give yourself a hard, mathematical tripwire to catch a disaster early. That's exactly right.
You catch the mistake in 90 days, instead of waiting 90 months for the Auditor General to expose it and the class-action lawsuits to hit. It forces absolute intellectual honesty on the entire procurement process, from the vendor all the way up to the board. It really does.
You know, going all the way back to the start of this deep dive, when we look at automation, we desperately want to think of systems like Michigan's Midas as an X-ray machine. Binary. Clean.
It either sees the broken bone, or it doesn't. Fraud or not fraud. But that's a fantasy.
Right. We are operating in highly complex, diagnostic, muddy waters. The real world is infinitely messy, and these combobolistic systems are not omniscient.
If you spend $47 million to build an X-ray machine, fire all the human doctors to save money, and let the machine automatically schedule amputations based on shadows it doesn't truly understand... You aren't innovating. You are building a tragedy. Exactly.
Price the doctors. Price the supervision. Price the entire life in the system.
Real cases
These examples show the decision, and its consequences, in real deployments. Real anchors are cited; the pattern each teaches is stated plainly.
Example 1: Michigan MiDAS, a build that deleted its supervision (2013 to 2015). Michigan paid a contractor about 47 million dollars to build a bespoke unemployment and fraud-detection system, and configured it to issue fraud determinations with no adjudicator review (Michigan Ford School STPP, 2024). It falsely accused about 40,000 people; the Auditor General found roughly 93 percent of a 22,000-determination sample were not fraud (Michigan Office of the Auditor General, 2016). The state settled Bauserman v. Unemployment Insurance Agency for 20 million dollars in 2022 (Michigan Attorney General, 2022). The decision failure was not "build versus buy." It was pricing the build without the supervision tax the stakes demanded. This is the module's anchor case and it owns the deep treatment here; other topics reference it.
Example 2: A commodity capability, correctly wrapped. Consider a mid-size logistics firm that needs to summarize inbound customer emails into a ticket. Summarization is a pure commodity: no customer chooses the firm because of it, the stakes of a wrong summary are low and recoverable, and speed matters. The six axes point cleanly at wrap: low differentiation, low control needs, low liability, high speed value. Building a summarization model here would be burning money to reinvent a commodity. The correct call is a thin wrap over a general-purpose model, with a light human check on the ticket before action. (Illustrative pattern, not a named incident; grounded in the axes in Section 3D.)
Example 3: A high-stakes decision that a buy cannot make safely alone. A hospital procuring an AI tool that flags patients at risk demands the opposite treatment. Even a bought product must be validated on the hospital's own population and kept under clinical review, because external validation has repeatedly shown vendor accuracy claims do not survive contact with a new population. (see Topic 4.6) The lesson: buying does not outsource the supervision tax for a consequential decision. The stakes, not the acquisition route, set the oversight. A buy here is defensible only with the Supervise bucket funded and staffed.
Example 4: The government build overreach that never worked as sold. A telehealth company built its business around a proprietary AI symptom checker and scaled on the promise it would replace clinical triage; the capability never matched the demonstration, and the company collapsed. (see Topic 3.1) for the deep treatment. The build-decision lesson: building your differentiation is right only if you can actually build the thing to the standard the stakes require. A build you cannot deliver is worse than a buy you can, because you have spent the money and still do not have the capability.
Example 5: The vendor whose "autonomous AI" was mostly people. A restaurant-technology vendor marketed drive-thru voice ordering as autonomous AI while offshore staff handled a large share of orders; a securities regulator charged it with misleading statements. (see Topic 3.3) for the deep treatment. The buy-decision lesson: the sticker on a buy is only honest if the product does what the deck claims. Interrogating that claim before you sign is the subject of the next topic, and it is why "buy" is never "buy and stop thinking."
Example 6: The reversible wrap as a first move. A common and defensible 2026 pattern (established) is to wrap a general-purpose model for a new feature precisely because it is the most reversible option: a small first commitment, a low switching cost if the provider disappoints, and real production learning about whether the capability is worth a later build. The organizations that use wrap well treat it as a two-way door and a data-gathering exercise, not a permanent architecture. (see Topic 0.4) The lesson: matching the size of your first commitment to the thinness of your evidence is itself a governance decision.
Example 7: The buy whose sticker hid a build. A recurring pattern, seen across public-sector and enterprise procurement, is the low-sticker product whose real cost lives in integration: connecting it to legacy identity systems, reformatting data, and building the glue that keeps it talking to everything else. The buyer then maintains that glue forever. The subscription looked like a buy; the total effort was a build hiding inside a subscription. The decision lesson: price integration inside the Acquire and Maintain buckets, not as a free afterthought, or the cheapest-looking option quietly becomes the most expensive. (Illustrative procurement pattern, grounded in the TCO buckets of Section 3C.)
Example 8: The GPAI upstream you inherit without choosing it. When you wrap a general-purpose model, you inherit the model provider's data choices and legal exposure whether or not you examined them. A downstream deployer's risk is shaped by decisions made far upstream, in the training data and terms of a model it did not build. (see Topic 5.5) for the deep treatment of what a foundation-model vendor's choices expose a downstream deployer to. The build-buy-wrap lesson: the wrap column of your memo must price not only the API bill but the inherited exposure, because renting the model rents its upstream problems too.
Example 9: The differentiator that justified building. Organizations whose core product is an AI capability, where the model itself is the reason customers pay, routinely and correctly build, because owning the model is owning the moat. The build-decision lesson is the mirror of the commodity case: when a customer genuinely chooses you because of this specific capability, renting it means renting your own differentiation from a provider who can rent it to your competitor tomorrow. Differentiation is the one axis that can justify the full cost and responsibility of a build, provided you can deliver and maintain it. (Illustrative pattern, grounded in the differentiation axis of Section 3D.)
Example 10: The self-hosted wrap forced by data law. An organization handling data that cannot lawfully be sent to an external provider, but wanting the flexibility of a general-purpose model, may run an open-weight model on its own infrastructure: a self-hosted wrap. The Acquire and Run buckets grow, because you now operate the model yourself, but the data never leaves your control, which is the only lawful path for that data. The lesson: sometimes the law, not the budget, selects the option, and the memo must say so plainly rather than let a cheaper but unlawful external wrap sit in the comparison. (see Topic 5.5) (Illustrative pattern, grounded in the data-sensitivity axis.)
Where people go wrong
- "The cheapest sticker is the cheapest option." The sticker is only the Acquire bucket. A cheap buy with a compounding subscription and a costly integration, or a cheap wrap with a Supervise bucket the stakes demand, can be the most expensive option over its life. Compare on total cost of ownership, never on the sticker.
- "If we buy or wrap, oversight is the vendor's job." The stakes set the supervision tax, not the acquisition route. A bought fraud detector or a wrapped model still needs your humans reviewing consequential outputs. Renting the model does not rent you an exemption from watching it.
- "Renting the model rents the blame." In a wrap, you are almost always the party the public and the regulator hold responsible, because you deployed it under your name. Michigan could not blame the contractor to the 40,000 people it wrongly accused. Deployers answer for deployments. (see Topic 8.4)
- "Build gives us the most control, so build is the safe choice." Build gives the most control only if you can actually build the thing to the standard the stakes require. A build you cannot deliver, or cannot maintain, is more dangerous than a competent buy, because you have spent the money and still lack a safe capability. (see Topic 3.1)
- "We need it this quarter, so speed decides." Speed is one axis of six, and the one most often used to stampede a bad decision. "We need it now" justifies choosing the faster option among safe ones; it never justifies skipping differentiation, control, liability, or the supervision tax.
- "Accuracy in the demo is the accuracy we will get." Demo accuracy is measured on the vendor's chosen data, not your real population. External validation repeatedly shows vendor claims shrink on contact with reality. (see Topic 4.6) Price the supervision that catches the errors the demo did not show.
- "A wrap is basically free once it is built." The API bill scales with usage and can surprise you badly at production volume, the model can change or be deprecated under you, and the Supervise bucket is set by stakes, not by how thin the layer is. Price the run cost at real volume and pin the model version as a dependency.
- "This is a one-time decision." A commodity today can become differentiating tomorrow as you accumulate proprietary data, and a wrap you chose for speed may justify a build later. Revisit the decision when the differentiation math changes; treat it as reversible where you can. (see Topic 0.4)
- "If the vendor calls it a custom model, it is a build." Some "custom AI" is a thin wrap over a general-purpose model wearing a build's price tag, carrying the lock-in of both. Ask what base model it uses and what happens when that model is deprecated. (see Topic 3.3)
- "The excluded options belong in the comparison to be fair." If an option is unlawful for your data or jurisdiction, it is not a live option, and including it to improve the numbers is how you end up defending an illegal deployment. Name the exclusion and its ground; compare only the lawful options. (see Topic 5.5)
- "The proposal on my desk already did the analysis, so I just approve it." A proposal usually prices the sticker and leans on the one axis its author cared about. Approving it as written means inheriting its blind spots. Find the missing Maintain and Supervise buckets, the over-weighted axis, the assumed volume, and the absent exit before you sign.
- "A pilot that worked proves the option works." A pilot runs at low volume, on friendly data, often with the team watching closely. Production is high volume, real data, and ordinary attention. A wrap that was cheap and accurate in a pilot can be costly and error-prone at scale. Treat the pilot as evidence about feasibility, not as the deployment cost or accuracy.
Questions people ask
- What is build?
- Acquiring an AI capability by developing and owning it end to end, either in-house or through a contractor building to your specification, including the model, code, and data pipeline. Gives the most control and differentiation and carries the most cost and responsibility. Michigan's MiDAS was a build.
- What is buy?
- Acquiring an AI capability by licensing a finished vendor product, typically as software as a service (SaaS), where the vendor hosts and maintains it and you pay a subscription. The fastest path to a working capability and usually the one you understand the least, because the product is often a black box.
- What is wrap?
- Acquiring an AI capability by building a thin application layer (prompts, guardrails, your own data, a user interface) around a general-purpose AI model you reach by application programming interface (API) and do not own. You own the layer and rent the model. Faster than build, more flexible than buy, and dependent on a model provider whose behavior, price, and terms you do not control.
- What is total cost of ownership (TCO)?
- The full lifetime cost of an AI capability across four buckets: Acquire (getting it working once), Run (operating it daily), Maintain (keeping it working as it drifts and models change), and Supervise (the recurring human-oversight cost). The sticker price is only the Acquire bucket; the decision lives in the other three. More on Total cost of ownership (TCO)
- What is supervision tax?
- The recurring cost of keeping a human meaningfully in or on the loop of an AI system, reviewing outputs, sampling decisions, and being ready to intervene. It is set by the stakes of the decision, not by whether you built, bought, or wrapped, and it does not go away by choosing to rent the capability. Removing it to lower cost, as Michigan did, is what turns a cheap capability into an expensive disaster. More on Supervision tax
Keep going
This lesson builds Buying AI well, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.