Build, buy, or kill: the quarterly portfolio review with real numbers
The short answer
A portfolio review judges every system together, on a schedule
You cannot see which system is quietly rotting or quietly carrying the business by staring at one; you see it by lining them all up. A fixed quarterly cadence inserts a cold look in the calm middle of a system's life, between the optimism of launch and the fury of failure, which is the only time an honest verdict is possible. Skip the schedule and the only reviews you ever get are the launch and the crisis.
What you will be able to do
- Run a quarterly portfolio review of your organization's AI systems: take the AI systems inventory you built in Module 0, put a real number on each system's cost and return, and produce one decision (build, buy, or kill) for every system in the same sitting.
- Distinguish the three verdicts precisely: build (keep investing to develop or scale a system in-house), buy (replace or acquire the capability from a vendor), and kill (stop and decommission), and name the fourth quiet option (hold steady) for what it usually is, a weak build that is avoiding a real decision.
- Calculate a system's honest total cost of ownership by adding the build cost, the run cost, the supervision tax you measured in Topic 8.2, and the risk cost, and set it against the honest return you established in Topic 8.1. (see Topic 8.1) (see Topic 8.2)
- Evaluate a kill decision on its merits by separating the sunk cost (money already spent, which is gone whatever you decide) from the forward cost (what keeping it alive will cost from today), and recognize escalation of commitment when you or a colleague argue from the sunk cost.
- Set a kill gate: a pre-committed threshold, agreed before anyone is attached to the system, that converts a future kill from an emotional fight into a decision that was already made.
- Judge the cost of a late kill: why killing a system at a gate before deployment is cheap and killing it after it has shipped and harmed people, as Ofqual was forced to, is catastrophic, and why a reversible system can be scaled boldly while an irreversible one must be shipped slowly.
- Defend each verdict in your portfolio review to a hostile board with the numbers behind it, so the review becomes the backbone of the investment memo in Topic 8.6 and survives the board inspection in the capstone. (see Topic 8.6) (see Topic 13.2)
The lesson
August 2020. The pandemic has closed England's schools, and hundreds of thousands of teenagers are waiting on A-level results that will dictate their university admissions. To replace the cancelled exams, the educational regulator, Ofqual, reaches for a statistical model to standardize teachers' submitted grades.
This diagram shows how their direct-center performance model operated. It pulled a school's grade distribution from the previous three years, and forced current students into that exact historical shape. But it contained a mathematical blind spot.
For large classes, the norm in state schools, it imposed the historical pattern. For small classes of five or fewer, it left teachers' generous grades untouched. The model's execution was devastatingly precise.
Nearly 36% of A-level grades dropped one full letter below the teacher's assessment. Another 3% fell two grades. Capable students at ordinary schools watched their university placements exaperate instantly, simply because of the historical record of the building they happened to study in.
The backlash was immediate. Crowds gathered outside the Department for Education, chanting slogans aimed directly at the mathematics that derailed their futures. Within four days, the government collapsed under the pressure.
The prime minister labeled the system a mutant algorithm. The results were entirely scrapped in favor of the original teacher assessments, and the chief regulator resigned a week later. The tragedy of Ofqual is that the algorithm did exactly what it was programmed to do.
The downgrade rate and the severe skew favoring private schools were highly visible in the regulator's own testing long before results day. The numbers that doomed the system were known. The failure was not a lack of data.
The failure was the absence of a strict governance mechanism to look at that bad data and kill the system at an early, cheap gate, before it was deployed to the public. Organizations with highly competent engineering teams routinely push flawed AI systems into production. They build robust testing pipelines, but fail to build the one mechanism that actually stops a deployment, a scheduled process to kill a system from the inside.
The corrective instrument is the quarterly portfolio review. This ledger forces a reckoning for every single AI asset an organization runs. You cannot govern an AI system by looking at it in isolation.
Every model in an organization competes against every other model for the exact same finite pool of capital, engineering, and human supervision capacity. The review forces a scheduled evaluation of the entire portfolio at once. It requires leadership to assign a definitive, binary verdict to every single system, forcing decisions that teams naturally want to delay.
Without this rigid ledger, failing systems are allowed to drift, and a system that should be killed will eventually be killed, by a regulator, by the market, or by a public crowd. This mechanism ensures you execute the kill on your own schedule. The review must run on a strict quarterly schedule.
This inserts a cold, routine examination into the calm middle of an AI system's life cycle. Without a hard schedule, systems are scrutinized at exactly two moments, the day they launch, when everyone is deeply optimistic, and the day they fail publicly, when everyone is panicking. This scale illustrates the core mathematical foundation of the review.
On the right is the honest return. On the left is the honest total cost of ownership, or TCO. The TCO side begins with the obvious expenses.
We calculate the ongoing build costs for future integrations, plus the run costs, the compute power, vendor fees, and data pipelines required to keep the model operational. The scale usually tips when we add the hidden expenses, beginning with the supervision tax. This is the ongoing human cost of keeping the system safe, the full-time reviewers, the escalation handlers, and the incident response teams.
Finally, we add the risk cost. This calculates the tail risk exposure, the specific financial impact of a regulatory ban, a discrimination lawsuit, or a headline-generating failure. On the other side of the ledger, the honest return must be calculated net of the work the AI creates.
You subtract the time agents spend overriding bad outputs and apologizing to customers from the gross efficiency the system initially promised. Leaving the supervision tax, or the risk cost, off this ledger allows a dangerous system to earn an unearned pass. A model that looks cheap to run often hides a catastrophic tail risk and a small army of human babysitters.
Running a portfolio review based on impressions or gross efficiency is corporate theater. Running it on these two honest numbers creates a governance record that survives a hostile board inspection. Once the numbers are established, we filter the risk by evaluating reversibility.
This tag dictates exactly how cautiously a system must be treated. A two-way door is a system that can be walked back cheaply if it fails, like an internal drafting tool. Because the penalty for an error is just flipping a switch, these systems are scaled boldly.
A one-way door is a system that determines outcomes for human beings, who gets a loan, a medical diagnosis, or a passing grade. By the time you detect an error, irreversible harm has already locked in. These systems ship slowly, and you kill them at the first sign of wavering metrics.
Even with rigorous numbers and explicit reversibility tags, human emotional attachment will try to keep failing systems alive. The partnership between MD Anderson Cancer Center and IBM's Watson for Oncology illustrates exactly how this distortion scales. A university audit revealed that this project consumed more than $62 million before being terminated.
It never guided the care of a single patient. This is the sunk cost trap, or the escalation of commitment. At every review, the loudest argument for spending more money was the massive amount of money they had already spent.
Money and engineering hours already spent are mathematically irrelevant to a live decision. That capital is gone regardless of whether you continue or stop. To neutralize this trap, ask the room one flat question.
If we were starting today with none of that money spent, knowing what we now know, would we build this? If the answer is no, the verdict is kill. Sunk cost is the enemy of the ledger, and explicitly stripping it from the room forces an honest decision. This chart illustrates the economic reality of ending a system.
The cost to kill is not fixed, it rises steeply the longer you wait. As a model moves from pilot to deployment, it accumulates dependent users, downstream workflows, and data integrations. Killing a system at the idea stage costs almost nothing.
Killing a deployed system that people rely on requires an expensive catastrophe management protocol. To halt this compounding cost, organizations implement a kill gate. A kill gate is a specific, mathematically measurable threshold, like the supervision tax exceeding the honest return for two consecutive quarters.
You agree to it in advance, while the team is cold and entirely unattached to the system. Removing these gates to move faster does not remove the eventual kill decision. It merely ensures that every kill happens at the most expensive possible moment, in full public view.
These rules combine to force the ledger's final output. Every system on the board must leave the quarterly review with one of three absolute verdicts. Firmly, the first is build.
You continue investing in-house capital, because the honest return clears the total cost of ownership, and the system provides a proprietary market advantage. The second is buy. You replace the internal system with a vendor product, because the market caught up and offers a lower cost.
You strictly maintain all internal oversight and monitoring duties. The third is kill. You stop and decommission the system because it fails the forward-looking math or it breached a pre-committed risk gate.
There is no fourth option. Hold steady or keep an eye on it is strictly banned. A hold is simply a build decision too timid to name its budget, or a kill decision too attached to say the word.
By forcing exactly one of these three verdicts, the portfolio review transforms ambiguous risk anxiety into explicit defensible business decisions. Organizationally, we reframe how we view these outcomes. A well-gated early kill is never a personal failure for the engineering team.
It is a critical portfolio success. Keeping a marginal system alive carries a massive opportunity cost. A doomed model actively starves your winning systems of finite capital and scarce human supervision.
This ledger's ultimate purpose is reallocation. When killed, a system's capital and reviewers are immediately routed away from failure and injected into high-performing builds. Executing this is a complete project.
A safety commission requires migrating workflows, archiving training data, and reassigning the workforce. Consider a logistics company executing a pre-agreed kill gate on a customer routing algorithm the moment it shows a bias towards small accounts. They end the project internally at the cost of a two-week decommission, completely preventing a disastrous public relations crisis.
Build the master ledger, document every AI system you operate, calculate the honest numbers, tag the reversibility of the doors, and enforce pre-committed gates without emotion. If you do not ruthlessly kill failing AI systems early on, your schedule, the market, regulators, or the public will inevitably do it for you.
The ideas, one by one
Every system leaves with exactly one verdict: build, buy, or kill
Build is a decision to spend more and must earn it; buy trades control for a lower total cost of ownership but never removes your oversight duty; kill stops and decommissions. "Hold steady" is not a fourth verdict; it is a build too timid to name its budget or a kill too attached to say the word. A deferred decision on a failing system is how you kill it too late.
Run it on two honest numbers, not impressions
Each system gets an honest total cost of ownership (build, run, supervision tax, risk cost) and an honest return (real value net of the work it creates). The supervision tax and the risk cost are the two most often omitted and the two most likely to flip a verdict. A review run on anecdote is theater; a review run on numbers is a judgment you can defend to a board. (see Topic 8.1) (see Topic 8.2)
Sunk cost is irrelevant, and saying so out loud is the skill
Money and time already spent are gone whatever you decide, identical on both sides, so they cannot belong in a live decision. When the argument for keeping a system leads with the past, it is escalation of commitment, and the cure is one flat question: if we were starting today, knowing what we know, would we begin this? MD Anderson's 62 million US dollar Watson project is what happens when that question is never asked.
Set the kill gate while you are cold
A pre-committed, specific, measurable threshold, agreed before anyone is attached, converts a future kill from a fight into the execution of a decision the team already made together. Ofqual's disaster read through this lens is the absence of a gate: the damning numbers were visible before results day, and with no line pre-agreed, the kill had to be improvised under a crowd four days after the harm.
The cost of a kill rises every gate you let a doomed system pass
Killing at the idea stage costs nothing; killing a deployed one-way door that people now depend on is catastrophe management, the Ofqual stage. So kill early or pay compounding interest to kill late. A gate is not a hoop, it is a cheap exit, and removing gates to move faster only guarantees every kill happens at the most expensive possible moment.
Reversibility decides how boldly you move
Scale two-way-door systems boldly, because being wrong costs a quick reversal; ship one-way-door systems slowly and kill them readily, because being wrong costs irreversible harm and no reversal can rescue you. Ofqual's algorithm was a one-way door dressed as routine: once grades issued, university places moved faster than any U-turn could catch.
A kill is a portfolio success, not a personal failure
Done at the right gate it frees capital, frees your scarcest reviewers, and removes uncompensated risk. Treating kill as failure is what makes kills late and therefore ruinous. Celebrate the reasoned kill, separate ending a system from judging its builders, and handle the human consequences with care rather than avoiding the kill to dodge them. (see Topic 9.1)
If you do not kill it, someone else will, on worse terms
A system that should die will be killed by the market, a regulator, the public, or a two a.m. incident if you do not kill it yourself. The review is how you keep the kill decision yours, made at a cheap gate on your calendar instead of at an expensive crisis on someone else's. That is why it is a governance instrument, not an accounting chore.
The biggest cost of a loser is the winner it starves
The opportunity cost rarely appears on a system's own row, and it is often the largest cost of all: the value of what the same money and, above all, the same scarce supervision could have done instead. "It is not costing us much" is usually wrong: even a cheap-to-run system consuming a reviewer is starving your best systems of the attention they need. A complete kill names where the freed resources go, because reallocation from losers to winners is the whole point of a portfolio.
A kill is a project, not a switch
The verdict is the decision; the decommission is the work, and it must handle the dependents, the data, the people whose roles change, and the record of why the system was killed. A kill agreed in the room but never assigned an owner and a plan is a kill that quietly keeps running, and a decommission done carelessly turns a good decision into a fresh incident. (see Topic 9.1)
The review is built to be defended
Its verdicts feed the investment memo a CFO stress-tests next, and the board inspects them line by line in the capstone. Every verdict must trace to a number a hostile reader can check, and every kill must be defensible on forward numbers alone. Write the reasons down now, because the board's hardest question, "why did you keep running that," must already have an answer that is a dated, gated decision. (see Topic 8.6) (see Topic 13.2)
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 62 of the podcast.
Read the full conversation
So in August of 2020, there was this this mathematical algorithm that quietly stole the futures of thousands of British students Yeah, and it was entirely because of the physical size of the classroom. They sat in right. It's wild Welcome to this deep dive everyone We are we're skipping the small talk today and getting straight into the mechanics of governing artificial intelligence in the real world Exactly.
I mean consider this an executive education session for those of you out there who are leading managing or Well, just trying to survive the current landscape of AI integration. Yeah, because it is a landscape today We're examining the architecture of a very specific corporate mechanism The topic is build buy or kill the quarterly portfolio review with real numbers, right? and it's focused entirely on budgets ROI and Rift management. Our mission here is to equip you with the exact framework to evaluate every AI system your organization owns We want to force hard decisions based on brutal economics rather than you know Just vendor hype and really to master the hardest skill in governance, right? How to kill a doomed system on your own terms before the market or a regulator does it for you? Exactly, because the reality you are facing right now is a sprawling estate I mean you have sanctioned software from enterprise vendors internal engineering builds and you almost certainly have shadow IT Projects running on corporate credit cards right now.
Oh, absolutely So to manage all of that you need a systematic method to judge what lives and what dies And we're gonna start not with the theory of that method But with the consequence of lacking it right a vivid case study of what happens when a kill decision is forced by a massive Public crisis rather than you know executed quietly at a scheduled gate So let's go back to the summer of 2020 in England The pandemic had closed the schools right and hundreds of thousands of teenagers were waiting on their a-level results Which are you know, the exams that dictate university admissions over there, but the actual exams were canceled, right? So the regulator which is awful the Office of qualifications and examinations regulation. They were staring down this impossible logistical nightmare Totally impossible teachers across the country had submitted an estimated grade for each student Along with a rank order of students within each class, but awful couldn't just rubber stamp the human grades, right? no, because the fear was national grade inflation if Every single teacher gave their students the benefit of the doubt the resulting grades would be meaningless to the universities, right? The universities wouldn't know who to admit So they turned to an algorithm to standardize the results. They called it the direct center performance model Yeah, the direct center performance model Can you walk us through the actual mechanics of what this model did because the technical failure here is it's incredibly specific Yeah, it really is So the underlying logic of the model was statistical enforcement the algorithm pulled a specific school res grade distribution from the previous three years Okay, it then took the current year students and roughly forced them into that exact same historical shape Wow, so if a state school historically produced say 5% a grades, but the teachers predicted 15% a grades for the current year the algorithm just crushed the surplus It systematically downgraded students to fit the historical performance of the physical building they studied in that is brutal But the fatal mechanism the detail that decided everything here was how it handled sample sizes Exactly because statistical models require sufficient data to make confident predictions for a class of 15 students or more The algorithm aggressively imposed that historical pattern and largely ignored the individual teachers judgment, right? But for a small class like five students or fewer the statistical confidence just vanishes You can't map a three-year bell curve onto three kids.
No, the math doesn't work So, what did it do for the small cohorts for those small cohorts? The algorithm essentially gave up it left the teachers predicted grades completely untouched Wow And when you step back and look at the physical reality of the British education system small intimate class sizes are Overwhelmingly found in fee-paying private schools. Yes Large packed classes are the norm in state-funded public schools, right? so by pure statistical design The model quietly handed the benefit of the doubt to the wealthy and it brought a mathematical hammer down on everyone else That's exactly what happened when results day arrived on August 13th 2020 the numbers delivered exactly what the model was designed to do which was devastating nearly 36% of all a-level grades came out one full grade below what the teacher had assessed Roughly 3% plummeted to entire grades Unbelievable bright high-performing students at historically underperforming state schools. Just watch their university offers instantly evaporate Yeah, and the fallout was immediate I mean students gathered in the streets protesting outside the Department for Education chanting slogans aimed directly at a piece of code, right? And four days later on August 17th under completely unmanageable public pressure the government panicked and reversed course They threw out the algorithm entirely and reverted to the original teacher estimates The Prime Minister publicly blamed a quote mutant algorithm and days later the chief regulator resigned So the critical lesson here for an executive listening to this is that the final decision To kill the algorithm and revert to human judgment was the correct one Yes, the correct decision was made the governance failure was the timeline the when and the how exactly That exact kill decision was available to a school the day before the results were published It was available a week before a month before but if they had executed the kill then I mean it would have cost them Huge internal frustration, wouldn't it? Sure, but it would have prevented the devastating political and human damage that followed I want to push back on classifying this purely as a governance failure though I mean, it's easy to play Monday morning quarterback here fair enough The engineers and statisticians at aqua were trying to solve a mathematically impossible problem under a crushing pandemic time constraint Isn't this ultimately just a tragic math error? I see why you'd say that but no like a bad waiting in the algorithm.
They couldn't possibly have foreseen. It was not a surprise That's the key the numbers damning the model were known internally well in advance. Oh wait, they knew yes off-chlorine Testing they could see the downgrade rate in their own internal sandboxes They could see the demographic skew benefiting private schools before they ever deployed it to the public Wow So they did not lack data not at all.
What they lacked was a scheduled mechanism to force a reckoning with that data They lacked a systemic willingness to read those testing numbers as a hard kill signal and act on it I see so killing a system late after the harm is propagated when the decision is forcefully taken out of your hands by an angry public or Regulator that is the ultimate failure of executive governance So an organization cannot just rely on having smart engineers or you know Hoping the testing phase catches everything. You can't rely on heroic crisis management No, absolutely not If you want to avoid an off cool style disaster You need a ruthless scheduled mechanism to evaluate the reality of your state which brings us to the core tool Ad-hoc panic is the enemy of sound governance, right? Yes, and the cure is a fixed standing court You have to move your organization from the chaos of retroactive crisis management to a cold Operational cadence and we define this as the portfolio review It is a single rigorous document and a single meeting that judges the entire set of your AI investments together Simultaneously rather than one project at a time, right? So this is the first major spine point of our framework today a portfolio review judges every system together on a schedule Let's dissect both halves of that concept Yeah, why judge them together because if I'm the VP of operations Why can't I just review my supply chain predictive model on Monday and then look at the HR screening tool next month? What's the practical business value of forcing them into the same room exactly the value comes from confronting the reality of fixed resources? We are talking about capital engineering talent and specifically your supervision capacity Okay Let's define that supervision capacity is the ongoing human cost required to keep an AI system safe accurate and compliant, right? Exactly every human reviewer you have assigned to babysit a marginal failing system is a reviewer who has actively stolen From securing a highly valuable Critical system. I see so an AI project does not merely compete against its own baseline return on investment No, it fiercely competes against every other AI system in your organization for that exact same scarce pool of human supervision And you just cannot see that zero-sum competition if you review them in silos You only see the resource drain when you line every system up side-by-side.
That's why they have to be together Okay, let's talk about the schedule part our underlying framework insists. This must be a quarterly exercise Yes quarterly, but for a busy executive team dragging every single AI system into a massive quarterly review Sounds like an unbearable administrative burden. It does sound heavy Why not just rely on continuous automated monitoring through a dashboard or to save time a deep comprehensive annual review? Well continuous monitoring provides the illusion of governance It produces dashboards alerts and red flashing lights, but it does not produce decisions.
That makes sense You just watch the red light flash right a team can stare at a degrading metric on a dashboard for six months Rationalizing the failure without ever actually pulling the plug Well about annual reviews an annual review is dangerously slow a doomed AI system left alone for 12 months Will accrue a massive web of technical dependence Integrations and entrenched user habits, right? If you wait a year the system becomes too entangled to kill cleanly exactly So a quarterly cadence strikes the necessary balance It is frequent enough to catch model drift because an AI system that looked flawless at launch might be quietly Deteriorating in accuracy six months later, but it's infrequent enough that the review remains a considered authoritative judgment Rather than a daily nervous reaction without a strict schedule reviews only seem to happen at two emotional extremes Let me guess at project launch when everyone is blindingly optimistic Yep Looking at vendor promises and the other extreme is total failure when everyone is furious and trying to manage the fallout So the quarterly schedule forces a cold Dispassionate look in the calm middle of the life cycle exactly and it elevates your process from a mere inventory, right? A lot of companies boast about having an AI inventory just a spreadsheet Listing what models they're running but a static list of systems with no attached verdicts and no financial numbers is just a wish list an Inventory commits you to nothing. It just allows the underperforming systems to keep burning resources a true portfolio review forces a definitive outcome Because sitting in a boardroom every quarter is a massive waste of expensive executive time if the outcomes are vague Oh, absolutely. You can't look at an underperforming customer service bot nod and say let's monitor it closely Our framework provides incredibly rigid boundaries for the decision that can be made in this meeting.
The strictness is entirely intentional Vagueness is the exact mechanism by which weak bloated systems survive So this is our next core rule of the room every system leaves with exactly one verdict build buy or kill There are absolutely no gray areas permitted. Let's unpack the realities of those three verdicts Starting with build this means the organization commits to keeping the system in-house Actively investing engineering effort and capital to develop or scale it, right? But this shouldn't be the default state just because an internal team wrote the code should it know Build is a verdict that must be aggressively earned When you declare build you are actively choosing to allocate more money more compute and more supervision capacity to a specific project Right, you're doubling down exactly, you only choose this path when the system demonstrably returns significantly more value than it costs and Crucially when the system is fit to your highly specific business problem provides a proprietary Competitive edge that you absolutely must own and control If it does not provide a unique edge building it in-house is a waste of engineering talent, which leads directly to the buy verdict Okay, so this means you replace the internal effort or acquire the capability outright from a third-party vendor. Yes You are trading total internal control for speed of deployment and lower immediate engineering costs But when you outsource the technology, you don't outsource the liability, right? Never buying is a valid strategic verdict, but it is never an escape hatch from corporate responsibility So if you purchase an off-the-shelf generative AI tool from a major vendor You still completely own the monitoring the risk management and the supervision tax precisely if that vendor system hallucinates a fake refund policy to a customer or discriminates against an applicant The regulator and the public will hold your brand accountable not the software vendor, right? The legal and reputational liability remains entirely on your balance sheet always and finally we arrive at the verdict that organizations actively run from kill you stop the system you turn off the compute and you decommission it you arrive at this verdict when the Financial return no longer clears the operational cost or when the systemic risk has vastly outgrown the business value It is the core action this entire methodology exists to make possible And look if the executive team lacks the courage to say kill The market or a regulator will eventually say it for them usually with catastrophic timing But in practice the most dangerous threat in that boardroom is not an overt refusal to kill a system Is it no it's not the danger is the fake fourth option the verdict that managers use when they want to avoid doing their jobs, you mean hold steady or Let's just keep an eye on it for another quarter exactly hold steady is a fog that obscures accountability It is almost always a build verdict from a team to cowardly to admit They are continuing to spend company money Or it is a kill verdict from a team to emotionally attached to their code to say the word out loud Yeah, it makes me think of well if a couple is having relationship issues, and they say they are taking a break Everyone knows that's not a real permanent state right right it's either a deferred breakup because they lack the courage to end it or Disguised reconciliation, but with absolutely zero accountability to the actual relationship in the interim That's a perfect analogy hold steady and a corporate review functions the exact same way So what if a system is genuinely worth keeping but it currently requires very little modification Then you must classify it as an explicit build verdict you assign it a small strictly defined Maintenance budget and a defined review date got it Everything else labeled hold is simply a failure to make a decision by forcing the absolute Discipline of exactly one of the three verdicts build buy or kill you you prevent average systems from devolving into corporate zombies that silently drain your supervision capacity and the dynamic between build and buy is also fluid right because a Proud custom in-house build from two years ago.
Let's say a bespoke natural language processor Might have been the right choice then oh sure, but the market moves violently fast Today an API from an external vendor might offer a vastly superior capability for a fraction of the compute cost So clinging to the legacy in-house version at that point is irrational the competitive advantage You thought you built has eroded if you refuse to shift the verdict from build to buy you are just protecting an engineer's ego At the expense of the company's bottom line But you cannot successfully hand down a kill or a buy verdict based on executive intuition no, absolutely not you can't make these decisions because a VP likes the aesthetic of the dashboard or Because the data science team is intensely proud of their architecture You have to run this court on hard math which brings us to the operational ledger of the meeting This is our next core principle run it on two honest numbers not impressions Yes, if you are going to defend a strict kill verdict against a hostile defensive internal champion You need an unassailable evidentiary baseline So how do we actually calculate these numbers without letting teams? You know lie with math the entire review rests on number one the honest total cost of ownership or TCO and number two The honest return let's dissect the honest TCO first because project teams Systematically manipulate or ignore portions of this to make their systems look profitable constantly an honest TCO is composed of four distinct layers Skip any one of them and you are artificially flattering the system Okay layer one is the straightforward build or acquisition cost the salaries of the engineers to develop it or the upfront license fee to buy it Right and layer two is the run cost the ongoing AWS compute bills the cloud infrastructure The vendor API fees, but the framework notes that teams constantly hide the data pipeline costs here Don't think oh all the time the ongoing unglamorous work of sourcing labeling and cleaning fresh data Just to keep the model from degrading right, but even so the pipeline costs are substantial But they pale in comparison to the two costs that are almost universally ignored Yes, layer three is the supervision tax the permanent human cost required to keep the AI system operating safely exactly It includes the human in the loop reviewers the escalation handlers the compliance auditors and the incident response teams So if you have a cheap automated loan approval model that flags 30% of its applications for manual review That is not a cheap model Not at all If you do not account for the salaries of the human underwriters swamped by those flags your TCO is a fiction Wow Yeah, and then layer four is the risk cost. This is the financial cost of the potential harm. The system can inflict Let's pause here because this is where people get tripped up Yeah, if I am an executive looking at a spreadsheet, I can calculate cloud compute costs easily I can look up the salaries of three human reviewers to calculate the supervision tax, right? But risk cost feels dangerously subjective if I'm evaluating an AI HR screening tool How do I accurately quantify a headline risk or a regulatory fine without just inventing a number to justify a kill verdict? The key is you are not trying to produce a flawless actuarial table You are aiming for a rough honest estimate based on highly specific worst-case scenarios Okay, the common failure mode is that risk teams try to compute AI risk as a blended average They multiply a low probability of failure by an average cost of harm, but AI risk it does not behave like traditional software bugs No, AI risk is tail-shaped.
It is highly improbable, but astronomically severe So if an HR tool exhibits a massive biased outcome at scale or triggers a total regulatory ban Averaging that catastrophe out across thousands of successful mundane interactions mathematically washes away the actual threat Precisely. You cannot price the average you have to price the tail You evaluate a typical failure case and you explicitly evaluate a specific worst-case scenario a rough Intellectually honest number that invites aggressive challenge from the room is Infinitely more useful than a precise dishonest number Which is usually zero right a zero that completely obscures the tail risk exactly and this risk cost is highly dependent on what we call a jurisdictional multiplier the physical location where the system operates Drastically alters the financial risk. Oh that makes sense If that HR screening tool is deciding the fates of applicants in the European Union It falls under the EU AI Act which carries massive compliance requirements and staggering financial penalties for violations Phasing in over the next few years and if you look purely at the United States, you are navigating a fragmented state-by-state patchwork It's a maze.
Colorado has enacted its own AI disclosure statutes Texas has specific intent focused prohibitions Illinois treats discriminatory AI and employment as a strict civil rights violation with severe penalties Yeah, and New York City requires mandatory independent annual bias audits for employment tools Wow, so the risk cost for a u.s. Facing system cannot be a flat national average It must reflect the strictest applicable penalty across every state the system touches, correct? So that gives us the honest total cost of ownership build run Supervision and risk the second number balancing the ledger is the honest return This is the real Tangible value produced by the system net of the work it creates, right? If a vendor promises a million dollars in efficiency gains, you cannot just log that gross benefit You have to subtract the cost of the rework the human corrections the customer churn caused by bad automated interactions and that massive supervision effort Once those two numbers are forcefully displayed on the screen the honest TCO and the honest return The arithmetic dictates the reality if the honest return clears the honest cost with a sufficient margin You build if the cost eclipses the return you kill it should be that simple, but corporate dynamics are rarely that clean Yeah, I mean you have the math on the screen the math unequivocally proves the system costs more than it returns The verdict should logically be kill that it won't be the manager who built it will sit in that boardroom and violently rebel against the arithmetic Why does the math lose to the room? Because at that moment you are no longer dealing with accounting you were dealing with psychology The math is simple the human resistance to the math is the core bottleneck of the entire portfolio review The ultimate enemy of the kill decision is some cost and the escalation of commitment Which is our next crucial rule some cost is irrelevant and saying so out loud is the skill It's the hardest skill to master So some cost in traditional terms is the money time or engineering effort already spent that you cannot recover Regardless of what decision you make today, right? But how does this specifically manifest in AI projects in AI the sunk cost fallacy is often disguised as technical optimism You hear things like the model is still learning or we just need to feed it one more quarter of clean data escalation of commitment is the behavioral flaw where an organization pours Increasingly more resources into a demonstrably failing course of action Specifically because of the heavy resources they have already sunk into it Exactly the verbal tell in a review meeting is any defense that starts by looking backward after everything we've invested We are so close. We can't just throw away two years of work There is a staggering historical example of this the MD Anderson Cancer Center as project with IBM's Watson for oncology Oh, yes a classic case The goal was to build an AI system that could ingest massive amounts of medical literature and patient data To guide complex cancer treatment and a later University audit revealed that MD Anderson spent over 62 million dollars on the effort 62 million It was eventually terminated while still in the development phase and it never guided a single live patients treatment The critical takeaway for an executive here is not to debate whether the underlying natural language processing was flawed, right? The governance lesson is studying the internal dynamics of a massive failing project Because at every single review meeting along the way that escalating sum of money the first 10 million then the 30 million Then the 62 million undoubtedly became the loudest most persuasive argument for spending even more Exactly the sheer weight of the past investment paralyzed the kill decision when mathematically the 62 million should have carried Absolute zero weight in the forward-looking strategy now if I am a VP and I walk into a board meeting to report on a struggling initiative and I say some cost Doesn't matter just forget the 62 million dollars. We spent the board is gonna fire me.
You'd think so, right? Yeah, the instinct is that it is deeply irresponsible to ignore massive past capital expenditures But that instinct is financially illiterate keeping a failing AI system alive does not recover the 62 million dollars It just sets more millions on fire to protect the ego of the executives who championed it precisely The 62 million is gone. It is ash Whether you continue the project or terminate it that money does not return So the only financial figures that have any place in a live decision are forward costs and forward returns What will keeping this system alive cost us starting today? What tangible value will it return? Starting today to break that psychological deadlock in the boardroom. The framework provides a specific tactical move The flat question this is a great tool when a project owner makes an emotional sunk cost argument The decider must look at them and ask if we were deciding today with none of that previous money spent knowing exactly what we know Right now about its performance Would we authorize starting this system from scratch if the intellectually honest answer is no The verdict is kill a mature board of directors will ultimately respect the executive who stops the bleeding based on cold forward realities The true skill of governance is possessing the professional ruthlessness To look at a team that has bled for project for two years and state that flat question out loud in a tense room But human beings are biologically wired to be terrible at ignoring sunk costs in the heat of a defensive confrontation Oh completely it is human nature to protect your territory and your creations.
So if we know we are going to be psychologically Compromised and defensive in that quarterly meeting. How do we bypass our own flawed psychology? We bypass it by making the decision before the emotional moment ever arrives. This is the next rule We set the kill gate while you are cold.
The kill gate is a pre-committed Mathematically defined threshold it is debated and agreed upon at the inception of the project before anyone has spent years Coding it before anyone is emotionally attached, right? It functions to convert a highly emotional future kill fight into the simple administrative execution of a prior agreement But a poorly constructed gate is entirely useless Like what if a team defines their gate as we will kill the system if it stops being valuable to the users They have built no gate at all. It must be a hard indisputable measurable line Okay So an effective gate sounds like if the honest return falls below the total cost of ownership for two consecutive Quarters the system must be killed or replaced by a vendor Exactly or targeting the supervision tax if the cost of the human reviewers required to correct the AI's output exceeds the gross value the AI produces we kill it or a strict fairness metric like if the error rate on any specific demographic group exceeds the baseline error rate by more than 5% The system is suspended immediately and killed unless the deviation is fixed within one sprint Let's return to the off-goal disaster for a second The ultimate tragedy of off-goal was the total absence of a kill gate If in May while they were relatively calm in designing the architecture they had agreed on a hard gate something like if the model downgrades more than 20% of the national cohort or The variance skews heavily against state schools based on class size We will not deploy it right the numbers generated in their own internal testing would have cleanly automatically tricked that gate It's like if you were trading equities you use a stop-loss order Yes Perfect analogy you buy a stock at $100 and you place an automated system order to sell it immediately if the price drops to $90 You don't have a committee meeting about it when it hits $90 the system just executes Exactly you do that precisely so you don't panic hold when the market plummets convincing yourself that it will bounce back tomorrow A kill gate is a stop-loss order for corporate AI It completely removes the need to summon fresh courage in the middle of a crisis when a gate is crossed The executive chairing the portfolio review is not acting as a villainous executioner terminating a team's dream They are simply reading aloud the rules that the team themselves wrote and agreed to when their minds were clear exactly But let's say a system does waiver and trip a gate Why is there such an urgency to execute the kill immediately? Why can't the decider say? Okay, we hit the gate Let's give the team three months to retrain the model and see if it recovers because waiting is a highly destructive financial act We have to examine the temporal physics of a kill decision The foundational rule here is the cost of a kill rises every single gate you let a doomed system pass Let's trace that escalating cost along the timeline at the initial idea gate killing a concept cost absolutely nothing You just erase a whiteboard right at the prototype gate You have lost a few weeks of engineering salary, but once it reaches full deployment killing it requires unwinding complex technical integrations migrating angry users and absorbing a reputational hit and if you let it survive to the final stage the off-goal stage where the Algorithmic outputs are directly impacting the material lives of the public killing of the system is no longer governance It is pure catastrophe management So a gate is not a bureaucratic hoop for a team to jump through a gate is an opportunity for a cheap exit Organizations that boast about removing governance gates to move fast and break things do not actually avoid kill decisions They simply guarantee that every kill they eventually make will happen at the most expensive most highly publicized moment possible And this urgency ties into a critical framework for understanding the nature of the decisions. We are authorizing Yeah, we have to distinguish between one-way doors and two-way doors.
This is all about reversibility Reversibility dictates your risk tolerance in the portfolio. Okay, let's define a two-way door It's a decision that you can walk back cheaply and quietly Consider an internal generative AI drafting tool deployed to help marketing employees write first drafts of emails That is a two-way door, right? If the model starts hallucinating or degrading in quality you turn off the server access on Friday afternoon Yeah, and nobody outside the building even notices exactly for two-way doors You can experiment iterate and scale boldly, but a one-way door is a deployment that is incredibly expensive painfully slow or virtually impossible to reverse once it touches real people an AI system that Algorithmically decides who gets approved for a mortgage who receives municipal welfare benefits or who is assigned a university placement grade? Like a more those are one-way doors. You must test these endlessly deploy them incredibly slowly and be willing to kill them at the slightest Provocation because with a one-way door by the time you're monitoring dashboard informs you that the AI was biased It has already altered human lives in ways that a software patch cannot unchange Local treated a massive one-way door as if it were a routine two-way administrative process Even though the UK government reversed the awful decision in just four days, which you know It sounds like an incredibly rapid technical rollback The university places had already been formally reallocated based on the initial algorithmic grades a fast technical reversal Does not undo the real-world harm of a one-way door Once the bell is rung the sound travels regardless of how fast you try to unring it the damage to those students Trajectories was already done for any system functioning as a one-way door an acceptable average return over time Does not justify the tail risk if the performance metrics even waiver if a fairness gate is even brushed against in a single Demographic you kill the system immediately.
You never grant a one-way door the benefit of the doubt So the numbers waiver the pre-agreed gate is tripped You assess that it is a one-way door and the decider in the boardroom pronounces the verdict But a kill verdict is not just a word typed into a spreadsheet It triggers a complex physical action and it has profound implications for how the organization's resources are subsequently managed This brings us to the mechanics of the aftermath. We have to completely demystify what a kill actually looks like operationally Culturally within the organization a kill must never be framed as a failure a kill is a massive portfolio success Yes, it frees up locked capital. It frees up that fiercely contested supervision capacity We defined earlier and it immediately eliminates a compounding corporate risk But the physical execution of the kill is a highly structured project not a light switch You don't just instruct an engineer to unplug the server a safe compliant decommission must meticulously handle four distinct vectors first vector Dependence what downstream workflows vendor integrations or secondary internal models will instantly break when this primary system banishes You have to map all of that second vector data What are the legal retention versus deletion obligations? Oh, right because if it's an EU system GDPR might require you to purge the personal data the model generated Well financial regulations might require you to securely archive the decision logs for seven years third vector people Which users and specifically which internal staff have entire roles built around managing the system? How are their roles transitioned and fourth the record? You must compile an immutable evidence pack detailing precisely why the system was killed based on what specific metrics? Tripping which specific gate this is built for a later inspection by the Board of Directors or an external regulator a lot of executives View decommissioning as mundane IT plumbing.
Yeah, if I am sitting in the c-suite focused on high-level strategy Why should I care about data logs and downstream API dependence? Shouldn't I just issue the kill order and let the technical teams handle the pointing because a technical dependency that is severed without warning causes a Massive operational outage. Oh if you execute a brilliant governance kill poorly on a technical level you transform a sound strategic decision into a brand new chaotic IT incident and a kill executed without a meticulously documented record looks exactly like negligence to an auditor a Regulatory will demand. Why did the system disappear? What were you hiding if you lack the data record you appear guilty of covering up a failure? Rather than executing a governed process and understanding the aftermath requires understanding Opportunity cost the system you are actively choosing not to fund Project owners often defend zombies by claiming it's practically running itself.
It's not costing us much to keep it alive that is an economic lie a low maintenance system that consumes even a fraction of one human reviewers time is actively Silently starving your most successful winning system of that reviewers critical attention Exactly a complete properly executed kill mandate must explicitly name exactly where the newly freed resources are being reallocated Shifting capital and human talent away from losers and funneling it aggressively into winners is the entire Economic purpose of running a portfolio review. Okay, we have established the theory. We have mapped the math of the ledger We understand the psychology of sunk cost and we have built the gates now.
Let's operationalize it let's step inside the room on the second Tuesday of the quarter and run a Highly realistic simulation of exactly how this meeting should sound and feel we will walk through the Harbor Line scenario This is a synthesized fictional company, but it represents a flawless execution of the mechanics first we must define who is permitted in the room a Functional review requires four distinct voices and only these four one the system owners They bring the necessary context and technical history, but they are understood to have inherent bias and attachment to the finance voice They enforce strict cost discipline and validate the TCO But they do not possess the authority to hand down verdicts three the risk or governance voice They verify the metrics against the pre-agreed gates and assess the reversibility of the doors and for the decider an executive leader holding real budgetary authority Because a review committee that is only empowered to produce recommendations is corporate theater The kills must be immediately executable the moment the meeting adjourns To achieve that the meeting order must be ruthlessly rigid step one Put all the hard numbers on the screen cold before anyone is allowed to speak step two Propose verdicts based strictly on forward-looking numbers only the decider must ruthlessly enforce the flat sunk cost Question the moment someone mentions the past step three read the pre-agreed kill gates aloud step four Write the defensive record and the decommission plan on the spot before the room empties Let's enter the room Camille is the AI governance lead in the decider at Harbor Line Her committee is reviewing an 18 month old AI system called the priority router Its function is to score incoming customer support tickets and decide which clients jump the queue for immediate human assistance So Camille puts the numbers on the screen The honest return is approaching zero because the human support agents are spending massive amounts of time Manually overriding the AI is nonsensical routing calls and the supervision tax is massive because those overrides take time But the governance voice highlights a far deeper issue The router is systematically burying small new customers at the bottom of the queue simply because they lack a deep Historical record of interaction with the company. It is the old school pattern manifesting in the b2b environment And it is a one-way door because a buried support ticket translates directly to a lost customer Confronted with the grim reality on the screen the lead engineer who built the router immediately deploys the sunk cost defense He argues the f1 accuracy score is trending up slightly. We just need to ingest one more quarter of clean data We have two years and over a million dollars invested in this architecture.
We cannot just throw it away Camille immediately deploys the flat question She looks at the engineer and asks if this router did not exist today And I described its current performance to you right now a near zero return Actively burying our smallest customers and requiring massive daily human intervention Would you ask me for a million dollars to build it from scratch the room goes silent the engineer cannot intellectually justify a yes Camille then pulls up the foundational document a year ago. They established a kill gate If the fairness metric the routing delay between enterprise and small customers crosses a 5% deviation We suspend immediately and kill unless fixed in one sprint The risk voice confirms the router is currently at an 8% deviation The gate is crossed Camille hands down the verdict kill notice the analytical trend here Camille Isn't just reacting to a static snapshot of today's terrible numbers. She's evaluating the slope of the degradation She executes the kill on the float long before the system hits rock-bottom and triggers a mass exodus of small clients and the formal defense Record she writes for the board relies entirely on the forward-looking metrics the explicitly crossed gate and a managed transition plan For the support agents the document makes absolutely zero mention of the two years of engineering investment The execution is flawless, but the human element is raw if you are that engineer you are devastated Two years of your professional life erased in a 10-minute discussion How does an executive stop this level of ruthless operational efficiency? From completely destroying the morale of the engineering teams They rely on to build the future by forcefully decoupling the kill verdict from the concept of personal failure The failure was not the engineering effort to build the router The success is the governance framework catching a degrading system early Camille publicly frames the kill as the reviewer Executing their exact job perfectly you care for the human element by ensuring robust workforce transition processes But you absolutely never dodge a necessary kill verdict simply to spare an engineers feelings Because you must control your destiny in this space or the market will inevitably Control it for you.
If you lack the discipline to kill a doomed system in your own boardroom Someone else is going to do it and they will do it on infinitely worse terms the market will simply render the tool obsolete and drain your budget a Regulator will ban it and fine you the public will protest it or an unmonitored Technical incident will blow it up at 2 in the morning The ultimate choice facing an executive is not whether a failing AI system dies The choice is whether it dies quietly on your calendar for cheap or whether it dies on the front page of the news Costing you millions in market cap and reputational ruin mature AI governance is never about avoiding all technical risk It is about aggressively pricing that risk Trapping it behind rigid gates and possessing the cold calculated discipline to act on the math when the time comes So here is the single most valuable action You can take this Monday morning when you return to the office take your sprawling AI Inventory and force every single system onto one page force your leadership team to assign exactly one word next to every project Build buy or kill base those words purely on forward-looking costs and returns Explicitly ban the phrase hold steady from the room Identify which of your zombie systems are silently starving your actual winners and pull the plug Thank you for joining us on this deep dive. Go force those hard decisions
Real cases
These are real, documented cases. Each shows what a portfolio review is for by showing what a build, buy, or kill decision looks like when it is made well or made too late.
Example 1: Ofqual's A-level grading algorithm (England, 2020). This is the anchor. With exams cancelled by the pandemic, Ofqual built the Direct Centre Performance model to standardize teacher-submitted grades against each school's three-year history (BBC News, "A-levels and GCSEs: How did the exam algorithm work?", 2020). The method leaned on the historical pattern for classes of fifteen or more, left the smallest classes' teacher grades entirely untouched, and blended for the cohorts in between, which systematically advantaged smaller private-school cohorts over larger state-school ones. On results day, 13 August 2020, nearly 36 percent of A-level grades were one grade below the teacher assessment and about 3 percent were two grades lower (BBC News, 2020). After public protest, the government reversed course on 17 August and reverted to the teacher-assessed grades; the chief regulator resigned on 25 August. The portfolio lesson is exact and has two halves. First, the kill (revert to human grades) was the right verdict and was available, far more cheaply, before results day, because the damning numbers were visible in Ofqual's own testing. Second, because there was no pre-committed gate and no willingness to kill early, the decision was made four days after national harm, at the maximum cost, with the choice taken out of the regulator's hands by public fury. A quarterly review with real numbers and a fairness kill gate is the instrument designed to force that decision at the cheap gate instead of the expensive crisis.
Example 2: MD Anderson and IBM Watson for Oncology (United States, terminated 2016, audited 2017; referenced). The deep treatment of this case belongs to Topic 3.1, where it anchors the gap between a vendor's demo and a deployed reality (see Topic 3.1); here it carries exactly one lesson, the sunk cost trap. MD Anderson Cancer Center partnered with IBM to build a Watson-based system to guide cancer treatment. A University of Texas System audit found the project cost more than 62 million US dollars, was terminated while still in development, and never guided the care of a single patient (STAT News, 2017; University of Texas System audit, 2017). It is the standing example of the sunk cost trap: at each review the mounting sum already spent argued for spending more, which is the one argument that should carry no weight in a live kill decision. The forward question, "if we were starting today, would we begin this," would have said stop long before the total reached 62 million. (Marker: the audit figures are as reported by the University of Texas System and contemporary reporting; the teaching point about sunk cost stands regardless of the precise final number.)
Example 3: The build-or-buy decision made with real budget (referenced). The Michigan MiDAS unemployment system is the owned anchor of Topic 3.2, where a state built an automated fraud-adjudication system that, run without human oversight, wrongly accused tens of thousands of people. (see Topic 3.2) It is referenced here only to mark the through-line: the build-or-buy choice that topic teaches at a system's birth is the same choice this review reopens every quarter, now with the added verdict of kill that MiDAS itself so plainly earned and was denied for years.
Example 4: The honest ROI that a review must use (referenced). Topic 8.1 anchors the honest-return discipline to the Indian company Dukaan, whose chief executive publicly cut about 90 percent of support staff for a chatbot and claimed large cost savings. (see Topic 8.1) It is referenced here because the honest return you established there is one of the two numbers this review runs on. A portfolio review is only as truthful as the return figures fed into it; a headline saving that ignores the work the system created will keep a losing system alive.
Example 5: Scoping against reality before you build (referenced). Topic 3.1 anchors the gap between an AI system's demo and its real capability to the collapse of the telehealth firm Babylon Health, whose symptom checker never matched its hype. (see Topic 3.1) The link to this topic is direct: a system whose demonstrated value never materializes is a kill candidate at the first review, and catching that gap early, at a cheap gate, is precisely what a scheduled review with honest numbers is designed to do.
The pattern across these cases. Read together, the examples show one recurring shape, and naming it makes the review's purpose concrete:
- The correct kill decision was almost always available earlier and cheaper than the moment it was actually made.
- The loudest argument against killing was the money already spent, which is the one argument a sound decision must ignore.
- Where the decision was delayed, it was not avoided; it was simply made later, by a regulator, a market, or a crowd, at far higher cost.
- The systems that did the most harm were one-way doors, where by the time the failure was undeniable, the harm could not be reversed.
- The biggest cost of the systems kept too long was rarely their direct run cost; it was the attention and supervision they drained from the systems that deserved it.
- In every case, a scheduled review with real numbers and a pre-committed gate would have converted a public catastrophe into a routine line item.
That shape is why this topic exists: the review is the difference between "we ended it ourselves, early, on the numbers" and "it was ended for us, late, in the news."
Where people go wrong
- "We already put so much into it, we cannot kill it now." This is the sunk cost fallacy, and it is the single most expensive sentence in AI governance. Money and time already spent are gone whatever you decide; they are identical on both sides of the choice and therefore irrelevant to it. The only costs that count are forward costs. The fix: whenever the argument for keeping a system leads with the past, ask the flat question, "if we were starting today, knowing what we know, would we begin this?" If the answer is no, the verdict is kill, and 62 million already spent, as at MD Anderson, changes nothing.
- "Kill means we failed." Wrong, and the belief is what makes kills late and therefore catastrophic. A kill at the right gate frees capital, frees your scarcest reviewers, and removes uncompensated risk; it is a portfolio success. An organization that never kills anything is not disciplined, it is hoarding half-alive systems, each drawing supervision and each a latent incident. The fix: name a well-reasoned kill as the reviewer doing their job, separate the kill of a system from any judgment of the people who built it, and handle the real human consequences with care rather than avoiding the kill to dodge them. (see Topic 9.1)
- "We will just keep an eye on it." "Hold steady" is almost always a build verdict too timid to admit it is spending money, or a kill verdict too attached to say the word. A deferred decision on a failing AI system is how you end up killing it four days too late, like Ofqual. The fix: force exactly one of build, buy, or kill for every system; if a system is genuinely worth holding, make it an explicit build with a small named budget and a defined next review, not a fog.
- "The review is an accounting exercise, so finance should run it." No. The portfolio review is a governance instrument that happens to use numbers. Its purpose is to keep the kill decision yours instead of a regulator's, a market's, or a crowd's, and that is a governance judgment about risk and defensibility, not a spreadsheet. The fix: run it as a governance review with real numbers, pairing the honest return and cost with the reversibility tag, the risk cost, and the defensibility of the system, none of which a purely financial view captures.
- "Sticker price tells us build or buy." A naive build-buy comparison pits an in-house build's engineering cost against a vendor's license fee and misses the two biggest numbers: the supervision tax on each, and the risk cost of each. A cheap-to-license vendor system you must still heavily oversee may cost more than it looks; a proud in-house build may hide a full-time reviewer. The fix: compare total cost of ownership to total cost of ownership, each with its full supervision tax (see Topic 8.2) and risk cost included, not sticker price to sticker price.
- "A high return justifies keeping it, whatever the risk." Not for a one-way door. A system that decides who gets a loan, a place, a benefit, or a grade can produce a fine average return and still be a kill, because its downside is irreversible harm to the people it gets wrong, and there is no cheap reversal to rescue them. Ofqual's model would have "worked" on aggregate statistics while wrecking individual futures. The fix: tag every system by reversibility, and when a one-way-door system's fairness or safety numbers waver, kill it rather than betting on an average.
- "We can decide to kill it if and when it goes wrong." Deciding to kill in the moment it goes wrong is deciding under attachment and panic, exactly when judgment is worst, which is how the kill arrives late and expensive. The fix: set a kill gate in advance, when you are cold, a specific measurable threshold (return below cost for two quarters; supervision tax above value; a fairness metric across a line) agreed by the whole team, so that when the gate is crossed the kill is the execution of a prior decision, not a fresh fight.
- "The build-or-buy question was settled when we launched." Stale. The market catches up, the supervision tax reveals itself, and a capability's advantage erodes; a build that was right two years ago can be a buy today, and clinging to the in-house version is sunk cost in an engineer's badge. The fix: reopen the build-buy line every quarter on current numbers (see Topic 3.2), and be willing to retire your own creation when a vendor now delivers the same honest return for less.
- "If we remove the gates, we will move faster." Removing gates does not remove the kill decisions; it guarantees every kill happens at the most expensive possible moment, in public, on someone else's timing. A gate is not a hoop, it is a cheap exit, and its whole value is that killing at a gate is affordable while killing after it is not. The fix: treat each gate as an option to exit cheaply, and remember that the cost of a kill rises every gate you let a doomed system pass.
- "We reverted the system quickly, so no harm done." A fast reversal does not undo the harm a one-way door has already caused. Ofqual's four-day U-turn was correct and still could not restore the university places already reallocated in the hours after results day. The fix: judge a late kill not by how fast you reversed but by how much irreversible harm occurred before you did, and let that judgment push one-way-door systems to be shipped slowly and killed at the earliest wavering, not the last.
- "The numbers were not clear enough to act on." Usually the numbers were clear and the will was absent. Ofqual could see the downgrade rate and the private-school skew before results day; MD Anderson could see the cost curve. The signal to kill is rarely missing; the courage to read it as a kill is. The fix: pre-commit to the gate so that reading the number and acting on it are the same act, and the review is where the reading happens on schedule, not the crisis.
- "A kill is the end of the story." A kill is a decision that must itself be defensible and documented, because it will be inspected. A system killed without a recorded reason looks, to a later auditor, like a system that vanished, and a system kept without a recorded reason looks like negligence. The fix: write every verdict, especially every kill, with its forward numbers, its gate, and its human plan, into the portfolio review, so the board inspection in the capstone finds a reasoned record, not a gap. (see Topic 13.2)
- "The kill decision ends the system." No; the kill verdict is the decision, but the system only ends when the decommission is done, and a decommission is a small project with real risks of its own. The fix: assign every kill an owner, a date, and a plan for its dependents (migrate or gracefully cut), its data (retain or delete lawfully), and its people (support through the workforce process); a kill agreed but not executed is a system still quietly running. (see Topic 9.1)
- "This system barely costs anything, so leave it." A small direct cost hides a large opportunity cost: the scarce supervision and management attention the lingering system consumes is attention your best systems are denied. The fix: judge a keep by what the same resources could do elsewhere, not by the modest figure on the system's own row, and free the resource for a named better use when the system is not earning it.
- "We run everywhere, so one risk number fits the system." Wrong; the same system can carry a very different risk cost in different jurisdictions, because the law that turns a harm into a liability differs. The fix: read each system's risk cost through the actual legal and reputational exposure where it runs (a high-risk load under the EU AI Act, a lighter formal cost but real civil exposure under a principles-based regime, a specific operational cost under strict content-labeling rules), rather than defaulting to a single home-country figure.
Questions people ask
- What is portfolio review?
- The periodic, evidence-based examination of every AI system an organization runs, judged together rather than one at a time, that assigns each system a single verdict (build, buy, or kill) on real numbers. It is the mechanism that keeps the kill decision in the organization's hands.
- What is build (verdict)?
- The decision to keep investing your own effort to develop or scale an AI system in-house. It is a decision to spend more, justified when the honest return clears the honest total cost, the fit is worth owning, and you have the supervision capacity to keep it safe as it grows.
- What is buy (verdict)?
- The decision to acquire or replace an AI capability from a vendor rather than carry it in-house, justified when a vendor's total cost of ownership is lower or the capability gives no ownership advantage. It shifts part of the cost but never removes the deployer's oversight and monitoring obligations.
- What is kill (verdict)?
- The decision to stop an AI system and decommission it, justified when the honest return no longer clears the honest total cost, the risk has outgrown the value, or the system cannot be made defensible at an acceptable price. A kill at the right gate is a portfolio success, not a failure.
- What is total cost of ownership?
- The full forward cost of keeping a system: build or future-build cost, run cost, the supervision tax, and the risk cost. Omitting any part, especially the supervision tax or risk cost, flatters the system and corrupts the verdict.
Keep going
This lesson builds ROI and value measurement for AI, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.