The Model Inventory and Tiering That Survives an Examiner
The short answer
An examiner asks for the inventory before any single evaluation report
One strong report proves one system was checked. A complete, tiered inventory proves the organization knows what it runs at all, which is the artifact Earnest Operations could not produce after eleven years.
What you will be able to do
- Analyze why a model inventory, not a single system's evaluation report, is the first artifact an examiner or auditor asks for, using the Earnest Operations settlement as the case where its absence became a legal finding.
- Define the fields a defensible inventory row must carry: a named owner, a stated purpose, a materiality score, a complexity score, a resulting tier, a current validation status, and a next review date.
- Apply a two-factor tiering rule, materiality by complexity, to sort a mixed set of real systems into three tiers, and justify each placement in one sentence.
- Distinguish your organization's own materiality-and-complexity tier from a legal risk category such as the EU AI Act's high-risk classification, and explain why the two exist for different reasons and can disagree.
- Diagnose the shadow-model and undocumented-override problem: models and human bypass patterns that exist in production but never made the inventory, and explain why an inventory with silent gaps is worse than no inventory, because it looks complete.
- Argue a tier up or down for a specific model when its stated risk framing (for example, "it only drafts, a human decides") understates the harm a wrong output could still cause.
- Assign validation depth, revalidation frequency, and sign-off authority to each of the three tiers, and explain why the current US banking model risk guidance excludes generative and agentic AI from its own scope, and what that means for who must extend the discipline anyway.
- Critique an inventory handed to you by a colleague or a vendor, identifying missing fields, silently excluded systems, and tier assignments that do not match the rule applied to the rest of the list.
- Defend a completed model inventory against a hostile examiner's core question: how do you know this list is the whole list.
The lesson
On July 10, 2025, the Massachusetts Attorney General released the findings of a multi-year investigation into the student loan lender Ernest Operations LLC. The investigation resulted in a $2.5 million settlement targeting the company's automated underwriting and internal governance controls. For over a decade, the lender relied on a three-stage algorithm to process every applicant.
This system decided who received a loan and who was rejected before a human ever read the file. One specific variable, the cohort default rate, penalized applicants based on the average default rate of others who attended the same school. And this specific rule disproportionately impacted graduates of historically Black colleges and universities.
The investigation also surfaced a deeper systemic failure. Human underwriters were routinely bypassing the model's outputs without clear documentation. Because these overrides were unrecorded, the company could not say which model made which decision, who was accountable, or why a human ignored it.
The penalty was tied to a specific organizational inability. Ernest could not produce a current, owned account of its own automated decision-making ecosystem. At its root, this was an inventory failure.
No single verified list of the models running inside the building existed anywhere in the organization. For a risk manager like Elmer at Larkspur Community Bank, the Ernest settlement serves as a prompt to evaluate his own internal ledger. Many teams operate under the assumption that because their most important system is validated, the organization as a whole is protected from scrutiny.
In a regulatory audit, the examiner requests the model inventory first, before they ever ask to see a single evaluation report. An evaluation report proves one system was checked. Only a complete, tiered inventory demonstrates that the organization understands the full shape of its automated risk.
The remedy the attorney general imposed on Ernest makes the path forward explicit. The lender must now maintain a written corporate governance system for its AI models. To survive an audit, this inventory row must be composed of exactly seven non-negotiable fields.
The first two columns establish the name and version of the system, followed by the owner. This owner must be a specific named person, not a team or a department. Field 3 is the purpose, a one-sentence statement of exactly what decision the system makes and for whom.
Fields 4 and 5 are the scores that drive the entire governance plan, materiality and complexity. The final two fields record the resulting risk tier and the validation status, which includes the date for next mandatory review. Missing even a single field from this row transforms the inventory from a defensive artifact into a legal finding.
The inventory uses a mechanical sum to prevent individual teams from subjectively downscoring their own systems. Materiality asks what it costs the organization or the public if the model is wrong. This is scored 1 to 3 based on the worst realistic outcome, never the average case.
Complexity measures how difficult it is to trace the system's reasoning and how much its output can vary without a human choosing that variation. A score of 1 is reserved for simple, fixed calculations and arithmetic logic that a person can verify by hand. A score of 3 identifies opaque or generative AI models whose logic cannot be fully traced by inspection.
The scores are added. A sum of 5 or 6 forces a Tier 1 critical designation. 3 or 4 creates Tier 2, significant.
A sum of 2 is a Tier 3, limited system. This formula replaces intuition with arithmetic. It ensures that critical systems receive high-depth validation, regardless of whether a team finds that scrutiny convenient.
One common argument used to lower materiality scores centers on the human reviewer. The reasoning suggests that if a person reads every draft before it is sent, the model itself carries only minor risk. We can test this reasoning by looking at a generative assistant that drafts customer support replies.
Under high ticket volume, a human agent has only seconds per draft. They stop reading for nuance, scanning only for obvious errors before hitting send. A human in the loop changes how a system is monitored, but it does not lower the potential harm of a wrong output.
Review quality inevitably degrades as volume increases. If we score the assistant honestly, it carries moderate materiality for factual errors and high complexity as a generative model. 2 plus 3 is 5. It lands in Tier 1. A system reaches the highest risk tier based on what it is capable of producing, even if a human nominally manages the button.
Practitioners also risk confusing internal operational risk with external legal classifications. The EU AI Act is the current global benchmark for a legal risk classification system. The Act defines statutory obligations for high-risk systems, but those legal categories are distinct from an organization's internal need to determine validation depth.
In the United States, banking supervisors follow a different framework, known as SR-26-2. SR-26-2 currently excludes generative and agentic AI from its own scope, describing these systems as too novel for the existing banking rules. A regulatory carve-out defines the reach of a specific rule.
It is not a finding of safety. Internal tiering discipline must be applied to these systems, even if a the inventory's value depends on its completeness. The primary threat is the shadow model, any system shaping decisions that never made the official list.
These often enter as business tools, like a spreadsheet regression. Others arrive through vendor software, purchased as just a license without separate governance review. As the Earnest case showed, undocumented overrides function as shadow models, creating a gap between intended logic and actual outcomes.
An inventory with silent gaps creates an illusion of coverage, inviting confidence the organization hasn't earned. Closing these gaps requires active discovery, mandatory intake for new tools, procurement sweeps, and reporting undocumented spreadsheets. Organizations must adopt a version reopening rule.
Any change to a model's training or inputs triggers a fresh scoring of materiality and complexity. When an examiner arrives, they are testing exactly one thing. Whether the organization truly knows the complete shape of its own risk, they will not wait months for a list to be reconstructed from memory and emails.
The tiered inventory functions as the map of enterprise consequence. It is the artifact required to identify high-risk systems, establish legal conformity, and answer the opening question of a regulatory audit.
The ideas, one by one
A defensible inventory row has seven fields
Name and version, owner, purpose, materiality score, complexity score, resulting tier, and validation status with a next review date. A row missing any of these cannot be defended under examination.
The tiering rule is materiality by complexity, summed
Materiality (1 to 3) scores the worst realistic outcome if the system is wrong. Complexity (1 to 3) scores how traceable and how variable the system's output is. The sum, 5 to 6, 3 to 4, or 2, produces Tier 1 Critical, Tier 2 Significant, or Tier 3 Limited, mechanically, without room for a preferred outcome to bend the score.
A human in the loop does not lower materiality by itself
Review quality degrades under volume the same way clinician attention degrades under a flood of false alarms. Score the worst realistic outcome regardless of who technically clicks the button, which is exactly why a generative customer-reply assistant can land in Tier 1 despite making no decision on its own.
The inventory is not a legal risk classification
(see Topic 5.3) The EU AI Act's high-risk category and this topic's materiality-and-complexity tier answer different questions and can disagree; a system can be legally low-risk and organizationally Tier 1, or the reverse.
A regulatory carve-out is not evidence of low risk
SR 26-2 expressly excludes generative and agentic AI from its scope. The discipline this topic teaches, inventory, tiering, validation, review dates, transfers to those systems as a matter of practice, whether or not a specific rule currently requires it.
Shadow models and undocumented overrides are the two ways an inventory quietly fails
A shadow model never entered the list at all. An undocumented override means the model's stated decision and its actual real-world outcome are two different, unreconciled things, exactly what the Earnest investigation surfaced.
Each tier's requirements are non-negotiable and deliberately unequal
Tier 1 gets a full evaluation report, independent challenge, and annual revalidation. Tier 3 gets basic documentation and a review every few years. Treating every model the same, whether by over-scrutinizing or under-scrutinizing, defeats the purpose of tiering at all.
The inventory decays without active maintenance
Overdue review dates, deleted rather than retired rows, and version changes that never reopen a row all quietly turn a once-accurate inventory into a document the organization treats as current long after it stops being true.
The inventory travels forward
This artifact feeds the conformity file, the evidence annex, and every model and system card the organization produces. Build it complete now and every downstream artifact that depends on it inherits that completeness. (see Topic 5.6) (see Topic 10.2) (see Topic 10.3)
Tiering errors run in two directions, and both are costly
Under-tiering hides real risk behind a lighter review than a system's true consequence warrants. Over-tiering makes the inventory unaffordable to maintain and dilutes the meaning of the highest tier. Both come from letting a preferred outcome, rather than the rule, decide the score.
The inventory itself needs an owner, distinct from any single model's owner
That role is accountable for the discovery process, the overdue-review count, and consistent rule application across rows, and is the person an examiner will ask to speak with first about the list as a whole.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 33 of the podcast.
Read the full conversation
For 11 years, a student loan lender called Earnest Operations ran the exact same three-stage algorithm on every single applicant. Right, and you really have to let that time frame sink in. 11 years.
Yeah, over a decade. I mean, this system operated entirely in the background. It was deciding who got a loan, whose application got rejected, and it was doing all of this long before a human ever set eyes on the file.
It operated with absolute silent authority. Exactly, and one of its rules was a hard-coded knockout. It just automatically denied any applicant who lacked at least a green card.
Just a blanket denial. Right, no secondary review, no consideration of alternative credit data. Just a hard stop, but the mechanics get even more complicated than that.
They do, yeah, because another rule in that same algorithm folded in a variable called the cohort default rate. Right, and we should clarify what that is because it wasn't a measure of the individual applicant's credit worthiness. No, not at all.
It was the average default rate of historical borrowers who had attended the identical school as the current applicant. Which means we are dealing with a system that shifts the burden of historical statistics directly on to the individual. Exactly.
As a direct result of that specific variable, you could have two individuals possessing identical income, identical credit scores, identical flawless repayment histories. But they'd get totally different answers. Right, entirely different answers simply because one of them went to a different college.
And here is the fatal flaw in their governance. Nobody at the company had ever tested that variable for disparate impact on the specific demographics of the applicants it was grading. Never.
So welcome to the deep dive. If you are a sharp, busy professional, say you're an executive, a risk officer, or a compliance leader, and you're preparing to navigate the highest levels of enterprise AI governance, well, you might think your primary challenge is building better, more accurate algorithms. Which is what everyone thinks, right? Yeah, exactly.
But today's stack of research shows that what you actually need, before anything else, is a better map of the algorithms you already have. Because on July 10th, 2025, that silent 11-year reality finally caught up with earnest operations. It absolutely did.
The Massachusetts Attorney General announced a $2.5 million settlement with the company. Wow. And, you know, if you read the press releases, the public allegations focused squarely on those knockout rules.
They highlighted that cohort default rate variable and its disproportionate discriminatory effect on Black and Hispanic applicants. Right, many of whom were graduates of historically Black colleges and universities. Exactly.
So the bias finding is the headline that the media ran with. But as a professional guiding enterprise AI governance, the detail that should stop you cold is actually a quieter one. It's buried much deeper within the investigation's findings.
Yeah, because if I'm reading this investigation correctly, the algorithm itself was really only half the problem. The underwriters of the company routinely bypassed the models, right? Uh-huh. Yes.
They overrode the automated decisions without clear standards. And worse, without documentation. So after 11 years of these automated and semi-automated decisions, when the regulators finally came knocking, the company could not produce a clean answer to the most basic governance question there is.
They couldn't. They couldn't say which model made which decision, who owned that model when it was last checked for safety, or who overrode its output. And this brings us to the core mission of our deep dive today.
Yeah. I mean, this is not just a bias story, is it? No, it is fundamentally an inventory story. The remedy imposed on Ernest by the Attorney General wasn't just a financial penalty.
The state forced them to build a written corporate governance system of internal controls and risk assessments. Right. And if you strip away the legal freezing, that mandate describes a single, highly specific artifact.
And today we are building exactly that artifact for you. We are mapping out the model inventory and the tiering system that actually survives an examiner. And we're going to do it step by step.
Let's start with the foundational reality of regulatory scrutiny. If you take nothing else away from this discussion, write this down. This is our first core principle.
An examiner asks for the inventory before any single evaluation report. Yes, it is inevitably the first request they make. And the reason is entirely structural.
How do you mean? Well, let's look at it from the examiner's perspective. If you confidently hand an examiner one incredibly thorough, robust, perfectly documented evaluation report for your flagship underwriting model, you have proven exactly one thing to them. That you checked one system.
Precisely. You proved you checked one system. That report tells them absolutely nothing about the other 40, 50 or maybe 100 AI systems running in the shadows of your organization.
I want to pause on that because I think the natural executive instinct is always to showcase your best work. Sure. It's human nature.
Right. Like if I'm the head of risk, I want to slide my thickest, most comprehensive audit across the table to prove, you know, that we take this seriously. Are you saying that actually backfires if I don't have the broader list? Well, it doesn't necessarily backfire, but it is deeply insufficient.
It is a defense of a single node, right? When the examiner is asking for a map of the entire network. OK, that makes sense. We saw this exact dynamic play out in the healthcare sector with the Epic sepsis model.
Oh, right. Tell us about that. This was an algorithm designed to predict which hospital patients were at high risk of developing sepsis.
Hundreds of hospitals trusted the vendors performance numbers for years, and they ran this model on live patients. Real patients in ICUs? Yes. But for a long time, nobody at these individual hospitals wrote an independent evaluation report on how the model was actually performing on their specific local populations.
Let me make sure I understand the mechanics of that. Why wouldn't a hospital, which obviously has massive compliance and risk departments, evaluate a medical algorithm? It comes down to how it arrived. The root cause was that the model arrived bundled into their existing electronic health record system.
Yeah, it was pushed through as a software update. So it was never separately identified and entered into an internal model inventory. And because it wasn't on an inventory, the governance teams didn't even know they needed to evaluate it.
Because it just looked like regular software. Exactly. It ran as software, not as a governed model.
And this is why a weak or absent inventory is considered a structural finding by an examiner. It basically means the organization cannot demonstrate that it even knows the shape of its own operational risk? Precisely. Ernest could not answer the regulators questions because no single current owned list of its models existed anywhere in the building.
OK, so the mandate is clear. We need the list. But to survive that examiner's opening question, I'm assuming handing over just, you know, a basic Excel spreadsheet with some project names is not going to cut it.
Not even close. A spreadsheet with names like MarketingBotV2 and CreditScore is practically useless. You must build a living register where every single entry provides absolute crystalline clarity to a stranger reading it.
A regulator should be able to read a single row and understand exactly what the system does, who is responsible for it, and how dangerous it is. Which brings us to the actual architecture of this artifact and our second core principle. Let's state it plainly.
A defensible inventory row has seven fields. Seven specific fields. No more, no less.
And if you are building this for your enterprise, you need to enforce these with rigorous precision. OK, let's break down the anatomy of this row, starting with the first field. The name and version identifier.
Right. Now, I know from looking at software development inside big companies, naming conventions are usually a disaster. You've got project code names that change every quarter, chat channel nicknames.
I'm assuming that doesn't fly here. It absolutely does not fly. The name must be stable across the enterprise.
But crucially, the identifier must include the specific version of the model running in production today. Why is the version number so critical? Well, this is where organizations fail constantly. A vendor can fundamentally change a model's underlying behavior.
They can swap out the training data or alter the logic through a routine update all without ever renaming the core product. Oh, wow. Yeah.
Right. So if your inventory does not track the specific version number, your inventory cannot tell a reader whether the validation documents you have on file actually describe the software that is making decisions right now. That makes total sense.
If I test version 1.0, and the vendor silently pushes version 1.5, my safety audit for 1.0 is legally and operationally meaningless. Exactly. Okay.
So that's the first field, known aversion. Let's move to the second field, which is the owner. And I really want to push back on how organizations typically handle this, because I see this in corporate documentation all the time.
Oh, I'm sure you do. Yeah. You see a row where the owner is listed as the data science department or, you know, the risk team or retail banking.
Why can't an entire department own a model? After all, it takes a whole team to build and maintain these systems. It takes a team to build it, absolutely. But teams diffuse accountability.
When a governance document lists a department as the owner, it is essentially a psychological trick to ensure nobody is individually on the hook when things go wrong. Let me try to picture this practically. It feels a bit like a commercial kitchen.
If the executive chef says the entire kitchen staff is responsible for checking the internal temperature of the chicken, then nobody actually checks it. Everyone assumes someone else grabbed the thermometer. When everyone owns it, no one owns it.
That analogy is exactly right. If the chicken is undercooked and a customer gets sick, you cannot fire the abstract concept of a department. Right.
You need to know the specific sous chef who was stationed at the grill and was supposed to probe that exact cut of meat. In enterprise governance, a named owner is the specific individual person an examiner calls when a row needs defending. So does this person need to be the lead engineer, like the person writing the Python code? No.
And in fact, it usually shouldn't be. The owner is the executive or senior manager who is accountable for the business outcomes of that model. They are the person who signs the tier assignment.
They're the one who has the political authority to turn the model off if it starts misbehaving. If your inventory lists as a department, it is declaring a silent gap in accountability. Got it.
Field one, name inversion. Field two, a specifically named owner. Let's look at field three, the purpose.
I imagine this is another area where corporate speak creeps in. Oh, it is the most common place for marketing fluff. The purpose field needs to be one or two highly specific sentences stating exactly what decision the system makes and for whom.
Can you give an example? Sure. Something like, this model provides automated approval and denial recommendations for unsecured personal loans under $50,000. Clear and mechanical.
Contrast that with the marketing fluff. What does a bad purpose statement look like? A bad one reads like a Silicon Valley pitch deck. This system leverages AI to drive better customer outcomes and optimize lending synergies.
Right. Which means nothing. Exactly.
And if the purpose is vague, the system cannot be accurately scored for risk. The examiner needs to know the precise mechanical function of the model to understand its potential blast radius. Which perfectly transitions us into fields four and five, which measure that exact blast radius.
Right. Field four is the materiality score, which is the cost of the system being wrong. And field five is the complexity score, which measures how hard it is to know why the system produced a specific output.
Right. And both of these are scored on the scale of one to three. And we're going to dive deeply into the mechanical engine of those scores in just a moment, because getting those numbers wrong is catastrophic.
Definitely. But before we do, we have to finish the row. Field six is the tier and validation status.
The tier itself is mathematically derived from the materiality and complexity scores, which we'll cover shortly. But the validation status is critical, right? Yeah. It cannot just be a binary checkbox that says, yes, validated.
No, it can't. Why not? If the compliance team checked it, why isn't a checkmark enough? Because validated means very different things, depending on the risk level. The status field must state the specific dated kind of review it has received.
Like what? Like, did it undergo a full independent evaluation by a third party? Or was it just a documented self-assessment by the internal team? Or has nothing been done at all yet? An unqualified validated on an inventory is just an empty claim. An examiner wants to see the specific standard of evidence. And that brings us to the final piece of the architecture, field seven, the next review date.
Yes. Every single row expires. Every single one.
Every single one. An AI system is not a static barrage. It is interacting with shifting data, shifting consumer behaviors, and shifting markets.
Without a next review date, the inventory is just a snapshot that immediately begins to decay the second you hit save. It's obsolete almost instantly. Exactly.
If you do not have review dates mandating when the system needs to be re-audited, you're holding a historical document, not a living governance tool. Okay. Let's step back and summarize this architecture for the listener.
To survive the examiner, your inventory isn't a list. It is a matrix. Seven fields for every single system.
Name and version, a named owner, a specific purpose, a materiality score, a complexity score, the resulting tier and validation status, and an expiration date. That is the defensible row. And the engine that drives that entire row is the tiering math.
We have established that the row must contain materiality and complexity, but how do we guarantee this isn't just arbitrary guesswork? Right. Because if I'm a project manager and my bonus is tied to launching this new AI feature by Q3, I'm highly incentivized to grade my own homework favorably. Of course you are.
I might look at those one-to-three scales and say, yeah, you know, this isn't so risky. Let's give it a low score so I can skip the six-month safety audit. How do we stop that? We stop it by relying on a mechanical engine that completely removes subjective optimism.
And that is our next core principle. Write this down. The tiering rule is materiality by complexity summed.
Materiality by complexity summed. Yes. It is a non-gamable calculation.
Let us dissect materiality first. So materiality is basically asking, what does it cost the organization, a customer, or the public if this system is wrong? Precisely. And the strict non-negotiable rule here is that it must be scored on the worst realistic outcome.
Not the average case and not the most likely case. The worst realistic outcome. Okay, let's walk through the one-to-three scale.
Right. A score of one means minor consequence. If the system produces a wrong output, it costs a little bit of time.
It is an easily corrected error with absolutely no legal or financial exposure. Give me an example of a materiality one system in a typical corporate environment. Think of an internal AI system that reads employee calendars to estimate meeting room capacity.
Oh, okay. If the algorithm hallucinates and recommends a room that is off by two seats, the employees have to stand, or maybe they just grab an extra chair from the hallway. It's annoying, but the harm is minor and instantly correctable.
That is a one. Exactly. Okay, so step it up.
What is a score of two? A score of two is moderate consequence. The wrong output has a real measurable cost, financial or reputational, but the harm is boundable. Boundable meaning? Meaning it is reversible or it is virtually guaranteed to be caught downstream before permanent damage occurs.
So maybe like a marketing recommendation engine. That's a great example. Let's say a retail bank uses an algorithm to suggest which credit card to pitch a customer in an email.
If the system goes haywire and suggests a premium travel card to a college student who clearly won't qualify, well, that's a bad customer experience. Right. It wastes marketing spend, but it doesn't fundamentally destroy the student's life.
Exactly. It's a failure. It costs the bank money, but it is a boundable harm.
Now contrast that with a score of three, which is severe consequence. And what does severe look like? This represents significant financial loss, legal liability, discriminatory harm, physical safety risk, or lasting reputational damage. The key differentiator for a three is that the harm lands on a real person and it may not be reversible.
Which brings us right back to our opening story. Earnest operations credit underwriting model. Yes.
If that algorithm incorrectly denies a qualified applicant alone, because they went to a historically black college, that applicant misses out on buying a home or funding their education. That is a severe harm. That is a three by definition.
It is. And I really want to address the most common mistake organizations make right here. Executives will look at a system like an underwriting model and say, yes, the worst case scenario is bad, but our model is incredibly accurate.
It has a clean track record. It behaves perfectly 99.9% of the time. Therefore, we are lowering the materiality score from a three to a two.
Let me jump in here with an analogy because the source material frames this perfectly. Lowering the materiality score because the model has a good track record is a massive logical flaw. It's completely backward.
It is like scoring a smoke detector. You do not score a smoke detector's importance based on the fact that your house hasn't caught fire in three years. The absence of an incident is not evidence that the worst case outcome won't occur.
That smoke detector analogy is vital to understanding governance. You score the fire, not the frequency. Score the fire, not the frequency.
I love that. A highly reliable underwriting model that behaves perfectly almost all the time is still a severe materiality system because that one fraction of a percent failure means a qualified individual is wrongfully denied credit. Right.
Lowering a materiality score due to a clean track record is a fatal governance error. It confuses the probability of failure with the consequence of failure. Materiality measures consequence.
Full stop. OK. I think we have materiality locked in.
Consequence of failure. Worst case scenario. Scored one to three.
Now let's look at the other half of the engine. Complexity. And this is completely independent of how the model is used, right? This is just about the math itself.
Correct. Complexity asks a structural question. How hard is it for a human to know why the system produced the exact output it produced? So trace this one to three scale for me.
What is a complexity one? A score of one is deterministic and interpretable. We are talking about fixed rules, lookup tables, simple arithmetic. So if you put x in, you get y out.
Yes. You put the exact same input in, you get the exact same output every single time. And a human being with a calculator can trace the math by hand.
The hard-coded rule at earnest that denied applicants without a green card. That is a complexity one rule. It is binary.
It is fully traceable. Wait, I want to clarify that. The green card knockout rule was highly discriminatory, but you are saying it is low complexity.
Do not confuse complexity with risk. A rule can be incredibly damaging and legally disastrous, but mechanically, very simple. Complexity just measures traceability.
Understood. OK, so what pushes a system to a complexity two? A score of two is statistical, but still largely interpretable. This includes traditionally trained models like a logistic regression or a standard credit scorecard.
Let me stop you there because logistic regression is a term that gets thrown around a lot in boardrooms, but I want to make sure we know exactly how it works under the hood. How does it differ mechanically from a simple rule? Well, in a simple rule, a human writes the logic, right? If credit score is below 600, deny. In a logistic regression, the machine looks at thousands of past loans and assigns mathematical weights to different variables to predict a probability.
It might decide that income gets a 40% weight and debt to income ratio gets a 60% weight. OK. It is more complex than a simple rule, but it is still a two because a data scientist can open up the model, look at the weights, and explain exactly why a specific person was denied.
The relationship between the inputs and outputs can be inspected and explained. Which brings us to the black box. Complexity three.
A score of three is opaque, adaptive, or generative. The reasoning cannot be fully traced. The behavior might shift as it continuously learns in production or it produces open-ended, non-deterministic output.
So this is where large language models, generative AI, and agentic systems sit. Firmly at a three. Let's look at how an LLM functions mechanically to understand why.
When an LLM drafts free text, it is not retrieving a pre-written answer from a database. Right, it's not searching. No.
It is predicting the next word or token in a high-dimensional mathematical space based on probabilities. Because it relies on probability, it is non-deterministic. Two identical prompts asked five minutes apart can produce entirely different text.
Which makes it impossible to fully audit. Exactly. The specific microscopic path from your input prompt to the final output resists clean human explanation.
Even the developers who built the model cannot always tell you why it chose one specific adjective over another. That fundamental lack of traceability makes it a complexity three. Okay, so let me synthesize the math for the listener because this is the engine that drives everything.
We take the materiality score one to three and we add it to the complexity score one to three. Yep, sum them up. If your sum is five or six, that system lands in tier one, critical.
If your sum is three or four, that system is tier two, significant. And if your sum is exactly two, it's tier three, limited. But a tier is just a label.
It's meaningless unless it changes what happens next in the organization. What are the actual operational demands placed on a tier one system versus a tier three? The demands escalate exponentially and that is deliberate. For a tier one critical system, so a sum of five or six, the requirements are heavy, expensive, and non-negotiable.
Before that model can be trusted in production, it requires a full, rigorous evaluation report. Crucially, it requires an independent challenge. Meaning a second reader.
Not just a second reader, a completely independent reviewer who is conceptually and managerially distinct from whoever built or evaluated the model. So it can't just be the person at the next desk. Definitely not.
You cannot have one data scientist build it and their deskmate audit it. Tier one also requires revalidation at least annually or immediately upon any material change to the data environment. And finally, it requires formal sign-off from a senior accountable owner who actually has the political authority to say no to the business unit and halt the launch if the audit fails.
That is a massive amount of friction. What about tier two? Tier two. Significant systems, so a sum of three or four, still require scrutiny, but the burden is lighter.
The business needs a documented self-assessment. They must write down the purpose, the limitations, and conduct a spot check of the outputs. Does it need the completely independent audit? No.
This can ideally be done by someone adjacent to the builder rather than a strictly independent external audit. The revalidation cycle stretches out to an 18 to 24 month cycle and a mid-level manager sign-off is acceptable. And tier three, the sum of two, minor consequence, purely deterministic.
Tier three limited systems require almost no friction. You just need basic documentation on the inventory, what it does, who owns it, and the math showing why it scored so low. The review cycle stretches out to three years.
The gap between tier one and tier three is huge. And I imagine that's the whole point, right? Exactly. The whole point of tiering is to concentrate your organization's scarcest resource expert validation effort where consequence actually lives.
Right. If you give every single model the exact same light review, you haven't tiered anything. You're just pretending to do governance.
Conversely, if you give every simple arithmetic script the tier one treatment, you will make your inventory unaffordable to maintain. You will bankrupt your compliance budget. Okay.
So we've established the math, materiality plus complexity. But, you know, the moment you introduce this framework to a business unit, the moment you tell a product manager that their shiny new AI tool is a tier one and requires a six month independent audit, they are going to look for a loophole. Oh, immediately.
Right. And once you understand that materiality is based on the worst case scenario, the most common objection rises. Someone in that governance meeting will inevitably cross their arms and say, but a human reviews the output before anything happens.
The machine doesn't make the final call. So the risk is low. And this brings us to a critical warning.
This is our next spine element and it is the single most common tiering error in the entire industry. A human in the loop does not lower materiality by itself. I really want to dig into this because intuitively a human in the loop feels like the ultimate safety net.
If an AI writes a loan rejection, but a human loan officer reads it and has to like actually click approve before it goes to the customer, why doesn't that human's presence mitigate the worst case scenario? Because the assumption that a human reviewer is a perfect attentive safety net is a dangerous psychological fallacy. Human review degrades under volume. The defense mechanism completely breaks down when you scale it.
Can you give an example of that? Let's look at clinical medical settings and a phenomenon known as alert fatigue. In a busy ICU, monitors are constantly beeping with minor warnings. The human brain cannot sustain a state of high alert indefinitely.
So clinicians neurologically adapt to ignore the noise. They just stop hearing it. Right.
And the exact same neurological mechanism that allows false medical alarms to slip past fatigued doctors is the mechanism that allows low quality AI outputs to slip past fatigued corporate reviewers. Let's make this concrete with a real world example from our source material. Let's imagine row six on a company's model inventory.
It's a generative customer reply drafting assistant. When a customer emails a complaint, the AI reads it and drafts a response for the customer service agent to send. Now, a junior compliance analyst looks at this system and says, it only drafts.
A human agent reads every single message and physically clicks the send button. Therefore, the worst case scenario is impossible. So the materiality is a one minor.
But let's apply a rule and examine the operational reality. What is the worst case output of that system? I guess the generative model hallucinates. Yes.
It drafts a false corporate policy. It makes an unfulfillable financial promise to the customer or uses discriminatory language. Now, factor in the volume.
That human agent is processing 200 tickets a shift. Right. They have 15 seconds to review each draft.
They're experiencing cognitive fatigue. If they accidentally hit send on that hallucinated promise, the company is now bound by it. That is a real, measurable, financial and reputational harm.
That makes the true materiality a two. And because it is a generative language model, the output is non-deterministic. So the complexity is a three.
Materiality two plus complexity three equals a sum of five. Which means this generative customer service bot, which the junior analyst thought was completely harmless, is mathematically a tier one critical system. Wow.
It requires the exact same level of scrutiny, independent challenge and senior sign off as a credit underwriting model. The analysts completely missed this because they scored the human's presence rather than scoring the model's potential harm. I'm trying to picture the mechanics of this human in the loop failure.
It feels exactly like hiring a lifeguard. OK, let's hear it. If I hire a lifeguard for one small backyard pool with three kids in it, that human is a genuine safeguard.
But if I tell that exact same lifeguard to sit in the same chair and watch five crowded churning wave pools at a massive water park simultaneously, the presence of the lifeguard is basically an illusion. That's spot on. The water is just as dangerous, if not more so, because people have a false sense of security.
When you are tiering an AI system, you don't score the lifeguard. You score the depth of the water. That's a perfect visualization.
At low volume, a human is a safeguard. At high volume, human review is a purely theoretical mitigation. You score the water.
OK, so the business unit lost the human in the loop argument. The system is tier one. Now they execute their next defensive maneuver to avoid scrutiny.
They point to the government regulators. Ah, yeah. And say, hold on, current federal banking regulation expressly ignores generative AI, so we don't have to tier it internally either.
We have to address this head on because it is a massive trap. This is our next spine element. Write this down.
A regulatory carve out is not evidence of low risk. Let's get into the specifics of why this happens. Let's look at the actual regulatory landscape.
For over a decade, banks relied on a piece of guidance called SR 11 to 7 for model risk management. But on April 17th, 2026, the Federal Reserve and the Office of the Controller of the Currency issued an updated framework called SR 26 to 2, which superseded the old guidance. SR 26 to establishes rigorous modern standards.
But if you read the fine print, it expressly excludes generated and agentic AI from its current scope. Why would they exclude it? The regulators labeled these technologies as novel and rapidly evolving, essentially signaling that they need more time to study them before issuing specific rules. And how do corporate governance teams misinterpret that? They read the phrase excluded from scope and translate it in their heads to low risk.
They think if the Fed isn't worried about it yet, we don't need to govern it. But that's not what the Fed is saying at all. Exactly.
The regulation is merely stating the jurisdictional reach of one specific rule at one specific moment in time. It is a legal boundary, not a scientific finding about the underlying danger of the systems. So if your organization waits for a finalized federal rule to require you to tier a generative customer service assistant, you are actively choosing to run an untiered, high consequence system in the dark.
Exactly. And this brings us right back to the devastating lesson of earnest operations. Earnest was not penalized by a banking model risk supervisor, citing SR 11-7 or SR 26-2.
Right. They were sued by the Massachusetts Attorney General. Exactly.
They were penalized by a state attorney general who didn't care about banking regulations. The attorney general was wielding broad consumer protection and fair lending law. General law will reach your system and penalize you for harm long before AI-specific sector rules have caught up.
So if you only build your internal inventory to satisfy one specific regulator's checklist, you are completely exposed when a different legal door opens. Completely exposed. That is why the internal discipline of tiering materiality by complexity must operate entirely independent of external regulatory mandates.
You tier a generative system as critical because it is a Complexity 3 system by construction and because human review degrades under volume. You do it because it is dangerous, not because it is illegal. Well said.
The organizations that will survive the next wave of enforcement are the ones that tier these systems on their own operational judgment rather than waiting for a subpoena to tell them to do so. Precisely. Now let's unpack another major operational vulnerability.
We've talked extensively about how to tier the models you know about, but what about the ones you don't? The unknowns. Right. An inventory with silent gaps is vastly worse than having no inventory at all because a complete-looking list invites unearned executive confidence.
We have to talk about shadow models and undocumented overrides. Let's define the term first. A shadow model is any system making a material decision in your company that never got entered into the official inventory.
How do these things sneak in? They typically enter an organization through one of three doors. Door one is the most common, a spreadsheet with an embedded regression formula built by a clever business analyst to speed up their workflow. Ah, the classic spreadsheet.
Yep. Because a formalized data scientist didn't build it, nobody uses the word model, but it is absolutely driving financial decisions. Door two is presurement.
A vendor sells the company a massive software platform and bundled silently inside that platform is a proprietary scoring engine. And the procurement team just bought it as a software license. Exactly.
The procurement team bought it out of the software budget, so it never triggered a model risk review. What's the third door? Door three is the successful pilot. A team runs a proof-of-concept AI tool.
It works incredibly well. It works so well that people just start relying on it daily. It quietly transitions into production without anyone ever circling back to formalize its governance status.
You know, I know that for organizations with high technical maturity, there is a fourth, more advanced door for discovery. Oh. Yeah, they don't just ask managers what models they're running.
They run automated scans of their own code repositories and cloud billing records to catch AI artifacts that are drawing a compute budget but have no declared owner on the inventory. That's a very robust way to find them. OK, so those are shadow model systems you don't know exist.
Right. But let's look at the second failure pattern, which is the undocumented override. Yeah.
This takes us back to the explicit finding from the earnest settlement. Right. The human underwriters bypassed the automated underwriting model frequently, but they left no documentation as to why.
Now, let me play devil's advocate here. Go for it. If I'm an underwriter and the AI denies a loan, but I look at the file and realize the AI is missing vital context, so I override it and approve the loan.
And let's say the bearer pays it back perfectly. I was right. I saved the bank money.
Why does the compliance team care if I didn't log a paragraph about my reasoning? Because the absence of a record breaks the governance chain entirely. I'll use an analogy you might appreciate. Think about the autopilot system on a commercial jet.
Right. If a human pilot deviates from the autopilot because they spot turbulent weather ahead, that is excellent human judgment. That is exactly why they are in the cockpit.
But if the pilot turns off the flight data recorder when they make that maneuver, they blind the entire safety system. Nobody back at headquarters can ever reconstruct the pattern of why the autopilot was failing to see the turbulence in the first place. That makes perfect sense.
Yeah. An unlogged correct decision still blinds the governance system. Exactly.
Overrides are not inherently bad. In fact, human judgment is supposed to be the ultimate safeguard for a tier one system. But if you don't have a structured log of those overrides, you cannot audit whether they cluster around a specific demographic pattern that reveals hidden bias.
Like what happened at Ernest. Right. You cannot prove to an examiner how often the model's formal, mathematically validated decision was actually the decision executed in the real world.
A model whose decision is overridden a third of the time with no written record is a model that, from an audit perspective, barely exists. The organization is just guessing at how decisions are being made. So if I'm building this internal control, what does a defensible logging standard actually look like? What fields do I need to capture every time a human overrides the AI? A defensible log requires six specific pieces of data.
One, who overrode it? We need a named individual. Two, what is their specific authorization role to do so? Are they a junior clerk or a senior VP? Right. Three, the exact timestamp of when it happened.
Four, the original output of the model. So what did the machine want to do? Five, a specific case ID tying the override to the exact applicant or transaction. And six, crucially, a coded reason category.
Explain that last one. Why a coded category instead of just letting the underwriter type a note? Because you cannot systematically audit free text at scale. If you give someone a free text box, they will just type felt wrong or management discretion.
Which is impossible to analyze. Exactly. You need a structured drop down category like alternative income source verified or data entry error corrected.
So the governance team can pull a quarterly report and see exactly where the algorithm is consistently falling short. OK, so as we catalog all these systems, hunt down the shadow models and lock down or override logs, organizations inevitably crash into the reality of international loss. That they do.
The most prominent piece of legislation right now is the European Union's AI Act. When an executive sees the EU AI Act, which has its own very specific risk classifications, their immediate instinct is to save time. They say, let's just copy the EU AI Act's risk tiers and paste them directly into our internal inventory.
But that brings us to our next spine element. This is a crucial distinction. The inventory is not a legal risk classification.
It is a fundamental error to conflate the two. You cannot just adopt the EU AI Act's risk annexes and call it your internal tiering system. They serve entirely different purposes.
The EU AI Act answers a legal question about statutory obligations and market access within Europe. Your internal tier, which we defined as materiality by complexity, answers an operational question about internal validation depth and resource allocation. They can completely disagree with each other.
Walk me through a scenario where they disagree. How can a system be safe under the law but dangerous internally? Let's say you build a highly complex algorithmic trading model. Legally, a purely financial trading system might fall outside the specific high-risk annexes defined by the EU AI Act because those annexes focus heavily on fundamental human rights, employment, and biometric surveillance.
OK, so legally, it's not high risk. Right. From a strict legal compliance standpoint under that specific statute, it is not high risk.
But internally, that trading model has massive financial variance and extreme opacity. If it hallucinates, you lose $50 million in a millisecond. Internally, based on materiality and complexity, it scores out as a Tier 1 critical system.
So if I just blindly align my internal inventory to the legal statute, I am essentially declaring my trading algorithm to be low risk. Exactly. You are effectively outsourcing your company's operational risk judgment to a group of legislators who wrote a law that might not even cover your specific market, your geography, or your unique business use cases.
Wow. Legal risk and operational risk overlap, but they are not the same map. You must run your own materiality and complexity scoring independently of any statutory classification.
Getting this tiering right is a high-wire act because erring in either direction hurts the organization in entirely different ways. Let's look at the symmetries of error here. Let's do it.
Under-tiering is the obvious risk. This is what the examiners are hunting for. You take a critical Tier 1 system, the business unit pressures the compliance team, you massage the scores down to Tier 2 or 3, and you hide a massive operational risk behind a light, superficial review.
That is how you end up in the headlines with a multi-million dollar settlement like Ernest Operations. But there is a second failure mode that doesn't make the headlines. Over-tiering is the quiet failure.
Over-tiering destroys governance from the inside. When executives get scared of regulatory fines, their instinct is to label absolutely everything Tier 1 just to be safe. But if you label everything Tier 1, your inventory becomes completely unaffordable to maintain.
It's a resource allocation problem. I think of it like a municipal fire department. If the city council mandates that a burnt piece of toast in a toaster and a five-story collapsing building on fire are both classified as five alarm emergencies, requiring the maximum response, nothing actually gets safer.
The fire department just runs out of fire trucks. They exhaust their personnel responding to the burnt toast, and when the building collapses, they have no resources left to deploy. That is precisely what happens to compliance teams.
Over-tiering dilutes the entire meaning of the critical tier. If you have an inventory with 50 systems labeled Tier 1, but only 8 of them genuinely warrant that level of scrutiny, you will inevitably exhaust your validation staff. Because they're tied up with minor issues.
Exactly. You will spend your scarcest expert resources auditing predictable, low-complexity arithmetic models while genuinely opaque, high-risk systems get rushed. Ultimately, you will under-deliver on your internal promises of annual independent validation because you simply don't have the headcount.
So how do we balance this? How do we prevent the business units from under-tiering to save time and prevent the executives from over-tiering out of fear? To prevent these errors, you have to restructure the burden of ownership. We established back in Field 2 that every individual row needs a named model owner, but the inventory itself needs a distinct owner. Like a master owner.
Yes. The model owner and the inventory owner must be two entirely different roles held by different people. A model owner defends one row.
The inventory owner defends the structural integrity of the entire list. Because if a model owner is also the inventory owner, they are greeting their own homework. They are naturally incentivized to under-report the severity of their own models to avoid the grueling cost and delay of Tier 1 validation.
Exactly. The inventory owner's metric of success is the completeness, currency, and accuracy of the entire list. They are the ones asking the hard questions, running the shadow model discovery sweeps, and ensuring the tiering math is applied ruthlessly and consistently across every department, regardless of internal politics.
OK, let's assume an organization has done all of this perfectly. They have the perfect list. It is flawlessly tiered using the math.
They have independent owners and the logs are locked down. How do they stop this beautiful document from decaying into irrelevance the exact moment they publish it? You enforce strict maintenance disciplines. We touched on the next review date earlier.
A Tier 1 row whose review date has passed with no new validation document on file is not treated as a minor clerical documentation lapse. What is it treated as? From a governance perspective, that is a system operating in live production on expired evidence of safety. It must be flagged immediately as a critical operational risk.
What about systems that the company stops using? Do we just hit delete on the row to keep the list clean? Never. Second discipline. Retirement versus deletion.
You never delete a row from a governance inventory. You mark it as retired, you stamp it with a date, and you log a reason for its deprecation. Why keep it? You don't shred a patient's medical chart the day they are discharged from the hospital.
If an examiner launches an investigation today into a lending decision your model made two years ago, that row, its scores, and its validation history must be instantly retrievable. Makes total sense. And the third discipline.
Version reopening. If a vendor pushes a material update to the model's logic, that action reopens the row immediately, resetting the validation clock regardless of what the next review date says. I see.
So it's like a car's scheduled maintenance. The manual says to check the alignment every 30,000 miles, but if you get into a severe accident at 10,000 miles, the maintenance schedule changes immediately. You check the alignment now.
Exactly right. I want to tie all of these abstract mechanisms, the fields, the math, the maintenance, the carve-outs, together by stepping into an immersive, practical scenario. Let's walk through the exact executive actions required to build this artifact from absolute scratch.
Let's do it. Let's look at a hypothetical risk officer. We'll call him Elmer.
Elmer runs model risk oversight at Larkspur Community Bank, a mid-sized regional lender. It is a Monday morning, and his compliance director forwards him a news alert about the earnest operations settlement. Oh boy.
Yeah. The director attaches a single terrifying question. Could this happen here? Tell me by Friday.
Now, Elmer knows his bank doesn't have a centralized seven-field inventory. It's a mess. He has a decentralized spreadsheet from the data science team.
He has a disjointed procurement list of vendor features that mention the word AI, and he has a scattered folder of validation memos on a shared drive. He has four days. How does he start from zero without panicking? The cardinal rule for starting from zero is to begin with the highest stakes business function first.
Do not try to map the whole company at once. For Larkspur, their highest stakes function is credit decisions. Elmer knows he won't get a perfect list by Friday, and that is OK.
He accepts an imperfect, honest first pass, and he builds the discovery process alongside the list. He sets up the seven columns. On Tuesday, Elmer convenes his working group.
The head of small business lending immediately pushes back. He says, Elmer, why are we doing this? Our flagship underwriting model is already validated annually by an external auditor. We are perfectly safe.
And Elmer has to explain the structural reality to him. He explains that validating one flagship model perfectly doesn't answer the examiner's question about what else is running in the bank unseen. He asks the lending head about the custom Excel sheets the junior analysts use to prescreen applicants before they even reach the flagship model.
Which are shadow models. Exactly. Those spreadsheets are shadow models.
They go on the list. On Wednesday, Elmer meets with a customer service director. They are discussing a new generative AI draft reply tool the bank rolled out to handle customer complaints.
The director uses the classic human-in-the-loop defense. Elmer, this tool doesn't decide anything. It just drafts text.
A human agent reads every single draft before hitting send. The risk is minor. This is where Elmer has to force the issue.
He confronts the director directly. He doesn't look at the theoretical workflow. He looks at the operational reality.
He walks the director through the reality of a busy Tuesday shift. 200 complaint tickets. One tired agent.
15 seconds to process each draft. He asks the director, are they actually reading every word? Or are they just scanning for tone? Right. Because if that generative assistant hallucinates a false policy regarding fee waivers, human review will not reliably catch it at that volume.
It is a mathematical certainty that an error will slip through. Exactly. Elmer explains that a hallucinated fee waiver is a moderate financial harm.
That is a materiality too. Because it is a generative, non-deterministic system, it is a complexity three. He sums them up.
So it's a five. Yep. He forces the customer service director to acknowledge that this chat tool, which they thought was just a minor efficiency software, belongs in tier one.
It requires the exact same independent challenge as the underwriting models. By Friday morning, Elmer delivers his artifact to the compliance director. It is imperfect, but it is deeply defensible.
It lists the known models. It scores them honestly using the non-gamable math. It assigns a tier based strictly on that sum, independent of what the business unit wanted.
And crucially, he includes a final paragraph detailing exactly what areas of the bank they haven't checked yet. And that final paragraph is brilliant governance. Presenting an honest documented gap to an internal director or to an external examiner is a far stronger, more defensible position than falsely claiming perfection when you don't actually know.
It proves you understand the shape of your risk. OK, let's unpack the sheer magnitude of what we've covered today. We have built the definitive map that survives an examiner.
It requires seven distinct fields, completely devoid of marketing fluff. It is driven by a non-gamable tiering engine materiality by complexity. It is utterly immune to the human-in-the-loop fallacy, and it sees right through regulatory carve-out illusions.
And it requires the active, continuous disciplines of shadow model discovery and rigorous override logging to ensure that the map accurately reflects the reality of the territory. So what does this all mean for you, the professional listening right now? Here is your Monday morning move. First thing Monday, do not try to build the entire inventory for your enterprise.
You will overwhelm yourself and your team. Instead, run a targeted shadow model discovery step on yourself. Yes, start small but thorough.
Go directly to your procurement team and ask for a sweep of all vendor contracts containing AI labeled features. Pick just your single highest stakes business function. Create a defensible seven-field row for every system that touches that function, including those embedded Excel spreadsheets.
Accept an imperfect first pass, but write it down. Because the inventory is not merely a defensive shield to hold up against regulatory fines, it is the definitive unvarnished map of how your enterprise actually makes decisions today. Treat it with the reverence it deserves.
And think about this. What if building this inventory actually reveals that your company's true strategy is completely different from what the board thinks it is? Because if undocumented algorithms have been quietly making your highest stakes decisions for years, optimizing for metrics you never formally approved, well, then your inventory isn't just a map of risk. It's a map of your actual corporate strategy.
That is a terrifying but very real possibility. Remember, that x-ray machine of diagnostic clarity we wish we had for algorithmic risk doesn't exist out of the box. You have to build it yourself, row by row, score by score.
Without it, you are navigating muddy waters completely blind. Take the time, do the math, and build your map. See you next time on the Deep Dive.
Real cases
Example 1: Earnest Operations LLC and the inventory the settlement effectively ordered (United States, consumer lending). Covered in depth in Sections 3A through 3C. The governance lesson in one line: a regulator's remedy, requiring a written corporate governance system of fair lending testing, internal controls, and risk assessments for AI models, describes, in substance, exactly the inventory and tiering artifact this topic teaches, imposed after the fact rather than built before harm occurred.
Example 2: The undocumented override as its own finding (United States, consumer lending). Also drawn from the Earnest investigation, and worth isolating from the bias finding because it is a distinct failure mode. Underwriters who bypassed the model without recording it left the organization unable to answer how often, and for whom, the automated decision was actually the decision made (Section 3G). An inventory row's validation status is only as meaningful as the organization's ability to say the model's stated behavior is its actual behavior, and an unlogged override breaks that chain quietly.
Example 3: SR 26-2's own narrowed model definition (United States, banking regulation, standing instrument). SR 26-2 excludes simple arithmetic calculations from its own definition of a model, a regulator explicitly declining to apply Tier 1 style scrutiny to Row 9's kind of system. (see Topic 4.6) for the deep evaluation-report treatment of metrics and evidence; the point here is narrower: even a banking supervisor's own rule distinguishes complexity levels, which is the second axis this topic's tiering rule scores directly.
Example 4: A well-tiered generative deployment (illustrative, cross-sector). An organization deploying a generative assistant for internal knowledge search scores it for materiality (moderate, an employee could act on a wrong internal policy summary) and complexity (high, generative and non-deterministic), lands it in Tier 1 despite having no external customer contact, and builds a monitoring plan that samples outputs for factual drift rather than assuming a one-time evaluation covers a system whose behavior can change with each update. Contrast this with a deployment that scores the same category of system as Tier 3 because "it's just internal," conflating audience with materiality, and misses exactly the generative-complexity signal Section 3H names.
Example 5: An independent challenge that starts from the inventory, not the report (illustrative, cross-sector, pointer to Topic 4.7). (see Topic 4.7) owns the deep treatment of independent challenge; the inventory connection is that a Tier 1 designation is what triggers the requirement for a second reader in the first place. An organization that skips the inventory step and lets each team decide informally whether its own model needs independent review will systematically under-challenge the models whose owners have the least appetite for scrutiny, which correlates poorly with which models actually carry the most risk.
Example 6: A conformity file built from a current inventory versus one built from memory (illustrative, cross-sector, pointer to Module 5). (see Topic 5.6) owns the deep treatment of the conformity file; here the point is that an organization assembling evidence for a high-risk system under the EU AI Act moves fastest when its inventory already names the system, its tier, its validation status, and its owner, because the conformity file is largely an extraction from an inventory that was already current, not a fresh research project undertaken under deadline pressure.
Example 7: The Epic Sepsis Model, read as a missing inventory row (United States, healthcare, pointer to Topic 4.6). (see Topic 4.6) owns the deep treatment of the evaluation-report failure in this case; read through this topic's lens, the deploying hospitals' deeper problem was that a system arrived bundled into their electronic health record and was never separately entered into a model inventory with an owner and a tier at all. It ran for years as software, not as a governed model, and that categorization, more than any single missing metric, is why nobody assigned it the validation depth its materiality (a life-safety decision aid) and complexity (a proprietary, opaque scoring model) should have triggered from day one.
Example 8: A merger's shadow-model discovery (illustrative, cross-sector). During due diligence for an acquisition, a buyer's governance team runs the discovery process from Section 3G against the target company and finds an AI-driven pricing tool embedded in a licensed vendor platform that never appeared on the target's own inventory, because procurement had classified the purchase as "software," not "a model." The buyer scores it under its own materiality-and-complexity rule, lands it at Tier 1 given its direct revenue impact and opaque logic, and requires it added to the combined inventory with an owner and a validation plan before close. The lesson generalizes past mergers: any organization's discovery process should periodically re-run the same procurement sweep on itself, not only when a transaction forces the question.
Example 9: A regulator's document request answered from a current inventory (illustrative, cross-sector). An examiner requests, on short notice, a list of every automated system involved in a specific business line's customer-facing decisions. An organization with a current, tiered inventory filters its existing rows by purpose and produces a defensible answer within a day, each row already carrying an owner, a tier, and a validation status the examiner can follow up on individually. An organization without one spends the same window reconstructing the list from memory, emails, and hurried interviews, arriving at an answer that is slower, less complete, and, because it was assembled under pressure specifically to satisfy the request, far less credible than a list that existed before anyone asked. The difference between the two organizations is not the quality of their models; it is whether the inventory discipline in Section 3I was already running.
Example 10: An enterprise risk register mistaken for a model inventory (illustrative, cross-sector). An organization's enterprise risk team maintains a single line item, "AI and automation risk," inside its broader risk register, reviewed annually alongside cyber risk, market risk, and operational risk more generally. When a new compliance hire asks to see the model inventory, this line item is what gets produced. It names no individual systems, no owners, no materiality or complexity scores, and no tiers, because it was built at the granularity an enterprise risk function normally works at, not at the per-model granularity Section 3B requires. The gap surfaces only when a regulator asks a question the line item cannot answer, such as which specific systems make automated credit decisions, and the organization discovers it has to build the real inventory from scratch under exactly the deadline pressure Example 9 warns against.
Example 11: A first-year inventory that surprised its own builders (illustrative, cross-sector). An organization building its first formal model inventory, following the Section 3N approach of starting with its highest-materiality function, expects to find perhaps a dozen systems across the business. The procurement sweep alone surfaces thirty-one AI-labeled features bundled into existing vendor contracts, most never separately reviewed because they arrived as part of a larger platform purchase. Rather than treating this as evidence the process went wrong, the inventory owner treats it as the process working exactly as intended: the honest count was always thirty-plus, and the organization simply had not looked until the discovery process from Section 3G gave it a reason to.
Where people go wrong
- "We validated our most important model, so we're covered." One strong evaluation report proves one system was checked. It says nothing about the systems the organization did not mention. Earnest's failure was not a single bad model; it was the inability to produce a complete, current account of what it ran at all. An inventory, not a report, is what an examiner asks for first.
- "A human reviews the output, so the materiality is low." Section 3D and the Section 5 scenario both name this directly. A human in the loop changes what kind of monitoring is appropriate; it does not by itself lower the harm a wrong output could still cause, especially at volume, where review quality degrades the same way clinician attention degrades under a flood of false alarms. (see Topic 4.6) Score materiality on the worst realistic outcome, not on who is nominally watching.
- "If it's not called a model, it doesn't belong on the model inventory." A spreadsheet with an embedded regression, a vendor feature bundled into a larger platform, and a pilot that quietly became production are all shadow models by function, regardless of what anyone calls them. The inventory tracks decisions and outputs, not job titles or department budgets.
- "Our generative assistant isn't covered by SR 26-2, so it doesn't need to be tiered." SR 26-2 expressly excludes generative and agentic AI from its own scope, but that is a statement about the reach of one regulation, not a statement about risk. Section 3H is explicit: the discipline transfers by practice even where the current regulatory authority does not yet require it, and waiting for a mandate is choosing exposure, not avoiding it.
- "An override is fine as long as a person made the call." An override without a record is the specific failure the Earnest investigation surfaced. The problem is not that humans overrode the model; it is that nobody could later say how often, for whom, or why, which makes the model's stated behavior and its actual behavior two different, unreconciled things.
- "Tiering everything the same way is fairer." Giving every model light review under-protects the systems that carry real consequence. Giving every model full Tier 1 review makes the inventory too expensive to maintain, which predicts decay: rows go unreviewed, and new systems get added off the books because doing it properly feels too costly. Tiering exists so Tier 1 rigor stays affordable by being reserved for the models that earn it.
- "The inventory is done once it's written." An inventory with no review-date discipline is a snapshot the organization will treat as current long after it stops being true. Overdue Tier 1 rows are systems running on expired evidence, the same failure this program names for evaluation reports with no review-and-revoke trigger. (see Topic 4.6)
- "Deleting a retired model's row keeps the list clean." Deletion erases the history an examiner, or an internal investigation, might need later, for instance to understand what logic governed a decision made two years ago. Mark a row retired, with a date and a reason, and keep it.
- "A vendor's bundled model doesn't need our own tiering, since the vendor already validated it." A vendor's validation is a claim about the vendor's testing, on the vendor's terms, the same population-transfer problem Topic 4.6 names for a vendor's marketed benchmark. (see Topic 4.6) Your inventory still scores materiality and complexity for how the system is used in your organization, and your validation status still states what you, not the vendor, independently confirmed.
- "Our EU AI Act risk classification already covers this." (see Topic 5.3) The Act's high-risk category and this topic's materiality-and-complexity tier answer different questions, one legal, one operational, and they can disagree. A system outside the Act's high-risk annexes can still be your organization's Tier 1 by this rule, and a system inside the Act's high-risk category still needs your own materiality and complexity scoring to decide its internal validation depth, because the Act tells you your legal obligations, not how much internal scrutiny is prudent.
- "A blended risk score is easier to read than two separate numbers." A single blended score hides exactly the reasoning a hostile reader will pull on: whether materiality was scored down because a human reviews the output, or whether complexity was scored down because a system is "mostly" deterministic. Two visible numbers, summed by a stated rule, let a reader check the reasoning; one blended number asks them to trust it.
- "Tiering is a one-time exercise at intake." A system's materiality and complexity are properties of what it does today, not what it was approved to do at launch. A model retrained, repurposed, or connected to a new decision workflow needs its scores revisited, the same version-reopening discipline Section 3I applies to validation status applies to the tier itself.
- "A high Tier 1 count means we did something wrong." Section 3L names this directly: the tier distribution is an output of the rule applied honestly, not a target with a preferred shape. An organization whose core function is lending, fraud detection, or another severe-consequence activity should expect a Tier 1 heavy inventory, and reshaping the distribution to look more comfortable is the identical error to under-scoring a single row to avoid its true tier.
- "The inventory is one team's job, so no single owner is needed for the whole thing." Without a distinct inventory owner, per Section 3M, different teams apply the tiering rule inconsistently, review dates slip in whichever team is busiest, and cross-functional systems that do not obviously belong to any one team are the most likely to fall through entirely. The inventory needs the same kind of single, named accountability the rows themselves require.
Questions people ask
- What is model inventory?
- A living register of every model and automated decision system an organization operates, with each row carrying a name and version, an owner, a purpose, a materiality score, a complexity score, a resulting tier, a validation status, and a next review date. The artifact an examiner asks for before any single system's evaluation report.
- What is materiality score?
- A 1 to 3 rating of what it costs the organization, a customer, or the public if a given model is wrong, scored on the worst realistic outcome rather than the average case, and never lowered solely because a human reviews the output.
- What is complexity score?
- A 1 to 3 rating of how traceable a model's reasoning is and how much its output can vary without a human choosing it to vary, from a fixed deterministic calculation (1) to an opaque, adaptive, or generative system (3).
- What is tier (Critical, Significant, Limited)?
- The three-level classification produced by summing the materiality and complexity scores (5 to 6, 3 to 4, or 2), which determines a model's required validation depth, revalidation cycle, and sign-off authority.
- What is shadow model?
- Any system making or materially shaping a decision that never entered the model inventory, commonly arriving through an unrecognized spreadsheet formula, a bundled vendor feature never separately reviewed, or a pilot that quietly became production.
Keep going
This lesson builds Model risk management and independent challenge, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.