Meet your systems: mapping the AI your own organization already runs
The short answer
You cannot govern what you cannot see
Every governance mechanism in this program (risk classification, evaluation, logging, incident response) operates on a known system. Applied to a system you never discovered, all of them do nothing. Discovery is not a preliminary to governance; it is the governance that is missing when things fail.
What you will be able to do
- Build an AI systems inventory for your own organization: a working list of the AI systems it already runs, with the fields a governor needs to act on each one.
- Distinguish what a system is marketed as from what it actually is, and record the real autonomy level, the human involvement, and the vendor-versus-built origin of each system.
- Find the AI you were not told about: embedded features inside tools you already bought, shadow AI that staff adopted on their own, and third-party AI hiding inside vendor products.
- Decide which fields belong in the inventory for governance (not for engineering) and why each one earns its place: the decision the system touches, the data it uses, who owns it, and how exposed the organization is if it fails.
- Triage the inventory by criticality so the systems that can hurt someone or the organization rise to the top, without pretending you can govern everything at once.
- Set up the inventory as a living artifact, not a one-time spreadsheet, so it stays true as systems are added, changed, and retired.
The lesson
For years, Amazon's Just Walk Out technology felt like the cleanest demonstration of artificial intelligence anyone had built. You picked up a sandwich and left. It was understood and copied as pure autonomous computer vision doing the work no human could.
Then, reporting revealed a different mechanism. Behind the cameras sat roughly a thousand people reviewing video to determine what was taken. The receipts were accurate, but the story was wrong.
AI quietly masked an offshore workforce, a massive privacy surface, and human-dependent economics. If you had been the executive responsible for governing that system, and all you had on paper was the label autonomous AI checkout, you would have been governing a fiction. You would have missed the real failure modes completely.
Before you write a single rule or policy, you have to know what you actually run. You cannot govern what you cannot see. Today, you are going to build the artifact that makes everything else in this program possible, your AI systems inventory.
Ask a leadership team how many AI systems their organization runs, and you will usually get a confident wrong number based on the two or three flagship models they officially procured. The reality sits underneath those flagships. The 2024 Work Trend Index captured the actual scale of AI.
75% of knowledge workers use generative AI, and 78% of them bring their own tools rather than using what the organization provided. This is not a rebellious minority. It is the mainstream.
The vast majority of AI use happens on tools the organization did not choose, vet, or even see on any org chart. To find these tools, you need a highly pragmatic boundary. For your inventory, an AI system is any statistical or learned model that influences a decision, a recommendation, a generation, or an action affecting your organization.
Drawing the boundary wide reveals a complete AI estate hides across seven distinct categories. Built models trained in-house, bought standalone products, embedded features inside your existing software, third-party AI inside vendor services, shadow AI adopted by individual staff, hybrid systems secretly run by people, and agentic tools executing multi-step actions. Because these systems evade standard oversight, active discovery must precede all other governance mechanisms.
You cannot classify, evaluate, or monitor a system you have never found. Finding these systems requires one primary discipline, distrust the label and record the mechanism. This AI system's inventory grid forces that discipline into a working record.
A tool sold as autonomous AI implies a clean, software-only risk profile, creating blind spots. But the actual mechanism might be a baseline model paired with a massive offshore human team handling exceptions. This gap occurs for three honest reasons.
Vendors simplify descriptions for a general audience. Operational drift causes human review layers to grow over time as edge cases pile up. Or occasionally, a vendor deliberately misrepresents their capabilities.
The inventory's job is not to prove intent. Its job is to force the operational truth into a documented field so the actual risks, the labor dependencies and the privacy services, can be governed on their true terms. When assembling this document, the most common failure is treating it like an engineering specification.
Explicitly exclude engineering metrics, like model architecture and training sizes. An inventory drowning in technical data governs nothing. The fields that earn their place are what a governor acts on, the decision touched, the data consumed, people affected, and the exact organizational exposure.
But the most critical element of the matrix is human. Every system must have a named, individual human owner accountable for its decisions. You do not assign a vague team or an engineering department.
An unowned system is an ungoverned system. Leaving a row explicitly marked unowned is a high-value finding. It strips away the dangerous illusion of oversight, acting as a blaring alarm that forces the entire organization to step up and assign real accountability.
Once your estate is mapped, you will likely have dozens of systems. Attempting to govern all of them equally guarantees you will govern none of them deeply. You must triage the list using a strict rule.
Criticality is based purely on consequence to people and the organization, not technological sophistication. Consider a basic keyword filter screening job applications. Despite being technically primitive, it decides who gets seen by a recruiter.
Touching human livelihoods and carrying severe discrimination exposure requires high criticality. Contrast that with a complex, multi-layered neural network used only to draft internal meeting summaries. Because this touches an internal draft that nobody relies on for consequential decisions or customer money, it ranks low.
The advanced architecture does not elevate the risk. Effective governance ranks the cut, not the steal. You allocate your deep attention strictly to the rows that can actually hurt someone.
A static snapshot of an AI estate decays within weeks. Vendors ship new features, and narrow pilots quietly transition into production to handle real customer cases. To survive, the inventory must be a living artifact.
You add the final columns, a date last verified stamp, and a refresh trigger. A refresh trigger ties mandatory updates to specific operational events, like the signing of a new vendor contract or a new software rollout. It catches the drift that no single review schedule can anticipate.
This ongoing discovery must address shadow AI. Treating unapproved tool usage as a punishable discipline problem only drives the behavior underground. It is a vital signal that official, safe tooling is missing.
Running an anonymous amnesty survey is the most effective way to find these shadow tools. It converts the invisible risk of confidential data leaving your network into a governable reality you can actually secure. By keeping the scope disciplined, this inventory acts as the index, not the encyclopedia of your governance program.
It stays short enough to finish and clear enough to act on. It points outward to the deeper, heavier documents, the risk registers, data processing records, and model cards owned by your engineering and legal teams. To build this, you must begin the unglamorous work of discovery immediately.
Start by pulling the last 12 months of software invoices from finance and following the money. Interview your business process owners and ask what helps you decide, without using the word AI. And actively audit your existing CRM and HR suites for newly enabled automated features.
The AI systems inventory is the foundational artifact of command. It turns a sprawling, invisible hazard into a named, triaged list you can finally act on.
The ideas, one by one
The estate is mostly hidden
Leadership reliably knows the flagship systems and misses the embedded features, shadow tools, and vendor-hidden AI that together are usually the majority of the estate. The 2024 Work Trend Index found 75 percent of knowledge workers already using generative AI at work and 78 percent of them bringing their own tools. If your count feels quick and complete, you found the tip, not the iceberg.
Record what a system actually is, not what it is called
The Amazon Just Walk Out case is the permanent reminder: a single label can conceal a workforce, a privacy surface, a cost structure, and a dependency. The "what it actually is" field, capturing real autonomy and human involvement, is the one most inventories omit and the one that most often turns out to matter.
Cast the net wide across seven places AI hides
Built, bought, embedded, third-party, shadow, hybrid, and agentic. An inventory that checks only "built" and "bought" misses most of what can hurt the organization. Over-include and prune later; do not under-include and get surprised.
The fields that earn their place are governance fields, not engineering fields
The decision a system touches, what it actually is, who owns it, who it affects, and what happens if it fails. Architecture and hyperparameters belong to the systems that own them. An inventory drowning in engineering detail is one nobody finishes and nobody reads.
Every system needs a named human owner
Ungoverned systems are almost always unowned systems. An unowned high-criticality system is one of the most important things an inventory can surface, because every later obligation needs a person to land on.
Triage by consequence, not sophistication
Criticality is about who is affected and what happens if the system is wrong, never about how impressive the technology is. A keyword filter that screens applicants outranks a dazzling model that rearranges an internal dashboard.
Shadow AI is a signal, not a crime
Most AI use in many organizations is on unofficial tools, which tells you official tooling is missing and real work is happening on unofficial rails. Surface it with amnesty, understand the data flowing through it, then decide what to sanction, replace, or stop. Punishment produces silence, not visibility.
The inventory is a living artifact
A one-time snapshot is stale within weeks. Give it an owner, a refresh trigger, and a date-last-verified stamp on every row, or it will be wrong exactly when a later module or a real incident reaches back for it.
Complete broadly, then govern in criticality order
A large estate cannot be deeply governed all at once. Give every system a row first (a thin row on a known system still governs; an unknown system does not), then spend your deepest attention on the high-criticality systems and let the low ones hold their place. Triage is the professional allocation of finite capacity, not an admission of incompleteness.
The inventory is an index, not an encyclopedia
It deliberately excludes the depth that belongs to the risk register, the data-processing record, and the model card. Its power is that it stays short enough to finish, current enough to trust, and clear enough to act on, while every deeper document hangs off it.
This is the artifact everything else hangs off
The Module 2 data audit, the Module 5 risk classification, the Module 12 successor briefing, and the capstone dossier all begin here. A weak inventory now is a weak everything later.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 2 of the podcast.
Read the full conversation
Welcome to the Deep Dive. So today, we are looking at a really comprehensive executive education module on mapping organizational AI. Yeah, and it's a critical topic right now.
It really is. Our mission for this Deep Dive for the next hour or so is transitioning your invisible, risky AI estate into a, well, a named, owned, and triaged list. Right, because right now, for most listeners, it's completely invisible.
Exactly. And the source material we're working through today, it treats this as a strict discipline of decisions. So think of the tone today as a Harvard Business Review meets a trusted mentor.
I like that, very practical. Yeah, we are dealing in credible, precise, actionable reality here. And understand why mapping the AI your own organization already runs is an absolute emergency.
I think we need to look at a system that was, well, for a few years, it was the envy of the retail world. Oh, you're talking about Amazon's just walk out technology? Yes. Let's look at the mechanics of this, because for years, walking out of an Amazon store without stopping to pay, it was marketed as the cleanest demonstration of artificial intelligence in retail.
It felt like magic, right? Totally. You walked in, grabbed a sandwich, put it in your bag, and just left. And the system relied on, you know, ceiling cameras and deep learning models tracking the items.
Then a receipt just arrived on your phone shortly after. Right, and Amazon called it just walk out. Yeah, and competitors were like scrambling to copy what they believed was pure computer vision operating at a scale that humans just couldn't match.
Well, and it was positioned as a masterpiece of autonomous AI checkout. Exactly. But from an executive governance standpoint, that label suggests a very specific risk profile.
I mean, if you are governing an autonomous AI, your primary concerns are server uptime, model accuracy, and maybe edge case misidentifications. Like confusing a turkey sandwich with a ham sandwich. Exactly, that's the risk you think you're managing.
But the operational reality didn't actually match that label at all. In early 2024, Bloomberg published reporting that summarized a very different narrative. It shattered the illusion, really.
It did. They revealed that behind those cameras sat roughly 1,000 workers in India reviewing video footage of shoppers to determine what had actually been taken. 1,000 workers.
Yeah, a significant portion of transactions were passing through this human review rather than being settled by the model alone. Now we should note, Amazon disputed the exact framing. They stated that the offshore team primarily trained the model and validated a minority of visits.
Right, they pushed back on the exact numbers. They did. But simultaneously, Amazon announced it was pulling just walk out from its larger U.S. fresh grocery stores moving to smart shopping carts instead.
Which is incredibly telling. And it brings us to the core lesson for the executive listening right now. Yeah, what's the takeaway here? Well, the customer experience, the receipt arriving on the phone, that was accurate.
But the label, Autonomous AI Checkout, it concealed a massive, complex operational mechanism. It was hiding the truth. Right, if you governed that system based on its marketing label, you governed a fiction.
Because if the reality involves, you know, a thousand people watching video feeds, the risk profile changes entirely. Oh, drastically. You now have a massive data privacy exposure involving live video of customers being routed across international borders.
Wow, yeah, privacy is a huge issue there. And you have offshore labor and workforce management obligations. Your cost structure scales with human headcount, not just server space, which means the economics of the system are fundamentally different.
Right, it's not just paying for AWS anymore. Exactly, if an executive doesn't know the mechanism, they're completely blind to the actual risks they're supposed to be managing. Which sets up the central thesis of our entire deep dive today.
We are covering six specific realities of AI governance, and the source material anchors everything on this single immovable axiom, you cannot govern what you cannot see. It is the plainest, most important rule in the field. It sounds obvious, but I guess it's not.
It's really not. I mean, discovery is not just some preliminary step to governance. It is the absolute prerequisite.
When things fail, your visibility dictates your defensibility. Okay, let's test that logic. Why is an undiscovered system completely ungovernable? Because, I mean, we have compliance teams, we have IT security protocols.
Right, you think you're protected. Yeah, so why do those standard corporate safety nets fail if the specific AI system isn't on a map? Because every subsequent governance mechanism, risk classification, bias testing, incident logging, compliance with new frameworks like the EU AI Act, they all operate exclusively on a known system. Oh.
Those are mathematical operations that require a target. Think about bias testing. You cannot audit an algorithm for discriminatory hiring practices if you don't even know the HR department turned on a resume screening plug-in last Tuesday.
Because it's just not on your radar. Exactly, or consider incident response. If an undiscovered customer service chatbot hallucinates a nonexistent return policy and offers it to a client, you can't reconstruct the prompt chain.
You can't roll back the model if the system wasn't logged in your architecture in the first place. If you can't see it, your governance is a null operation. A completely null operation.
And yet there is this pervasive confidence among leadership that they already possess this visibility. An executive might say, we know what AR we run, we went through a massive procurement cycle for it last year. Yeah, the source material calls that the confident wrong number.
The confident wrong number, I love that. Because leadership usually knows their flagship systems. Yeah.
They know the multi-million dollar data analytics platform they issued a press release about. Right, the big shiny things. Exactly, they know the custom generative AI tool they built for their coders.
Yeah. But they ignore everything operating below the waterline. So the analogy that comes to mind for me here is home electricity.
If you ask a homeowner how many things in their house draw power, they'll count the refrigerator, the oven, the television, the HVAC system. They'll hit the big appliances. Right, they'll stop at maybe 10 items.
But they completely miss the routers, the smart thermostats, the 30 different phone chargers left plugged into outlets, the microwave clock, you know, the standby power on their monitors. It's all the small things adding up. Yeah, the bulk of the dependency is hidden in mundane, everyday objects.
That is a highly accurate parallel. The flagship AI systems, those are the refrigerator and the television. But the actual risk surface of the organization is distributed across dozens of smaller, unmapped implementations.
So it's an iceberg situation. Yes, if an executive's mental list of their organization's AI is short, they haven't found the whole iceberg, only the tip. They're operating on a completely false sense of security.
And that false sense of security leads directly to the second core principle from the source material. The estate is mostly hidden. Mostly hidden, that's key.
And this isn't just a theoretical warning, by the way. To prove this isn't anecdotal, we have to look at the hard data from the Microsoft and LinkedIn 2024 Work Trend Index. That was a massive survey, wasn't it? It's huge.
It was a survey of 31,000 knowledge workers across 31 countries. And the findings show that 75% of knowledge workers are using generative AI at work. Which is a staggering number on its own.
But here is the critical data point for our discussion today. 78% of those users are bringing their own AI tools to work. It's BYOAI, bring your own AI.
Let the scale of that just sink in for a second. In most organizations, the vast majority of AI usage is occurring on platforms the organization did not select, did not vet for security, and absolutely cannot monitor. It's completely off the books.
Yes, this is shadow AI. It is entirely invisible on any organizational chart or IT audit. Okay, let's get into the mechanics of shadow AI and why it bypasses standard IT controls.
Because, I mean, we've dealt with shadow IT before. For decades. Yeah, like a decade ago, people were creating unauthorized Dropbox accounts to share large files because corporate email attachments were too small.
So how is shadow AI different or more dangerous than an employee using an unapproved cloud storage folder? The difference really lies in what the external system does with the data. With shadow IT, like a cloud folder, the risk was primarily unauthorized access or data leakage if the password was compromised. Right, someone hacks the folder.
Yeah, the folder itself was just a passive container. But shadow AI is an active participant. Oh, that makes sense.
When an employee takes a confidential supplier contract, proprietary source code, or a spreadsheet of customer complaints, and they paste it into a public generative AI model to summarize or debug it, they are not just storing data. They're giving it to the machine. Exactly.
They are feeding data into a model that may use those inputs to train its future iterations. Meaning that proprietary source code could theoretically be regurgitated to a competitor who prompts the same public model a few months later. Precisely.
You have data exfiltration without a trace. There are no logs, no audit trails, and zero accountability. That is terrifying.
And furthermore, shadow AI scales silently. Think about it. One employee figures out how to automate their weekly reporting using an unapproved web-based agent.
Right, they save three hours a week. So they share the trick with their team. Suddenly, an entire department's critical workflow is dependent on a consumer-grade tool running on a single employee's personal account.
Oh wow, yeah. If that tool goes down, or changes its terms of service, the department's productivity just collapses, and IT has absolutely no idea why. Okay, so if the risk is data exfiltration and fragile dependencies, the immediate executive instinct is usually punitive.
Always shut it down. Right, so shouldn't we just punish the employees? Like mandate a blanket ban on all unauthorized AI tools, block the IP addresses at the firewall, and just fire anyone who bypasses the controls? The source material is adamant that a punitive approach is a strategic failure. They push back strongly on this.
Shadow AI is a signal, not a crime. A signal of what? When 78% of your AI users are sourcing their own tools, they are communicating a massive operational failure on the part of the organization. Oh, I see.
They are screaming that the official tooling provided to them is either absent, terrible integrated, or just too slow to acquire through normal procurement channels. They're just trying to do their jobs. Exactly.
Employees are not using BYO AI to be malicious. They're using it because they are under pressure to be productive and the official rails are failing them. The text uses this really cool concept of a desire line or a desire path from urban planning.
Have you heard of this? Yeah, the dirt paths on college campuses. Right, when you look at a corporate park or a campus with a large grassy quad, you will often see a dirt path worn straight across the grass, totally bypassing the paved sidewalks that run in neat right angles. Because the paved sidewalks are inefficient.
Exactly, that dirt path is a desire line. It tells you the official sidewalk is in the wrong place. And if the administration just puts up a keep off the grass sign or finds people for walking on it, they haven't solved the underlying geometry problem.
People will just find a more hidden shortcut. Right, punishing walkers just makes them hide the shortcut. It's the exact same thing with shadow AI.
Banning the tools doesn't stop the behavior, it just drives the behavior further underground. Which makes it even harder to track. Yes, if you punish the employees, you destroy the psychological safety required to map your systems.
The governance move here is to offer amnesty. Amnesty, okay. You must surface the shadow usage by promising a blame-free environment.
Once you map these desire lines, then you can make intelligent decisions. So you bring it into the light first. Right, you might sanction the tool by purchasing an enterprise license that includes data privacy guarantees.
Or you might replace it by offering an in-house equivalent that performs the same function. Or I guess you could stop it if it's too dangerous. Exactly.
You might mandate stopping its use if the risk is truly uncontainable. But punishment produces silence. And silence is the enemy of governance.
So to break that silence, we need a systematic approach to finding where these systems actually live. And this is the third directive or pillar of the deep dive today. Cast the net wide across seven places AI hides.
Seven places. And before an organization can inventory its systems, it needs a functional definition of what belongs on the map. Which is tricky, right? Because everyone defines AI differently.
Very tricky. But the source text advocates for a deliberately broad boundary. The definition to use for this mapping exercise is, any system where a statistical or learned model influences a decision, a recommendation, a classification, a generation, or an action affecting your organization or its people.
Okay, that casts an incredibly wide net. It catches advanced generative language models, obviously, but it also catches very basic predictive algorithms. That breadth is highly intentional.
During the discovery phase, you do not wanna be a purist debating the semantic difference between machine learning, deep learning, and true AI. Right, you don't wanna get bogged down in the math. Exactly.
Over-include now, prune later. If you set the threshold too high, you miss the rudimentary systems that often cause the most immediate harm. That makes sense.
Now, it is important to note that this functional definition is distinct from legal definitions. Oh, like the EU AI Act. Exactly.
If you are complying with regulation EU 2024-16889, the EU AI Act, Article III of that legislation, has a specific, legally binding definition of an AI system. But for internal discovery and mapping, you use this broader operational net. Got it.
Okay, so the source material identifies seven specific hiding spots. I wanna walk through each of these in detail, really exploring the technical architecture and the specific governance friction they create. Sure, let's break them down.
The first category is built. Built systems are the models your organization trains, fine-tunes, or develops entirely in-house. So this is your own engineer's writing code.
Right, this involves your own data science and engineering teams. They are gathering the training data, defining the model architecture, and deploying the system on your own infrastructure. Can you give an example? Sure.
An example is a proprietary churn prediction model built by a telecommunications company to flag which customers are statistically likely to cancel their contracts next month. Okay, from a governance perspective, built systems seem like they should be the easiest to track because, well, the company controls the entire pipeline. You'd think so, right? No.
They are highly visible to the engineering department at first, but surprisingly, they often become completely invisible to executive leadership over time. Why is that? Technical debt accumulates. The original engineers who built the system leave the company, the documentation rots, and the model just continues running quietly on a server.
Oh, just humming along in the background. Exactly. It keeps making decisions based on data pipelines that haven't been audited in three years.
So even built systems need to be explicitly mapped. Okay, the second category is bought. This is standalone, licensed AI.
Right. Products your procurement team specifically went out and purchased because they are AI solutions. Like a dedicated tool.
Exactly. You are licensing a third-party applicant screening tool to parse resumes, or maybe a fraud scoring service for your transaction processing, or a dedicated AI demand forecasting platform. So with bought systems, you have the advantage of a procurement paper trail.
There was an RFP, a vendor assessment, a contract. Usually. But the third category is where that paper trail just evaporates.
And the text notes, this is the most under-accounted and insidious category, embedded. Oh, embedded is everywhere. Let's spend time here because this is where standard enterprise software silently transforms into an AI decision-maker.
Embedded AI is the primary reason executive maps are incomplete. These are AI capabilities switched on inside tools your organization already licenses for completely different traditional functions. You didn't buy an AI system, you bought something else.
Right, you bought a customer relationship management tool, an office productivity suite, an HR platform, or a help desk ticketing system. A system that might've been integrated into the company for a decade already? Exactly. Then the vendor pushes an overnight software update, and suddenly there's a new feature available.
An administrator clicks a checkbox to enable it, or sometimes, and this is worse, it is turned on by default. Wow, just instantly active. Instantly.
Your traditional database is executing learned models. What's a good real-world example of that? Well, Salesforce introduces Einstein, and suddenly it is reordering, which leads your sales team calls based on a predicted score. Or Microsoft rolls out Copilot, and suddenly a large language model has access to the semantic index of your company's emails, SharePoint drives, and Teams chats.
Oh, man. Or your help desk software, like Zendesk or ServiceNow, introduces an auto-summarization feature that reads inbound customer complaints and drafts replies. Let's trace the data flow here to understand why this is such a governance nightmare.
If I am, say, an HR manager, and I click enable smart resume ranking on my standard applicant tracking system, what is actually happening technically? Technically, you are initiating a whole new data pipeline. Previously, your HR software just stored PDFs in a structured database on a cloud server. When you enable the embedded AI, those PDFs are now being actively read, tokenized, and passed via API calls to a machine learning model.
So the data is moving. Yes, and is often hosted by a secondary subprocessor like OpenAI or Anthropic, depending on the vendor's backend architecture. The model computes a statistical ranking, returns the output, and alters the user interface to hide certain applicants and promote others.
And because it arrived as a simple feature update, it totally bypassed the IT procurement process, the security review, and the legal review of the data processing agreement. Precisely. The HR manager thinks of it as just a nice UI enhancement.
They don't realize they just deployed an algorithmic decision-making system that touches employment law. That is wild. Embedded AI sneaks through the perimeter because it rides on the trusted rails of existing vendors.
Which flows right into the fourth category, third party. This is hidden AI inside a vendor service delivery rate. This one requires careful vendor interrogation.
You hire a business process outsourcing firm, a logistics provider, or a marketing agency to deliver a specific service or outcome. You're not buying software, you're buying a service. Exactly.
You aren't buying software from them, you're buying the result. But unknown to you, their internal operations are powered by AI. For example, if I contract a logistics vendor for supply chain routing, I assume they have human dispatchers and traditional routing software.
Sure, that's what you'd think. But internally, they have deployed a predictive AI model to optimize routes, and they are secretly relying on offshore human reviewers to handle the edge cases where the AI fails. Like the Amazon example.
Right, and the sales deck you received never mentioned the model or the offshore reviewers. So you're exposed to the vulnerabilities of an AI system you didn't even know was in your supply chain. Sneaky.
Okay, so the fifth category is shadow, which we covered pretty thoroughly. The unapproved BYOAI tools and public LLMs employees use to circumvent official bottlenecks. That brings us to the sixth category, hybrid.
Yes, hybrid systems. These are systems marketed as fully autonomous artificial intelligence, but they actually rely on human intervention or review to function. So the Amazon just walk out scenario is the archetype here.
It's the perfect archetype. It's a combination of sensors, statistical models, and manual human labor. Why do these require their own separate category though? Because the marketing fundamentally misrepresents the mechanism.
They aren't necessarily a fraud. I mean, in many cases, human in the loop is the only economically viable way to handle the long tail of edge cases where the model lacks confidence. Right, the AI can't do it all itself.
But because the human element is hidden, the buyer completely misunderstands the system's scaling laws and privacy implications. Got it. And the final seventh hiding spot is agentic.
Agentic AI. This moves way beyond passive recommendation or generation. These are systems authorized to take multi-step actions autonomously without a human prompting each individual step.
It just goes on its own. Yes. You give the system a high level goal and it chains together API calls to execute it.
Give me a concrete scenario of an agentic system in a corporate environment, because that sounds very sci-fi. Imagine an autonomous procurement agent. The goal given by a human is simple.
Maintain a 30 day supply of office equipment. The agent monitors inventory databases. When it detects a shortage, it independently queries three approved vendors for pricing, selects the lowest bid, authenticates into your enterprise resource planning software, generates a purchase order, emails it to the vendor, and logs the expected delivery date.
Wow, it's literally acting as a digital employee. Yes, it is. And the risk profile is exponentially higher.
Because it's actually doing things in the real world. Exactly. If a passive generative AI hallucinates, a human reads a bad draft and just deletes it.
If an agentic AI hallucinates, it might issue 10,000 purchase orders in a recursive loop before anyone notices, because the human is out of the execution loop entirely. That's a massive financial risk. Absolutely.
So for the inventory map, you must meticulously record what an agentic system is authorized to do without asking permission first. Okay, so our net is cast across the seven places. Built, bought, embedded, third party, shadow, hybrid, and agentic.
And as we look at the hybrid category in particular, it perfectly illustrates the fourth major directive of the source material. Record what a system actually is, not what it is called. Distrust the label, record the mechanism.
That's the mantra. Distrust the label, record the mechanism. We discussed this with the Amazon anchor case, right? If you govern the label autonomous AI, you focus on server uptime.
If you govern the mechanism human in the loop video review, you focus on international privacy law and labor conditions. It changes everything. But I wanna introduce some skepticism here.
Put yourself in the shoes of a hard-nosed, results-oriented executive for a second. Okay, I'm with you. If I deploy a system to screen resumes and it accurately surfaces the best candidates, why should I care about the underlying mechanism? If the end result is accurate, if the output is correct, why does the how matter to governance? It's a fair question, but the answer is that the output only measures immediate utility while the mechanism dictates your organizational liability.
Ah, utility versus liability. Right, governance is the management of risk, not just the measurement of success. Let's look at your resume screening example.
If the system is a simple rules-based keyword filter, just looking for the word Python and MBA, your legal exposure is limited to ensuring those specific keywords aren't proxies for discriminatory traits. Okay, simple enough. Now assume the system is actually a deep learning model trained on historical hiring data.
Your legal exposure just exploded because the model might've learned latent biases from 10 years of human prejudice, discarding applicants based on subtle linguistic patterns that you can't easily audit. And if the system is actually a hybrid where a third-party company in another country is manually reviewing the resumes when the AI flags them as ambiguous? Then you have just introduced a massive data privacy breach because you are transmitting personally identifiable information of job applicants to an undisclosed third party without their consent. Wow, okay.
See, the accuracy of the output in all three scenarios might be identical. The candidates selected might be great, but the liability profile of each mechanism is totally different. The mechanism is where the danger lives.
The source text identifies four specific categories of operational truth that must be documented in your inventory. I'll list them, and let's analyze the implications of each. First, fully automated.
Right, a fully automated system has absolutely no human involved in its operating cycle. The model receives an input, computes the output, and executes the decision instantly. And the risk there? The risk here is scale and speed.
A fully automated algorithmic trading bot or a dynamic pricing engine can cause massive financial damage in milliseconds if the data drifts because there is no human friction to slow it down. Okay, second operational truth, human in the loop. Here, the system cannot execute the final action.
It computes a recommendation, but a human must review, adjust, or explicitly approve it before it takes effect, like a radiologist reviewing an AI-generated tumor highlight before making a diagnosis. That seems safer. It is, but the governance challenge here is automation bias, ensuring the human doesn't just blindly click approve on everything the machine suggests, essentially turning it into a fully automated system by negligence.
Oh, right, third truth, human on the loop. The system operates autonomously, but a human actively monitors the execution stream and holds an override switch. Think of a human supervisor watching the real-time telemetry of an autonomous drone flight ready to abort.
What's the risk there? The risk is attention decay. Humans are terrible at maintaining high vigilance over long periods when a machine is operating smoothly. We get bored.
So true, and the fourth operational truth, human-run but called AI. This is the Wizard of Oz mechanism. The model is a facade.
It might do basic routing, but the actual cognitive work is being performed manually by humans behind the interface. Now, when an executive discovers that a system they bought is mislabeled, the assumption is almost always malice, right? The vendor lied to us. Sure, that's the knee-jerk reaction.
But the source material outlines three entirely neutral reasons why these gaps between the label and the reality occur. Intent doesn't actually matter for the Mac. Exactly, this isn't just about vendors lying.
The first reason is simplification. How does that happen? Well, AI checkout is a brilliant, punchy marketing phrase. Computer vision plus sensor fusion requiring manual offshore video review of ambiguous edge cases is a terrible marketing phrase.
It really is. Right. Sales and internal communications naturally abbreviate complex concepts for a pitch.
Okay, what's the second reason? The second reason is drift. Let's explore software drift. How does a system's reality diverge from its label over time? It comes down to technical debt and the reality of edge cases.
An engineering team launches a fully automated customer service chatbot. It works perfectly in the sandbox. But in production, customers ask bizarre, highly contextual questions the model just wasn't trained on.
Real-world messy data. Exactly, the failure rate spikes. So to fix it quickly, the team quietly patches the system to route any low-confidence interactions to a human call center in the Philippines.
Oh, wow, so it changed. The system is now human in the loop. But the internal documentation, the corporate slide decks, and the executive understanding, they never update.
The system drifted and the label remained completely static. And the third reason is misrepresentation. Right, which is deliberately overselling the capabilities of a product to secure venture capital funding, win an enterprise contract, or inflate the valuation of a startup.
So whether the gap is caused by simplification, software drift, or deliberate misrepresentation, your job as a governor is identical. Ignore the pitch deck and record the operational truth. Exactly.
Which brings us to the actual construction of this map. The fifth mandate from the source text is about discipline. The fields that earn their place are governance fields, not engineering fields.
Yes. The guiding philosophy here is index, not encyclopedia. If you try to build an inventory that captures every technical nuance of every single system, the project will collapse under its own weight.
People will just give up. Nobody will finish it. An unfinished, unmaintained inventory governs absolutely nothing.
Think of a pilot's pre-flight checklist. The checklist works because it's short enough to complete. It covers flaps, fuel, critical instruments.
It's actionable. Right. If you buried those checks inside a 100-page engineering manual detailing the thermodynamic properties of the jet engine turbines, the pilot wouldn't read it, the checklist would be ignored, and the plane would crash.
That's a great analogy. Your AI inventory is a pre-flight checklist for organizational safety. It must be readable, and it must be finishable.
Which means we have to be completely ruthless about what to leave out. Engineering details do not belong on an executive governance map. Give me examples of technical details that IT might try to include, but that we must strip out.
Do not include the model architectures. Like what? You don't need to know if the system relies on a transformer, which is the architecture behind large language models like CHAT-GPT or a convolutional neural network or a gradient boosted tree. That goes out.
Exactly. Do not include hyperparameters, which are just the tuning variables engineers use to adjust how the model learns. Do not include the exact size of the training dataset or the specific mathematical weights.
Leave all of that out. Where does that information go then? If an auditor or a technical lead needs that depth, they can look at the system's dedicated model card or technical specification. It doesn't belong on the executive map.
Okay, if we strip all that out, what fields actually earn their place on the map? What does a governor need to see? Focus on the core fields. You need a plain language name and description. You need the origin, which of the seven hiding spots it came from.
You need a broad indicator of the data it ingests. You need to list the people affected by it. You need the exposure if it fails.
What is the worst case scenario? You need the date verified. And crucially, you need the purpose, defined specifically by the actual decision the system touches. Why phrase it as the decision it touches? Because the nature of the decision is the single best predictor of the system's potential for harm.
An AI system that touches a decision about what product to display next to a shopping cart affects the decision about user attention. Low risk. Right.
An AI system that touches a decision about credit limits affects people's financial stability. An AI system that touches a decision about resume screening affects people's livelihoods. And that leads directly into what might be the most important field on the map, criticality.
The text makes a phenomenal distinction here that counters basically every instinct a technologist has. It really does. Criticality is about consequence, not technological sophistication.
The guiding rule is rank the cut, not the steal. I cannot emphasize this enough. Let's compare two systems operating in the same company to illustrate this.
Let's hear it. System A is a bleeding edge, multimodal generative AI tool. It takes text prompts and hallucinates gorgeous, complex 3D design mock-ups for internal brainstorming meetings.
It uses massive compute power, billions of parameters, and the latest transformer architecture. Super advanced. System B is a rudimentary, almost archaic algorithmic keyword filter used by the HR software to automatically reject incoming job applicants who don't match specific criteria.
It uses a very basic statistical model. Okay, if you ask a software engineer to rank the criticality of those systems. They will flag the GNI tool as high criticality, without a doubt, because the technology is so advanced, complex, and unpredictable.
And they will flag the keyword filter as low criticality because the code is simple and, frankly, boring. But as an executive governing risk, you must invert that ranking entirely, right? Entirely. The GNI tool is low criticality.
Why? Because the people affected are internal employees and the exposure, if the system fails, say if it generates a weird image with six fingers, is just an imperfect draft in a private meeting. No consequential decision relies on it. But the simple HR keyword filter.
It is high criticality because it touches the livelihoods of external job applicants. If that basic statistical model is biased, or if the criteria uses proxy for protected classes like age or gender, your organization has a massive immediate legal exposure. Wow.
Regulators do not care how boring the code is. They care that it discriminates. Rank the cut, not the steel.
It's like a butter knife and a surgical scalpel are both just pieces of shaped metal. The technology is essentially the same, but only one is used to cut into human beings. And that intended consequence dictates the extreme level of sterilization, care, and precision required to handle it.
Exactly. The keyword filter deciding who gets an interview is a scalpel disguised as a butter knife. Consequence over sophistication.
That criticality flag is the fulcrum of your entire governance strategy. It tells you where to spend your finite time. But before we get to the triage part of things, there is one more essential field that requires its own directive.
And this is where a lot of frameworks fail in the real world. The sixth principle. Every system needs a named human owner.
The accountability gap. Governance theories fall apart in practice here. A system without an owner is a system without governance.
It is a rogue asset. The source material highlights a counterintuitive reality during the discovery phase. When an executive or a compliance officer is mapping the estate and they uncover a system, they will often find the owner field is completely blank.
It happens all the time. The instinct is to view this as a failure of the mapping process. But the text argues that finding a system and marking the owner as unowned flag for assignment is not a failure.
It is a massive victory. It is a high value operational finding. Ungoverned systems are unowned systems.
The greatest danger is creating a fiction of accountability. Right. The compliance officer just writes their own name or the IT director's name in the box simply to make the spreadsheet look complete.
They hide the vulnerability. Leaving it marked unowned sounds an alarm. It forces an uncomfortable conversation.
It forces the organizational hierarchy to have a difficult conversation and make a deliberate delegation of authority. Finding an unowned AI system is like finding an unattended piece of luggage at an airport. Oh, that's good.
You don't just ignore it or put your own luggage tag on it. You trigger a response protocol to figure out exactly who it belongs to. So when that conversation happens, who should the owner be? Let's go back to our HR resume screener.
When that is discovered, does the lead data scientist who actually understands the algorithm own it? No. Or does the IT systems administrator who configured the cloud server own it? Definitely not. The source text firmly corrects the instinct to assign ownership based on technical expertise.
Technical expertise does not substitute for business accountability. So who owns it? The owner must be the senior business leader who is accountable for the decision the system touches. So the head of talent acquisition owns the resume screener.
Yes. Because governance is about consequence and the authority to intervene. The data scientist understands the math, sure, but they do not have the organizational authority to weigh the system's operational efficiency against its legal risk.
They can't make the call. And they certainly don't have the authority to shut down the hiring pipeline if the system's biased. The head of talent acquisition possesses that authority.
And frankly, they are the one who will be fired or deposed if the system violates employment law. The business leader owns the risks, so they own the system. Exactly.
The IT and data science teams merely support the owner. Okay, so we have defined the fields. We have ranked the criticality.
We've assigned the business owners. We have a pristine map. But a map of a dynamic environment decays rapidly.
Fast. An inventory dies the moment it is filed. Right.
A static snapshot of your AI estate is out of date the moment you save the file. Vendors ship new embedded features overnight. Employees adopt new shadow tools on their lunch breaks.
Internal engineering pilots transition into production workflows. So you can't just do this once. If you build this map once, present it to the board, and file it in a digital drawer, it is functionally useless within three months.
It's like using a street map from 1995 to navigate a rapidly expanding city. You're gonna drive your car straight into a newly built canal. Precisely.
Yeah. So to prevent the map from decaying, the inventory process itself requires three operational rules. Let's hear them.
First, the inventory as a whole needs its own owner. Usually a chief risk officer, CIO, or a dedicated AI governance lead who is strictly accountable for keeping the list true. Got it.
Second, every single entry requires a date last verified stamp. If a system hasn't been verified in eight months, the reader immediately knows the data is stale and unreliable. That makes a lot of sense.
And the third? Third, you must establish defined refresh triggers. What constitutes a refresh trigger in a corporate environment? A trigger is an operational event that automatically forces an update to the inventory. It integrates AI mapping into existing corporate workflows.
For example, a contract renewal with a SaaS vendor is a trigger. When procurement renews the CRM license, they must verify if any embedded AI features were added over the last year. A new software rollout is a trigger.
The transition of an in-house pilot project into a live production environment handling real data is a trigger. So you tie the map update to things that are already happening? Yes. You combine these event-based triggers with a standing periodic review, perhaps the quarterly audit, to catch the software drift that no single event announces.
Quick question. What happens when a system is decommissioned? If the company decides a tool is just too risky and turns it off, do we just delete the row from the spreadsheet to keep the map clean? Never, ever delete the row. Never.
You change its lifecycle status to retired or decommissioned, and you preserve it as a historical record. Why is that so important? Because if a regulator audits your company three years from now regarding a discriminatory hiring practice that occurred last year, or if a new CIO takes over and wants to understand historical data flows, they need an intact audit trail. They need to know that a specific system existed, what its mechanism was, what decisions it touched, and the exact date it was terminated.
Destroying the row destroys your defensibility. Okay, we have covered the philosophy, the definitions, the fields, and the maintenance. Let's pivot to the final portion of the source material, the discovery playbook.
This is about actually taking command. This is where the rubber meets the road. Right.
We need to get incredibly practical for the listener here. How does a busy executive actually execute this discovery across thousands of employees and hundreds of software vendors? Because we firmly established that simply asking the IT department for a list will only return the confident wrong number. It will.
Discovery is unglamorous, manual legwork. It requires acting like an investigator, not just an administrator. The playbook outlines four concrete, actionable steps to force the hidden systems into the light.
Step one, read the invoices. Follow the money. Finance knows what you got, even if IT doesn't.
Makes sense. You pull the last 12 to 18 months of software subscriptions, vendor payments, and cloud computing licenses. And you aren't just looking for obvious AI companies.
You are looking for feature tiers within existing SaaS platforms. Because vendors monetize and go at AI. Exactly.
You will find embedded AI simply because a department manager authorized paying an extra $15 a user per month for the AI assist premium tier of their project management software. No technical network scan will find that, but the invoice reveals it instantly. Brilliant.
Step two of the playbook, plain language interviews. The phrase in the source material recommends here is very specific, and I feel like it really speaks to the psychology of how people interact with technology. It does, because if you walk up to a business process manager and ask, do you use artificial intelligence in your workflow? They will almost certainly say no.
Because they're picturing robots. Right. To them, AI means a humanoid robot or a massive supercomputer.
They don't think of their updated spreadsheet software as AI. So what is the question you ask them? You ask about the decision, not the technology. When you decide X, what helps you decide it? When you decide X, let's play this out.
When you decide what inventory to reorder for the regional warehouses, what helps you decide it? The manager will answer, well, the vendors demand forecasting software spits out a recommended number every Monday morning. But honestly, Laura on my team has been here for 10 years. She knows the seasonal trends better than the software.
So she reviews the software's number, adjusts it based on what she's seeing in the market, and we execute Laura's final number. Wow, that is incredible. With one plain language question, you just uncovered that a system marketed internally as an autonomous AI-driven demand forecast is operationally a human-in-the-loop system.
Exactly. The software drift occurred via human override. And now you know you need to govern Laura's interaction with the model, the data she uses to override it, and what happens to the supply chain if Laura goes on vacation.
You only find that by interviewing the students about their decisions. It's all about the decisions. Okay, step three in the playbook, vendor interrogation.
You cannot rely on the marketing website or the glossy sales deck. You must ask vendors directly, explicitly, and in writing about their architecture. What do you ask them? You ask, do you use third-party foundation models? Yeah.
Do you train your models on our proprietary data inputs? Do you rely on offshore human contractors to handle exception processing or model validation? The logistics vendor example applies perfectly here. Yes. You ask the delivery routing vendor, and they formally admit that when their routing algorithm encounters an anomalous address, the image is sent to a data center in another jurisdiction where a human annotator resolves it in real time.
And that data flow was entirely invisible until the direct adversarial question was asked. Exactly. And finally, step four, amnesty surveys.
This attacks the shadow AI problem we talked about. Run anonymous, explicitly blame-free surveys across the organization. You must frame it as operational research, not surveillance theater.
How do you phrase it? You say, we know our official tools are lagging. We want to understand what consumer tools you're using to stay productive and what kinds of documents you find most useful to process with them. Yeah.
You just need the truth about what data is leaving the perimeter. Okay, so let's synthesize this. You do the legwork.
You scour the invoices. You conduct the plain language interviews. You interrogate the vendors.
You run the amnesty survey. Suddenly, the confident wrong number vanishes. The illusion is gone.
Your mental list of two flagship AI systems balloons into a spreadsheet of 45 different AI systems scattered across built, bought, embedded, third-party, shadow, hybrid, and agentic categories. Overwhelming. For an executive, this is a moment of sheer panic.
The listener might be thinking, I cannot possibly govern 45 systems on Monday. I don't have the head count or the budget for that. That panic is real.
Yeah. It's the reason many executives subconsciously avoid discovery in the first place. Yeah.
But the source material provides the antidote, the triage strategy. You do not govern everything on day one. The rule is, complete the inventory broadly first, then govern deeply in criticality order.
Breadth first, depth second. Let's break that down. Breadth first means every single system you uncover gets a row on the map, no matter how sparse the information is.
Even if the row only says, name, unknown HR plugin, owner, unassigned, status, must verify. Just get it on the board. Exactly.
A known unknown is governable. An unknown unknown is a massive liability. Once everything has a row, you sort the entire list by the criticality flag.
Ranking the consequence. Right. You pull all the high criticality systems to the very top.
The systems that touch human livelihoods, financial assets, physical safety, legal rights, or systems that interact with customers unsupervised. A dangerous one. Those five or six high criticality systems receive your immediate deepest governance attention.
You assign the owners, you audit the data flows, you conduct the bias testing, the medium and low criticality systems. They just wait their turn in the queue. Triage is essentially the professional allocation of finite attention.
It is the ultimate defense. If an auditor or regulator or a board of directors asks you why a low risk internal categorization tool has a very thin governance record, you have a mathematically defensible answer. Which is? Because we deliberately prioritized our finite governance capacity on the systems that could actually cause harm to consumers or violate the law.
Triage is not a failure. It is the hallmark of mature risk management. Let's tie every single piece of this together with a comprehensive real world scenario.
The source material provides a really immersive case study about a fictional mid-sized regional retailer called Rivermark Outfitters. It's the perfect synthesis of the entire module. Let's walk through it.
We meet Austin. He is three weeks into a newly created role as the AI governance lead at Rivermark Outfitters. The chief operating officer calls him into an office and gives him a single blunt mandate.
Tell me what AI we run and tell me whether any of it can hurt us. That's the challenge. And if Austin was inexperienced, he would rely on the confident wrong number.
He would write down the two flagship systems everyone talks about. The custom built product recommendation engine on their e-commerce site and the massive enterprise demand forecasting platform they licensed last year. He would hand a list of two systems to the COO and fail his mandate.
But Austin follows the playbook. He starts with step one, invoices. He goes to finance, pulls the SAW subscriptions and finds three immediate surprises.
Which are? First, the customer support help desk recently upgraded to an AI assist premium tier that is automatically drafting replies to customer emails. That's embedded AI. Second, a regional HR manager expensed a resume linking add-on for their hiring portal, also embedded AI.
Third, he spots a line item for smart routing services from their third party logistics vendor, third party AI. He adds all three to the map. Nice.
Next he executes step two, plain language interviews. He sits down with the head of merchandising. It doesn't ask about AI.
He asks, when you decide regional inventory distribution, what helps you decide it? And he uncovers the exact scenario we discussed earlier. Exactly. The demand forecasting is actually a human in the loop system where an analyst routinely overrides the algorithm.
He records the actual mechanism, not the label. Then he moves to step three, vendor interrogation. He emails the logistics vendor providing the smart routing, asking specifically about human exception handling.
The vendor replies that when the routing algorithm fails, offshore contractors manually review the delivery manifests. So he records the hybrid mechanism and flags the international data flow. Finally, step four, the amnesty survey.
He sends out a blame-free questionnaire to all corporate staff. He discovers that marketing teams are pasting draft ad copy into public generative AI chatbots to brainstorm and procurement managers are uploading supplier contracts into unauthorized summarization tools to speed up their reading. Shadow AI.
The immediate risk is proprietary corporate data leaving the perimeter to train public models. So Austin has completed the broad discovery. He has a list of 11 systems, not two.
Now he applies the triage strategy. He looks at the flashy, highly advanced product recommendation engine on the website. The technology is complex, but the consequence is low.
It just shows a customer a different pair of boots. Criticality, medium. And he looks at the resume ranking feature the HR manager expensed.
The tech is basic, but it dictates who gets interviewed. Consequence is high. Criticality, high.
He ranks the cut, not the steal. Austin formats his map, keeping engineering details out so it remains an index, not an encyclopedia. He ensures every high-criticality system has a named business owner.
And he brings this table of 11 systems back to the COO. What does he highlight? He highlights three systems in red. The resume ranker because of employment law exposure.
The automated customer reply drafter because it speaks for the company unsupervised. And the shadow use of public chat bots leaking confidential data. The COO looks at the map and is stunned.
He says, I thought we only had two AI systems. And Austin replies, we have 11. And the two you knew about are not the ones that worry me most.
But now that we can see them, we can actually make decisions about them. That moment is the entire point of this exercise. It is the transition from organizational blindness to executive command.
By casting a wide net across the seven hiding spots, by distrusting marketing labels and recording the actual mechanism, by ruthlessly stripping out the engineering details, and by demanding human owners, you transform an unknown liability into a managed portfolio. Which brings us to a final provocative thought that looms over this entire methodology. The source text doesn't explicitly spell this out as a threat, but it is the grim reality of corporate governance.
Consider the immediate aftermath of an incident. Tomorrow morning, one of these embedded features or shadow tools you didn't know about fails catatrophically. The HR plugin makes a demonstrably discriminatory hiring decision that triggers a lawsuit.
Or a public chat bot regurgitates your proprietary source code to a competitor. The regulators knock on the door or the board of directors convenes an emergency session. What is the very first thing they're going to ask you for? They are not going to ask your engineers to explain the gradient-boosted tree architecture.
They are not going to ask for a lecture on neural networks. They are going to ask to see your map. They want the inventory.
They want documented proof that you knew the system was operating in your environment, that you understood its mechanism, that you accurately assessed its criticality, and that a named human being was accountable for it. And if you don't have it? If you cannot produce that map, you are entirely defenseless. In the eyes of a regulator, you cannot claim to be governing a system you didn't even know you had.
The absence of the map is the negligence. If you govern the label, you govern a fiction. If you don't map the estate, you are defenseless.
Which brings us to the Monday morning move. The single most valuable concrete action you, the listener, should take next week to step out of the dark. Do not wait for IT to build an automated scanning tool.
On Monday morning, bypass the technical perimeter. Go directly to your finance team or procurement department. Ask them to pull the last 12 months of SaaS software subscriptions, vendor contracts, and cloud licenses.
Follow the money. Go through that list line by line. Look at the software your company has trusted for years and highlight every single tool that might have quietly introduced an embedded AI feature, an upgraded AI tier, or a predictive module.
Don't debate whether it's true AI, cast the net wide, just highlight it. That highlighted list is row one of your inventory. It is the beginning of visibility.
In the world of organizational AI, nobody is gonna hand you an X-ray of your vulnerabilities. The vendors won't volunteer the risks and the employees will hide the shortcuts. You have to build the machine yourself.
You have to force the lights on. Because if you are standing in a store, marveling at the magic of an autonomous checkout, completely unaware of the thousand human beings in another hemisphere actually running the system, while you aren't governing the technology, you're just a tourist in your own company. Go build your map.
Real cases
These examples show the marketed-versus-actual discipline and the seven-category net applied to real system types an organization runs. The anchor case is examined first and in depth; the rest are illustrative categories, named where a real product family makes the point concrete.
Example 1: Amazon Just Walk Out (the anchor, examined in full). Amazon's cashierless checkout technology was launched with Amazon Go stores and marketed as computer vision that let shoppers take items and leave without checking out. According to reporting summarized by Bloomberg (2024), the operation involved roughly 1,000 people in India reviewing shopping video, with a meaningful share of transactions historically passing through human review; Amazon countered that these workers mainly trained and validated the model and checked only a minority of visits. In April 2024 Amazon announced it would remove the technology from its larger Amazon Fresh grocery stores in the United States in favor of Dash Cart smart carts, while retaining Just Walk Out in smaller formats and third-party venues. Inventory reading: the marketed label ("autonomous AI checkout") and the actual system (a human-in-the-loop hybrid with an offshore review workforce, a video-privacy surface, and human-dependent economics) diverge sharply. An inventory entry that recorded only the label would have governed none of the real risks. This is the case that justifies the "what it actually is" field for every entry you will ever write.
Example 2: The embedded feature you already bought (customer-relationship and office suites). An organization licenses a major CRM or office suite for reasons that have nothing to do with AI, then the vendor ships AI features (a Salesforce Einstein lead score, a Microsoft Copilot draft, an auto-summary in a help desk) that are enabled and quietly begin influencing decisions. Inventory reading: this is the "embedded" category, and it is the most under-counted because staff think of it as "just a feature of the tool." A lead score that reorders which customers a sales team calls first is an AI system touching a real decision, and it belongs on the list even though nobody ran an AI procurement to acquire it.
Example 3: The bought decision system (resume screening). A company licenses an applicant-screening tool that ranks or filters job applications. Marketed as "AI that surfaces your best candidates," the actual system might be a model, a keyword filter, or a hybrid, and it might run with or without a human reviewing its output. Inventory reading: origin is "bought," people affected is "job applicants," exposure if it fails is "discriminatory hiring decisions and legal liability," and criticality is high. The provenance and accountability sit with a vendor you will later need to interrogate, and the legal watch on hiring AI is active in several jurisdictions. (see Topic 5.3)
Example 4: The vendor's hidden AI (third-party inside a service). A logistics or customer-service provider sells you a service, and part of it is powered by AI the provider does not foreground: an automated triage model, a voice system, an offshore-plus-model hybrid. Inventory reading: this is the "third-party" category, and it is invisible until you ask the vendor directly. The Just Walk Out pattern (a hybrid presented as automation) recurs across the vendor landscape, which is why the marketed-versus-actual field applies to bought and third-party systems just as much as to your own. (see Topic 3.3)
Example 5: Shadow AI (bring-your-own tools). Individual employees use public generative-AI tools to draft emails, summarize documents, write code, or analyze data, on tools the organization never approved. By the 2024 Work Trend Index, this is the majority of AI use in many organizations. Inventory reading: origin is "shadow," owner is initially "unowned" (a red flag in itself), and the data field is the urgent one, because the governance risk is confidential data leaving the organization through an unsanctioned channel. Surfacing these converts invisible use into a decision: sanction, replace, or stop. (see Topic 2.1)
Example 6: The pilot that became production. A team runs a small proof-of-concept AI project with a narrow scope and clear caveats. It works, people like it, and without a formal go-live decision it quietly starts handling real cases. Inventory reading: the danger is that the governance attached to a "pilot" (light, experimental, low-stakes) never got upgraded to match the reality of a production system making real decisions. The inventory's refresh trigger exists partly to catch this transition, because the day a pilot starts touching real decisions is the day its criticality flag changes.
Example 7: The built internal model (churn prediction). A data-science team trains a model that predicts which customers are likely to cancel and feeds a list to the retention team each week. Marketed internally as "our churn model," the actual system is a built model whose output a human team acts on selectively. Inventory reading: origin is "built," so provenance and accountability sit with the internal team and the data it was trained on, not a vendor. People affected are customers who may be offered or denied retention incentives; exposure if it fails is unfair treatment of some customer groups and wasted spend. Criticality is medium: it shapes real customer treatment but does not decide rights, credit, or safety. The "built" origin matters because when the Module 2 data audit follows this system's data trail, the trail leads to the organization's own datasets and choices, with no vendor to interrogate. (see Topic 2.1)
Example 8: The translation or transcription feature (quietly consequential). A support team enables automatic translation of customer messages and automatic transcription of calls, treating both as harmless conveniences. The actual systems are AI models whose errors can distort the meaning of a complaint, misattribute a statement, or mistranscribe a critical instruction. Inventory reading: it is tempting to rate these low because they feel like plumbing, but criticality depends on what rides on the output. If a mistranslated complaint drives a wrong resolution, or a mistranscription enters a record that later matters in a dispute, the exposure is real. This example exists to warn against the reflex of rating "utility" features low without checking what decision depends on their accuracy.
Where people go wrong
- "We already know what AI we run." This is the confident-and-wrong answer, and it is almost universal. Leadership knows the flagships and misses the embedded features, the shadow use, and the vendor-hidden AI, which together are usually the majority of the estate. The cure is not confidence but discovery: invoices, interviews, tool audits, vendor questions, and a shadow-use survey. If your first count feels complete and quick, you have found the tip, not the iceberg.
- "Record what the system is called." The single most damaging inventory habit. The label ("autonomous AI checkout," "AI-driven forecasting," "smart routing") routinely conceals a different operational truth. Record the mechanism, not the marketing. The Just Walk Out case exists in your inventory as a permanent reminder that the label and the reality can be entirely different systems.
- "An AI system means our own trained model." Too narrow by far. Bought products, embedded features, third-party AI inside vendors, and shadow tools are all AI systems you run and must govern. An inventory that lists only home-built models misses most of what can hurt you.
- "More fields make a better inventory." No. An inventory drowning in engineering detail (architecture, hyperparameters, training-set sizes) is one nobody finishes and nobody reads. The governance fields are the ones a governor acts on: the decision it touches, what it actually is, who owns it, who it affects, and what happens if it fails. Hand the engineering detail to the systems that own it.
- "Shadow AI is a discipline problem to stamp out." Treating shadow AI as misconduct drives it further underground and destroys the trust you need to see it. It is a signal that official tooling is missing and real work is happening on unofficial rails. Surface it with amnesty, understand the data flowing through it, then decide what to sanction, replace, or stop. Punishment produces silence, not visibility.
- "Build the inventory once and file it." A one-time inventory is out of date within weeks, because vendors ship features, staff adopt tools, and pilots become production continuously. Design it as a living artifact with an owner, a refresh trigger, and a date-last-verified stamp on every entry, or it will be wrong exactly when a later module or a real incident reaches back for it.
- "Criticality is about how advanced the AI is." Criticality is about consequence, not sophistication. A simple keyword filter that screens job applications is higher criticality than a dazzling generative model that rearranges an internal dashboard, because the filter touches people's livelihoods and the dashboard touches nobody. Rank by the decision a system touches and the harm if it is wrong, never by how impressive the technology is.
- "We do not have AI, so we do not need an inventory." By the 2024 Work Trend Index, three in four knowledge workers already use generative AI at work, most of it on tools the employer never provided. An organization that believes it has no AI almost certainly has the least governed AI of all: entirely shadow, entirely invisible. "We have no AI" is usually a statement about visibility, not reality.
- "A utility feature like translation or transcription is too minor to inventory." The criticality of a feature is set by what depends on its output, not by how humble it feels. A mistranslated complaint that drives the wrong resolution, or a mistranscription that enters a record used later in a dispute, can carry real consequences. Inventory the feature and rate it on what rides on its accuracy, not on the reflex that "it is just plumbing."
- "The inventory is engineering's job, not governance's." Handing discovery to the engineering team alone reproduces the blind spot, because engineering knows the built and integrated systems and rarely sees the shadow tools, the embedded features enabled by a business unit, or the human-in-the-loop reality of a system in daily use. Discovery is a whole-organization activity that reads invoices, interviews business owners, and surveys staff. Governance owns the inventory; engineering is one of several sources it draws on.
Questions people ask
- What is AI systems inventory?
- A working document listing every AI system an organization runs, with the governance fields needed to act on each one (name, owner, purpose and the decision it touches, what it actually is, origin, data, people affected, exposure if it fails, criticality, and date last verified). It is the first artifact a governor builds and the one every later governance activity reaches back for. The plain principle behind it: you cannot govern what you cannot see.
- What is discovery?
- The active legwork of finding the AI systems an organization runs, using invoices, interviews, tool audits, vendor questions, and shadow-use surveys together. Discovery precedes every other governance activity, because risk classification, evaluation, logging, and incident response all operate only on systems that are known. More on Discovery
- What is marketed versus actual?
- The discipline of recording what a system actually is (its real autonomy level and human involvement) rather than what it is called. Central to inventory work because a single label can conceal a completely different governance profile, as in the Amazon Just Walk Out case.
- What is Embedded AI?
- AI features switched on inside tools an organization already runs for other reasons, such as a lead score in a customer-relationship platform or a draft-writer in an office suite. The most under-counted category in most inventories, because staff think of these as "just a feature" rather than as AI systems touching decisions. More on Embedded AI
- What is Shadow AI?
- AI tools that individual staff adopt without organizational approval, from public chatbots to browser extensions. In many organizations it is the majority of AI use. It is a signal that official tooling is missing, and its main governance danger is confidential data leaving the organization through unsanctioned channels with no owner and no logs. More on Shadow AI
Keep going
This lesson builds AI inventory and use-case intake, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.