Reading the primary source: a model card, a system card, and what they do not say
The short answer
The card is a primary source written by an interested party
Read it the way the International AI Safety Report (2025) read the science, marking confidence, except knowing the confidence markers were chosen by the seller. That makes reading them correctly more important, not less.
What you will be able to do
- Read a real vendor model card or system card end to end and locate its load-bearing sections (intended use, out-of-scope use, evaluations, limitations, training-data description, and change or version history) without being led by its framing.
- Separate three registers inside any such document: a firm claim backed by stated evidence, a hedge that sounds like a claim but commits to nothing, and a silence where an operator-relevant fact should be and is not.
- Analyze the "intended use" and "out-of-scope use" language for how it shifts the burden of safe operation onto you, the deployer, and decide what that shift actually obligates you to test.
- Apply the evidence-dilemma stance from the International AI Safety Report (2025): treat "the evidence does not establish this" as information you can act on, not as a gap to paper over.
- Detect staleness: find the date and version, decide whether the card still describes the system you are calling today, and connect this to documented behavior drift (see Topic 10.3).
- Cross-check a card's claims against independent sources rather than trusting the document alone, and know which claims you must reproduce yourself (see Topic 12.2).
- Produce a one-page source-reading note on the vendor card behind your own highest-stakes AI system, listing verified claims, unverifiable claims, and the specific questions to send the vendor (see Topic 3.3).
- Distinguish what a model card can tell you from what only a system card or your own testing can, so you never let a document about the model stand in for evidence about the deployed system.
The lesson
In January 2025, 30 nations that had gathered at the AI Safety Summit published a document that set the standard for evaluating general-purpose artificial intelligence. Led by machine learning researcher Yoshua Bengio and a group of 96 experts, the International AI Safety Report set out to answer a highly specific question. What does the scientific evidence prove about AI risk, and what does it not yet say? The report addressed what it called the evidence dilemma.
Policymakers are routinely forced to make sweeping decisions about artificial intelligence before the empirical proof of its safety or risk fully arrives. To solve this, the authors made no policy recommendations. Instead, they explicitly marked their own confidence on every claim.
They separated risks backed by robust evidence from those resting on theory, and they stated out loud exactly where the evidence was simply missing. That level of independent rigor is rare in applied AI governance. The documents that operators actually use to wire models into their organizations are almost never written by independent experts, marking their uncertainty honestly.
They rely on vendor model cards and system cards. A vendor card is a primary source, written by an interested party. It is drafted under severe deadline by marketing-adjacent teams, with a massive product launch riding on the outcome.
The document is engineered to persuade you to deploy, as much as it is to inform you of the risks. Extracting hard evidence from these documents requires an adversarial reading discipline. You have to read the source so closely that you can separate what the vendor can prove from what they hope, making the uncertainty surrounding the model precise.
Trusting a polished corporate document over reproduced evidence does not eliminate your operational risk. It merely leaves the deploying organization legally and technically blind to the failures the model will inevitably produce. This is a typical vendor system card.
It looks authoritative, prominently featuring a 91.2% accuracy score on the SafeReply benchmark. To tear this down, we must categorize every line into one of three registers. The first register is the claim.
This is a firm assertion backed by checkable evidence. By naming a specific number, a specific metric, and a defined test condition, the vendor has provided a claim that can be verified and independently reproduced. The second register is the hedge.
A hedge sounds firm, but commits to nothing falsifiable. Phrases like, designed to be helpful, describe an intention, not a result. They are engineered to survive scrutiny without carrying any actual information.
The expert tactic for handling a hedge is to mentally rewrite it in the margins as the specific question it is dodging. Designed to be helpful becomes an operational query, tested on which use cases, by what measure, and with what failure rate. The third and most dangerous register is the silence.
This is the operator-relevant fact that is entirely missing from the page. This text states no training data cutoff, names no evaluated languages, and provides no disaggregated failure rates for high-stakes edge cases. If a card is silent on a specific language or a known failure mode, the vendor has provided an absence of evidence.
For an organization making a deployment decision, an absence of evidence is operationally a no. Silences do not announce themselves on the page. Assuming that an unmentioned risk is an absent risk is a fatal failure in governance.
The silence list is exactly where your deployment decision actually gets made. A close reading requires comparing the different sections of a model card against each other. The most revealing findings live in the contradictions between what a vendor claims in one paragraph and quietly withdraws in another.
The intended use and out-of-scope use sections are the most consequential parts of the document, and the ones beginners most often skim. Intended use tells you exactly what bounds the vendor actually validated. The out-of-scope language is a direct transfer of liability, it puts on record the use cases the vendor explicitly refuses to stand behind.
If your actual use falls outside the validated bounds, and something goes wrong, the vendor's documentation is already on record stating you were warned. You own the risk. This chart shows an aggregate performance score, which is typical for the evaluation section of a model card.
But a single aggregate number excels at hiding severe failures in specific demographic groups, languages, or high-stakes conditions. The silence here is the breakdown. The limitation section reveals the vendor's willingness to state bad news.
A generic disclaimer telling you to review outputs helps you not at all. A specific limitation stating the model fabricates policy details for recent events gives you a testable failure mode. That specificity is a trust signal.
In the version and date section, you look for proof of stability. Vendors update models behind static product names, and behavior drifts invisibly over time. An undated card, or one older than the last system update, describes an endpoint you are no longer calling.
Together, these load-bearing sections function as a legal firewall. When read properly, they reveal the exact boundaries where the vendor's guarantees stop, and where undocumented risk is dumped onto the deployer's ledger. The demand for disaggregated quantitative analysis traces back to the original 2019 Google Model Card proposal.
It highlighted a face detection model whose impressive aggregate accuracy completely hid massive performance gaps across different skin tones and genders. Modern frontier system cards have expanded aggressively, sometimes running to hundreds of pages of red team findings. But extreme document length does not guarantee coverage of your specific edge case.
A 200-page document can still be entirely silent on your core deployment risk. Under Article 53 of the EU AI Act, this reading discipline carries legal weight. Providers of general-purpose AI models are now legally required to supply downstream operators with comprehensive technical documentation and a detailed summary of training content.
The enforcement reality is AI-washing. Regulators find companies providing no primary documentation because their automated systems secretly rely on heavy, manual human intervention. A missing card is a maximal silence, and often precedes a regulatory enforcement action over misrepresented capabilities.
A common mistake is assuming that a cited benchmark score in a vendor document proves the model is fit for your specific internal use case. Benchmarks can be unrepresentative of your data, or contaminated by having their answers leak into the training set. An even deeper error is relying entirely on a model card to approve a deployment.
A model card describes the model isolated in a lab. The real risk lives in the application layer, the prompts, the output filters, and the human oversight steps, which requires a system card to evaluate. Regulatory mandates assume you are applying this exact adversarial reading discipline.
You are legally expected to use the vendor documentation to meet your own compliance obligations. Reading the card properly is a codified duty. The theoretical markup phase must translate into an operational output.
Take your list of verified claims, rewritten hedges, and glaring silences, and separate every item into two distinct columns to build your action plan. Column A is Verify Myself. This holds the load-bearing claims your deployment decision rests on.
You cannot trust the vendor's benchmark data to reflect your specific reality. Operators must take these claims and independently reproduce them by testing the model directly on their own edge cases. Next, focus shifts to column B, Ask the Vendor.
This isolates the red margin silences, facts only the vendor holds, like training cutoffs, evaluated languages, and build version confirmation. Every question in column B must be put directly to the vendor in writing. Do not guess facts the vendor must track.
If a vendor evades your specific questions or flat-out refuses to answer, remember that this non-answer is never a dead end. It is critical evidence of the model's limitations, and you log their refusal word-for-word in your due diligence file to justify a highly restricted deployment. This two-column method forces an outcome.
It turns a vague unease about a polished corporate document into a precise, targeted list of actionable verifications that drive safe deployment. The ultimate artifact of this entire process is a one-page source reading note. It logs the document version, the verified claims, the silences, the test results, and a final sentence dictating what the evidence supports deploying.
This reading note goes directly into the organization's evidence annex. It serves as a permanent, inspectable record of due diligence, proving to auditors, regulators, or your successors exactly why a system was approved. Crucially, this artifact makes continuous monitoring incredibly efficient.
The frontier moves quickly, and vendors will routinely push silent version updates to their models. When that model inevitably updates, you do not need to start your risk assessment from scratch. You pull this reading note and simply retest the specific load-bearing claims you already isolated.
Read the primary source. Separate claims from empty hedges. Hear the hidden silences.
Then, step away from the document, and build the evidence the vendor card could not give you.
The ideas, one by one
Every line is a claim, a hedge, or a silence
A claim names something checkable. A hedge sounds firm and commits to nothing. A silence is the operator-relevant fact that is not on the page. Sorting a card into these three registers is the whole skill; the silence list is usually the most valuable output.
Silence is not safety
A card that says nothing about your language, your use, or your failure mode has told you nothing, which for a decision is the same as a no. Absence of a claim is never evidence of safety.
Out-of-scope language shifts liability to you
The uses a vendor refuses to stand behind are on record as your risk if you deploy them anyway. Read that section to learn whether your real use is inside the bounds someone actually validated.
Check the version and date first
Vendors update models behind stable names with no changelog, and behavior drifts (see Topic 10.3). A card older than the last update may describe a model you no longer call. That uncertainty belongs at the top of your silence list.
A close reading does not resolve uncertainty, it makes it precise
The output is not a yes or no; it is a sorted list of claims to reproduce yourself (see Topic 12.2) and questions to put to the vendor in writing (see Topic 3.3). A model card can never substitute for testing the deployed system.
The reading is an artifact, and it persists
Your one-page source-reading note goes into the evidence annex (see Topic 10.6), makes your standing frontier watch cheap (see Topic 12.3), and travels to your successor so your understanding outlives your tenure (see Topic 12.5).
A specific limitations section is a trust signal, a generic one is a red flag
"The model fabricates policy details for niche queries" hands you something to test; "outputs should be reviewed" protects the vendor and helps you not at all. Grade a card by how much of its bad news is specific enough to act on.
A close reading does not replace testing, it aims it
The whole payoff is a short, precise list of the few load-bearing claims to reproduce on your own data and the exact facts to ask the vendor, so you spend effort on what carries the decision instead of on everything or nothing.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 90 of the podcast.
Read the full conversation
So, on January 29, 2025, the most important document in the history of artificial intelligence basically dropped. Oh, absolutely. And what's fascinating is that it wasn't, you know, it wasn't some breakthrough algorithm and it wasn't a new piece of hardware.
Right. It was a report. Yeah, exactly.
Specifically, it was the first full international AI safety report. I mean, this thing was led by Yoshua Bengio. It was commissioned by 30 nations right after the Bletchley Park Summit and written by like 96 of the smartest minds on Earth.
Just a massive undertaking. Huge. It ran to hundreds of pages.
But if you talk to anyone in the governance space today, they don't even talk about what was in those pages. No, they really don't. They talk about what the report absolutely refused to do.
I mean, it made zero policy recommendations. It didn't tell governments what to ban. It didn't offer a single rigid framework or a definitive thou shalt not.
Which, honestly, completely broke the mold for how we expect the international summits to work. Right. Usually they want a soundbite.
Yeah. Usually when you get that many world leaders and experts in a room regarding a high stakes globally disruptive technology, the natural instinct of governance is to, well, draw a very clear, very heavy line in the sand. But they didn't.
No. Those 96 experts looked at the state of AI and realized that drawing a hard line was actually the most dangerous thing they could do because the ground was just moving too fast. Yeah.
So instead of dictating policy, they did something infinitely harder. They rigorously marked their own confidence on every single claim. I love that.
They separated the risks that actually had robust empirical evidence behind them from the risks that rested purely on laboratory modeling or theoretical math. And crucially, they stated out loud exactly where the evidence simply wasn't in yet. Exactly.
They embraced the muddy waters instead of pretending the water was clear. Which is so rare. And the report gave us this brilliant phrase for this reality.
They called it the evidence dilemma. The evidence dilemma. Yes.
It's just the fundamental problem of having to act like having to make massive consequential decisions before the definitive proof actually arrives. Because I mean, AI capabilities are scaling so much faster than the science of evaluating those capabilities. Right.
They're outpacing the science by miles. If you wait for perfect certainty, you're paralyzed. Right.
But if you pretend the evidence is stronger than it actually is, you're reckless. So the disciplined response to the evidence dilemma isn't to freeze and it isn't to guess. It's to size your deployment decision to the exact strength of the evidence you have in your hand at that exact moment.
Precisely. And that evidence dilemma is exactly what we are tackling in this deep dive today. Because if you're listening to this, you are likely that sharp, busy professional who actually has to make these high stakes deployment decisions.
You're in the franchise. Yeah. You're running governance or you're directing IT or leading a product team.
You have to sign your name on the line that says, yes, we are turning this AI on for our customers. And here is the catch. The documents you rely on to make that decision aren't written by 96 independent experts with a mandate to be intellectually honest.
Oh, not even close. You are reading vendor documentation. You are reading the materials provided by the very company selling you the software.
And navigating that transition from the academic standard of a safety report to the corporate reality of a vendor's technical documentation, that is like the most critical skill an operator can develop right now. It's the difference between actually governing your AI and just being a passenger on the vendor's hype cycle. OK, let's unpack this.
We're treating this as an executive master class today. Zero fluff. We need to define the actual materials you're going to be staring at on your screen when you make these calls.
Absolutely. So in the trenches, this usually comes down to two specific types of documents. First you have the model card.
Simply put, a model card is a short, structured document describing a trained model in the laboratory environment. It tells you what it is, what it was trained to do, where it excels, where it completely falls apart, and exactly how the vendor knows that information. It's essentially the spec sheet for the raw engine.
But you don't drive an engine, you drive a car. And that's where the second document comes in, the system card. While the model card describes that underlying engine on a test block, a system card describes the fully deployed product that is built around one or more of those models.
So it's got all the extras bolted on. Exactly. The system card is supposed to detail the guardrails, the output filters, the user interface constraints, the human oversight mechanisms, and all those residual risks that remain even after the vendor has tried to mitigate the dangers.
The engine versus the car, I really like that analogy. So if I'm buying an API to power, say, a customer service chatbot, the model card tells me how smart the raw language model is, but the system card tells me if the vendor actually put a filter in place to stop it from swearing at my customers. Precisely.
The system card represents the reality of the software you are actually integrating into your tech stack. Which brings us to a really uncomfortable reality about these documents. Because vendors are now legally mandated to provide them in many jurisdictions, they aren't just dry technical manuals anymore.
No, they're highly strategic assets. Right. And this is the first foundational rule we have to establish.
It's the lens through which you must read every single word. Here's the core idea. The card is a primary source written by an interested party.
Let's pull that apart, starting with the legal reality you just mentioned. Under the European Union AI Act, specifically Regulation EU 2024-1689 as of August 2, 2025, providers of general purpose AI are legally required to supply this technical documentation to you, the downstream deployer. They have to hand it over.
They have to give you a detailed summary of the training content, the capabilities, and the limitations. And the law expects you to actually read it. It's not just a box to check.
The FTC guidelines in the U.S. regarding deceptive claims operate on a very similar wavelength, right? Absolutely. Your entire compliance posture, your legal defensibility assumes you are receiving this documentation and using it as the foundation of your governance. So you can't just say, oh, well, I read a really great teardown of this model on a tech blog, so I figured it was safe.
No. A tech blog is a secondary source. A leaderboard is a secondary source.
The vendor's card is the primary source. It's the original document from which all those downstream summaries and hot takes are derived. And you are entirely responsible for the claims made in that primary source.
You are. But, and this is the massive flashing red light here, you are reading a primary source authored by the party with the strongest possible financial interest in you saying yes to their product. I mean, they have a product launch to worry about.
They have quarterly earnings. They have a marketing adjacent team that likely had a heavy hand in drafting or at the very least heavily editing that card right up until the midnight deadline. Yeah.
Even the most conscientious engineering driven vendor and, you know, they're phenomenal ones out there. They still have to run these documents through legal and PR. Which means the vendor gets to choose what to foreground, what to put in a giant bold font, what to phrase very gently and what to omit entirely.
They choose their own confidence markers. You know, it's funny, when I look at a restaurant menu and the menu says we use world famous pristine premium ingredients, my brain instantly translates that to marketing fluff. Right.
I don't treat the menu as a health inspection grade. No one does. I know the chef wrote it to sell me a $20 burger.
That's a phenomenal analogy. But because these model cards are formatted like scientific papers, I mean they have abstracts, they have charts, they use Greek letters for mathematical formulas, we completely drop our guard. We read a menu and treat it like an independent safety audit.
The formatting is literal camouflage. Because it looks like an academic paper, you assume academic neutrality. But you legally cannot skip reading the menu because the law says the menu is your starting point.
So you're stuck with it. You are. You just have to learn how to read the menu like a forensic accountant rather than a hungry diner.
You have to strip away the polish and isolate the actual operational truth. So how do we actually do that? If I have a 20-page PDF in front of me and it's beautifully formatted and full of really reassuring language, how do I break it down? You do it by fundamentally changing how you process the English language. Yeah.
You can't read it like a novel. You have to break every single sentence you read into one of three distinct registers. The three registers.
When you scan a paragraph, you must categorize every line. Every single line is either a claim, a hedge, or a silence. Okay.
Let's take those one by one. I really want to build a mental framework for this. What qualifies as a claim? A claim is a firm assertion backed by stated evidence that could, in principle, be checked or reproduced by you.
It is falsifiable. Okay. Give me an example.
So if a vendor card states, on the MMLU benchmark, the model scored 88.7 percent, that is a claim. Why? Because it names a specific known test, the MMLU. It names the specific number 88.7, and it implies a specific condition.
You could literally go look at the MMLU, see what questions it asks, and prove that number of true or false. But just because it's a claim doesn't mean it's practically useful, right? Like, if they just say 91.2 percent accuracy, my immediate thought is, on what? Accuracy on a math test. Accuracy on identifying images of dogs.
Exactly. And that is a very specific trap we call a masked claim. A masks claim.
Yeah. It is a cousin of the claim that you have to learn to spot instantly. A masks claim gives you a highly specific number, which makes your brain think it's rigorous, but it strips out the exact condition that makes it checkable.
So it's useless. Completely. 91.2 percent accuracy, with no benchmark named, no data split defined, and no demographic population specified is entirely unusable.
It's an illusion of precision. As an operator, you must treat a masks claim as a completely invalid data point until the vendor supplies the missing context. Okay.
So a true claim has the metric, the number, and the condition. That's our first register. What's the second? The second register is the hedge.
And this is where legal and marketing teams earn their salaries. A hedge is a sentence engineered to sound exactly like a firm claim, but it commits to absolutely nothing falsifiable. Meaning, if I test it? If you try to test it, you'll find there's nothing to actually test.
Give me an example. What does a hedge look like in the wild? You will see phrases like, the model is designed to be safe, or it demonstrates strong performance across a wide range of tasks, or we've invested heavily in reducing harmful output. Designed to be safe.
Man, I see that everywhere. It sounds incredibly reassuring. It sounds reassuring, but look at the mechanics of the sentence.
Designed to be is an expression of the vendor's intention, not a measurement of the model's actual outcome. Ah, right. A car can be designed to be safe and still have the brakes fail on the highway.
The statement wasn't a lie, they intended it to be safe, but it carries zero operational information for you, the driver. Same with strong performance. Exactly.
Strong compared to what baseline? A wide range of tasks. Which specific tasks? A hedge is built to survive public scrutiny and legal liability without giving you anything you can hold them to. I want to put this into practice.
Let's do a quick micro-reading exercise. Imagine I am sitting at my desk, and I'm looking at a fictional vendor card for an enterprise product. We'll call it SafeReply.
It's an AI that drafts responses to customer emails. Okay, SafeReply. I'm going to read a short paragraph from their executive summary, and I want you to break it down.
Let's do it. Okay, sentence one. Our model achieves 91.2% accuracy on the industry standard SafeReply benchmark.
That goes in the claim bucket. Yeah. It names a specific metric and a specific test.
If you were physically marking up this document, you would highlight that sentence in yellow. Yellow for claim. Got it.
Now, just because it's in yellow doesn't mean you automatically trust it. You still have to ask if the SafeReply benchmark actually resembles the kind of emails your real customers write. But structurally, it is a claim.
It is checkable. Okay, sentence two. It is designed to be helpful and safe across a broad range of customer interactions.
Pure hedge. Highlight it in pink. Designed to be is an intention.
Broad range is undefined. Helpful and safe are subjective adjectives with no mathematical backing. It's just a vibe.
It's totally vibe. Okay. If a customer uses the AI to generate a legally disastrous email, the vendor can point to this exact sentence and say, hey, we only said it was designed to be safe.
We didn't guarantee it would be safe in your specific interaction. Pink. Meaningless.
Sentence three. We have invested heavily in reducing harmful outputs. And honestly, if I'm reading this fast and I'm stressed, I read invested heavily and my brain just checks a mental box that says safety handled.
Which is precisely the psychological effect it was engineered to produce. But think about it operationally. Investment is an input.
It is not an outcome. Right. You can waste a lot of money.
Exactly. A company can spend $2 billion on a safety team and still reduce harm by absolutely 0% if their methodology is flawed. It is a hedge wearing the costume of corporate diligence.
Pink highlight. So when I have a page covered in pink highlighter, what do I actually do? I can't just throw the document away. As an executive, what is my physical action to neutralize a hedge? You neutralize a hedge by rewriting it in the margins as the exact question it is dodging.
You convert their vague reassurance into your precise interrogation. Oh, I like that. So when they say designed to be safe across a broad range of interactions, you grab a pen and write.
Were your specific interactions were tested for safety? By what exact measure and what was the failure rate? You don't accept the hedge. You isolate the missing data it's trying to hide. Which brings us perfectly to the third register.
We've talked about what is physically on the page, the yellow claims and the pink hedges. But the most dangerous register isn't printed in the document at all. No, it's not.
The core idea here is that silence is not safety. We define a silence as an operator-relevant fact that simply isn't in the document. And silences are lethal because absence does not announce itself.
A hedge at least gives you a sentence to highlight in pink. A silence gives you nothing. Your brain naturally assumes that if a massive, sophisticated tech company didn't mention a risk, the risk must not exist.
Exactly. Right. If I'm reading a card for a translation model and it doesn't mention how it performs in Arabic, I don't consciously think, wow, they have zero evidence for Arabic.
I usually think, well, it's a global model. They probably handled it. And operationally, for you making a deployment decision, that assumption is a massive liability.
If a card says nothing about non-English performance, it has not told you it is safe in other languages. It has provided zero evidence. And in governance, no evidence equals a no for deployment.
Let's go back to our safe reply example. Those three sentences sounded so confident. But what were they actually silent on? Think about what wasn't there.
It never stated a training data cutoff date. It never named the languages it was evaluated in. It never provided a disaggregated failure rate.
Meaning, like, does it perform worse on high-stakes messages like refunds compared to low-stakes messages like password resets? Exactly. And it never gave a version or build date. That's five critical silences hidden behind a wall of polished text.
Imagine if those 96 experts from the AI Safety Report had written that safe reply excerpt. If they were forced to mark their own confidence and break their silences, what would it actually sound like? Oh, it would be a completely different document. A rigorous Safety Report-style version of that exact same system would read like this.
On the safe reply benchmark, Build 2026-03 scored 91.2% overall, broken down. 94% on billing questions, 88% on complaints, and 79% on refund disputes, which is our weakest category. Validated in English only, we did not evaluate other languages, known failure.
The model will occasionally state a refund policy detail that the deploying organization has not configured. Wow, okay. That is night and day.
It shifts entirely. The accuracy line is now actionable because it names a specific build date. The language scope is explicitly stated English only.
But most importantly, the vendor is volunteering their weakest category, that 79% on refunds, rather than burying it inside that 91.2% aggregate average. And they explicitly name a testable failure mode, policy fabrication. I would infinitely rather have those two lines of genuine bad news than the whole paragraph of polished hedges.
Because if I know it fabricates refund policies, I can build a filter for that. Exactly. I can train my human reviewers to specifically look for fake refund policies.
I can manage a known risk. I cannot manage a hedge. Specificity is trust.
The more specific they are about their failures, the more you can trust their claims about their successes. But let me challenge this idea of hunting silences, because the reality of these documents is getting kind of absurd. If I pull a system card from a top-tier frontier lab, let's say the Anthropic Claude Opus 4.5 system card from November 2025, that document is literally 200 pages long.
Yes, it is. It has massive appendices. It has external red teaming reports.
If a document is 200 pages of dense technical evaluation, surely a silence just means the risk genuinely doesn't exist. I mean, they wouldn't just forget to test a major limitation in a book-length document. That is one of the most dangerous assumptions an executive can make today.
You are confusing length with completeness. Really? Yes. Length creates reader fatigue.
A 200-page card can easily bury an operator-relevant limitation inside 199 pages of entirely unrelated, highly complex evaluation. When a document is that long, the risk isn't that it's thin. The risk is that you get tired by page 40, you see how much work they did on, say, bioweapons safety, and you just assume they must have also covered your specific customer service use case.
Oh, wow. Yeah, you just tap out and trust them. So how do you combat the fatigue? How do you hunt a silence in a document that massive without losing your mind? You hunt silences by holding a fixed, uncompromising checklist of what a complete card must answer for your specific deployment, regardless of whether the card is three pages or 300 pages.
You don't read to see what they have to say. You read to see if they answered your specific list of questions. Which means we need to know how to build that checklist.
We need to know exactly where to look in these massive documents. And that brings us to the architecture of the cards themselves. There are specific sections that carry immense legal and operational weight.
Yes, the load-bearing sections. We can call them the load-bearing sections. And the core idea here is that out-of-scope language shifts liability to you.
This is critical. Every rigorous card will have a section for intended use and a section for out-of-scope use. The intended use tells you exactly what the vendor actually validated the model for in their own labs.
And the out-of-scope? The out-of-scope use tells you the specific scenarios the vendor explicitly refuses to stand behind. Right. But be honest.
Most people treat the out-of-scope section like the Terms of Service on a software update. They just scroll right past it because it feels like standard legal boilerplate. And that is exactly how you end up in front of a compliance committee or a regulator.
Out-of-scope language isn't just boilerplate. It is a quiet, deliberate transfer of liability from the vendor to you. Wait, explain that.
If you decide to deploy a customer support model to triage incoming clinical messages at a hospital, and clinical triage is listed as out-of-scope in the vendor's card, deploying it puts the full legal and operational risk entirely on your organization. Wait, so it's like if I go hiking in a national park and the park ranger puts up a massive wooden sign that says, Trail not maintained beyond this point. If I read that sign, ignore it, keep walking and fall into a ravine, I can't sue the park.
The park is already on the public record saying I was explicitly warned. Is that what out-of-scope does to a company? That is the perfect analogy. Walking past the sign doesn't magically make the rugged trail safe.
It just makes you entirely liable for your own rescue. When a vendor puts a use case in the out-of-scope section, they are hammering that wooden sign into the ground. If your use case fails, the vendor will point directly to that section and say we told them not to do it.
Okay, so intended use and out-of-scope use are massively load-bearing. What other sections actually carry weight? Let's look at the evaluations and metrics section. The single most important thing you are hunting for here is disaggregated evaluation.
You mentioned that briefly earlier with the refund example. Let's dig into the mechanics of that. What exactly is a disaggregated evaluation and why is it so vital? Disaggregated evaluation means the performance data is broken down by specific groups, specific demographic cohorts, specific languages or specific conditions, rather than being rolled up into one massive generalized aggregate score.
Because an aggregate score can hide a catastrophic failure in a minority of cases. Exactly. If I tell you a system has a 95% accuracy rate, you feel great.
But if that system processes data for 100 people and it is 100% accurate for 95 of them, but has a 0% accuracy rate for a specific minority group of 5 people, that is a catastrophic systematic failure for that group. The aggregate score of 95% completely masks the localized disaster. This reminds me of the Epic sepsis prediction case.
That was exactly this dynamic, wasn't it? Yes. The Epic sepsis prediction model is the textbook case study for why aggregate numbers are dangerous. Epic deployed a clinical model designed to predict sepsis in patients.
Based on its aggregate reputation and initial metrics, it looked stellar. It was widely deployed across hundreds of hospitals. But when independent researchers actually validated it externally out in the real world across different demographic and clinical conditions, they found that its real-world performance was far worse than the aggregate score suggested.
It was missing critical cases and generating massive amounts of false alarms in specific clinical contexts. Because the initial high score was an aggregate that smoothed over all the messy localized failures. Precisely.
When a model card gives you a single aggregate number for a test, the silence is the breakdown. You have to ask, show me the performance across different age groups, different skin tones, different language dialects, different hardware setups. If they don't provide that disaggregated data, you have a massive silence on your hands.
Let's move to the limitations section. I feel like this is another place where vendors try to hide bad news in plain sight. Counterintuitively, a long, highly specific limitations section is actually a massive signal of trust.
As we said before, specificity is trust. If a vendor writes, the model frequently fabricates policy details for events occurring after January 2024, that is a highly actionable claim. It is bad news, but it is useful bad news.
Because they can build a business process around it. Right. But if the limitation section just says, as with all AI systems, this model may produce inaccurate information, review all outputs.
That is not a limitation, it is a generic disclaimer. It's just covering their bases. It is a hedge that protects the vendor legally, but helps you, the operator, not at all.
If you cannot write a specific test that would trigger this stated limitation, it's just a disclaimer. What about the training data section? Especially with the new EU AI Act rules requiring summaries of training content, how do we read that section? You are looking for strict, unambiguous cut-off dates, and you are looking for the composition of the data. Does the card just say, trained on a large corpus of public internet data, or do they tell you the exact languages represented in the percentages of each? And increasingly, you must look for the synthetic share of the data.
Meaning data generated by other AI models, rather than collected from human sources. Yes. How much of this model was trained on the output of another model? We are seeing the consequences of training data opacity everywhere.
Look at the massive controversies around hardware vendors and data collection. We saw a stark cautionary tale with the Roomba images case a few years ago. Right.
Where sensitive, intimate images taken inside people's homes by autonomous vacuums ended up on social media because of the data annotation pipeline. Exactly. The hardware collected the data, but the pipeline involved sending those images to human gig workers for annotation to train the computer vision models, and those workers leaked the images.
Unbelievable. That case exposed a massive gap between what consumers assumed was happening and the reality of the training data supply chain. If a vendor card from 2026 is completely silent on where its training data comes from, how it was annotated, or what its synthetic data share is, they are withholding a fact that fundamentally shapes the model's blind spots and your liability.
Okay, so that covers the model card side. But if we are looking at a system card, the deployed product, what is the load-bearing section there? In a system card, you must interrogate the system guardrail section, and you have to be hyper-vigilant for what we call oversight theater. Oversight theater.
I love that term. I assume that means human-in-the-loop steps that look great in a PowerPrint presentation but are practically useless? Exactly. It's a process designed to satisfy auditors on taper, but it fails entirely in actual daily practice.
Vendors love to say, we ensure safety by keeping a human-in-the-loop for sensitive decisions. So how do you spot oversight theater in a document? What's the test? You run the three test questions against their claims of human oversight. Question one, what does the human actually see on their screen? Do they see the raw, full context of the model's output and the user's input, or do they just see a filtered, condensed summary? Okay, what's two? Question two, how much time do they have to make the decision? Are they expected to review a highly complex clinical or financial decision in three seconds, or do they have 10 minutes? And question three, can they actually overrule the system's decision, or can they only flag it for some downstream committee to look at later? So, if the reviewer is looking at a two-sentence summary, has a quota that forces them to act in three seconds, and can only click a flag button instead of actually stopping the action.
That is oversight theater. Statistically, a three-second review window is identical to zero review. The human is just a liability sponge.
A liability sponge? Wow. They are there so the vendor can blame the human when things go wrong. If the system cart doesn't explicitly define the parameters of the human review of what they see, the time allotted, and their veto power, that is a massive silence.
This is blowing my mind a bit, because you can do everything right. You can perfectly read all these load-bearing sections, you spot the hedges, you document the silences, you highlight the claims, you expose the oversight theater. But that brilliant, meticulous reading is completely useless if the document you are holding describes a reality that no longer exists.
Oh, absolutely. Which brings us to the fifth core idea you have to master. Check the version and date first.
If you do not do this, you are building a castle on sand. The core concept you have to understand here is behavior drift. Behavior drift is a well-documented phenomenon where a model's behavior changes dramatically between back-end versions, often while maintaining the exact same stable product name, and frequently with zero public changelog from the vendor.
So the API endpoint I am calling today might behave completely differently than the exact same endpoint I called last Tuesday. Yes. Stanford researchers conducted extensive studies on behavior drift over the last few years.
They proved conclusively that a stable product name, like GPT-4 or Quad-3, does not mean a stable model underneath. An API endpoint can shift its outputs, its formatting, its mathematical accuracy, and its failure modes overnight, because the vendor pushed a silent back-end update to optimize speed or safety. So a vendor card that doesn't have a specific build version and a publication date isn't a technical document, it is just a rumor.
It's a snapshot of a ghost. That is exactly what it is. The very first practical step you must take before you read a single word of the evaluation is checking the date on the card against the vendor's actual status page, or their release notes.
If the model card is dated March 1st, and the vendor pushed a silent update to the model on April 15th, that document might be describing an endpoint you are no longer even calling. Which means every single metric in that card is potentially invalidated. Yes.
That specific uncertainty version parity must go at the absolute top of your silence list before you do anything else. If you don't know what version you are reading about, the document is operationally useless. Let's take every abstract concept we've just discussed, the three registers, the silences, the shifting liability, the oversight theater, the versioning drift, and watch a sharp professional actually apply them to a high-stakes corporate decision.
I want to see what this looks like in the real world. Let's walk through an immersive scenario with Irene at Harborline. Let's set the stage.
Irene runs AI governance for a fictional mid-sized logistics firm called Harborline. She is about to route thousands of live customer service messages through a new LLM supplied by an outside vendor. This isn't a pilot.
This is production. So the stakes are real. The final deployment decision is hers to sign.
Her director hands her a 12-page vendor model card and explicitly tells her to read it with the same discipline as the International AI Safety Report. So it's a Friday afternoon, Irene sits down at her desk, she prints the card out, and she grabs three physical highlighters, yellow for checkable claims, pink for the hedges that sound firm but mean nothing, and she puts a blank piece of paper next to her keyboard to write down the silences. She starts with the intended use section.
She reads the first line. The card claims the model is validated for customer service assistance, including summarization and drafting in English. Irene highlights that in yellow.
It is specific. It is a claim. It is checkable.
But Irene knows her own business. Here is the operational reality. Harborline's logistics customers regularly write in English, Portuguese, and Cantonese.
So Irene immediately grabs her pen and turns to her blank page. She writes, Silence. No evaluation in Portuguese or Cantonese.
Card is silent on non-English performance. The vendor didn't say it was bad at Cantonese, they just didn't mention it, but the omission just told her exactly where her deployment is legally and technically uncovered. Exactly.
Next, she moves down to the evaluation section. She reads, The model demonstrates strong safety performance across our internal benchmarks. Irene takes her pink highlighter and strikes through the entire sentence.
It is a massive hedge. Strong compared to what? What constitutes an internal benchmark? She goes back to her blank page. She writes, Silence.
No disaggregated breakdown for high-stakes messages like customer complaints or refund requests. She knows that logistics is all about refunds and lost packages. If the card doesn't break out performance for refunds, she has zero evidence it works for her most critical use case.
Then she hits the limitations section. She finds a sentence that says, The model may produce confident false text about recent events. She highlights that in yellow.
That is actionable bad news. She knows she can build a guardrail for that. She can restrict the model from answering questions about delays that happened in the last 24 hours.
Nice. But the very next sentence says, Human review recommended for sensitive cases. Pink hedge.
Because sensitive is undefined. Recommending review without specifying the parameters is just oversight theater prep. Exactly.
Finally, Irene does the drift check. She flips to the front page to look for the version. The card has a date, but the version is just the product's marketing name, Harbor Assist Pro, not a specific back-end build number.
She pulls up the vendor's online status page and discovers that the model behind that marketing name was silently updated six weeks ago. The 12-page card she is holding predates the update. Wow.
So, at the very top of her blank page, she writes and underlines in red ink, This card may not describe the model we are calling today. Behavior drift is possible since the last update. And this is where the dynamic in the company fundamentally changes.
Think about what happens when Irene's director walks back into her office an hour later and asks if they are blocked from launching on Monday. If she hadn't done this reading, she would just say, Um, I feel nervous about this. It seems risky.
And she gets branded as the anti-tech compliance roadblock. She loses political capital. But because of her evidence dilemma stance, her political footing in the room is rock-solid.
She doesn't say they are blocked. She says they are precise. She is the only professional in the room armed with documented facts about what the primary source actually supports.
Her precision dictates her exact actions. He tells the director, We are officially halting the Portuguese and Cantonese deployment paths right now because there is zero evidence in the primary source to support them. Then, she takes the English refund messages and reproduces the safety claim herself on Harbourland's own historical data.
She runs a test. Yes. And in doing so, she discovers a hidden flaw where the model literally fabricates Harbourline's specific refund policies, a flaw the vendor's aggregate benchmark completely hid.
And finally, she doesn't just complain to the vendor. She sends four specific written questions directly to her vendor rep. 1. What is the training cutoff date? 2. What languages were formally evaluated? 3. What is the disaggregated failure rate specifically on refund logic? And 4. Does the current live build match this published card? Irene's story perfectly illustrates the final core concept we need to cover.
A close reading does not revolve uncertainty. It makes it precise. The goal of reading a 200-page system card isn't to give the vendor's prose a letter grade.
It is to generate an exact punch list of next steps. We call this workflow the verify or ask split. Every single gap, every pink hedge, every silence you wrote down on your blank page gets split into two distinct columns.
Let's break down those columns. Column A is Verify Myself. Column A is for the claims that you can and must reproduce on your own proprietary data.
If the vendor claims high accuracy on routing refund tickets, you do not trust their lab test. You run 500 of your own actual historical refund tickets through the model and you measure the real result. You verify it in your own environment.
And column B is Ask Vendor. These are the facts that only the vendor holds in their black box. Things you physically cannot test.
What was the exact training data cutoff date? Which specific languages were evaluated before launch? You send these questions to the vendor in writing. But as you were filling out these columns, you will inevitably run into a very common problem, conflicting information. What happens when you have multiple sources of information about a model and they disagree with each other? How do you rank them? That's a great point.
Because the marketing website might say one thing, the CEO's keynote says another, and the model card says a third. You must adhere to a strict hierarchy of truth. At the absolute top of the hierarchy is your own reproduced test on your own data.
That is ground truth. Nothing beats that. Directly below that is the vendor system or model card, because that is their most legally freighted document.
Below the card is the vendor's marketing material, their sales deck, their blog posts. And finally, at the absolute bottom are third-party leaderboards, which are often highly generalized, easily gamed, and sometimes contaminated by data leakage. So if the sales deck promises fully autonomous issue resolution, but the model card says designed to assist human agents, you believe the card.
We have seen massive enforcement actions around this exact discrepancy. Regulators are actively penalizing companies for AI-washing situations, where a company's marketing claims vastly outran their actual engineering reality. Regulators look at the gap between the glossy sales deck and the technical system card.
You must do the same. If the card doesn't support the sales pitch, the sales pitch is a liability. But let me ask you the most pragmatic question of this whole process.
What happens if you send your column B questions to the vendor in writing, and they just refuse to answer? What if the vendor rep emails back and says, we consider our training cutoff date to be proprietary competitive information, so we declined to share it? It feels like a roadblock, but a declined or evasive answer is still data. It is highly valuable evidence for your governance file. If they will not confirm the training cutoff, you log that explicit refusal in your due diligence document.
You literally write down, vendor explicitly refused to confirm cutoff date. So you just document the stonewall. Exactly.
That refusal dictates that you must size your reliance on the model accordingly. You must assume the worst case scenario for its knowledge cutoff, and build your guardrails around that assumption. Because if it ever goes to court, you have the paper trail proving you tried to find the edge of the evidence, and the vendor stonewalled you.
Exactly. All of this work, the highlighters, the blank page, the testing, the emails to the vendor, culminates in creating your final artifact. We call this the source reading note.
The source reading note. This is the output. This is a one-page document you create that summarizes your entire reading process.
It lists the verified claims, the remaining silences, the results of your column A tests, and the vendor's answers or refusals from column B. It goes straight into your governance evidence annex. It is your proof that you didn't just deploy and pray. And crucially, it goes to your successor.
Or to you, three months from now. Because when the vendor inevitably pushes a massive update to the model, you don't have to start from scratch. This single page tells you exactly which load-bearing claims you need to retest.
It tells you exactly what failed last time. It makes your standing watch over this technology cheap, efficient, and targeted, instead of a frantic scramble every time the API changes. We've covered massive ground today.
We've gone from the academic purity of the International AI Safety Report to the messy, liability-laden reality of corporate vendor cards. To summarize the core philosophy, the model card is the start of your evidence chain, not the end of it. Your close reading of the core tells you exactly what you need to test, and the results of your tests tell you whether you actually deploy.
But I want to leave you with a final reflection on the illusion of documentation. Just because a file exists in a repository and it has the title model card at the top tells you absolutely nothing about its utility. That is so true.
If you go look at public model repositories today, you will find thousands of auto-generated stub files. They have the bold headers for limitations and out-of-scope uses, but the paragraphs underneath are left completely blank. The template is there, but the bad news is missing.
A missing card, or a blank section in a stub template, is a maximal silence. It is the loudest silence possible. And regulators have repeatedly found, in case after case, that a missing or intentionally vague model card is almost always sitting right on top of a massive product misrepresentation.
If they won't write it down, they know it won't hold up. Which brings us to the single most valuable move you can make on Monday morning when you get back to your desk. I want you to pull the primary vendor card for your organization's highest stakes AI system.
The one that keeps you up at night. Look at the publication date on that document. Then immediately go to the vendor status page or release notes.
If that model has been updated since that card was published, even a minor point release, you must assume the document in your hand is stale. Put version parity at the very top of your silence list and email your vendor rep immediately to ask if the current live build actually matches the published card. Because if you don't demand the evidence, you are operating on an illusion of precision.
You're walking past the sign on the trailhead, hoping the bridge hasn't washed out. Have the discipline to state exactly where the evidence stops and act accordingly. Thanks for joining us on this deep dive.
Real cases
These are real documents and real events, used to show the reading discipline in action. Read each for the claim, the hedge, and the silence.
Example 1: The International AI Safety Report (2025), the standard to read toward. The first full report, led by Yoshua Bengio and backed by the 30 nations of the Bletchley process, is the clearest public example of a primary source that marks its own confidence honestly. It separated risks with robust empirical evidence (for example, harms from AI-generated media) from risks resting on modeling and theory (for example, some future-capability scenarios), and it stated plainly where substantial uncertainty remained. It made no policy recommendations, an act of discipline, refusing to claim more authority than a synthesis of evidence can carry. (UK Government, "First International AI Safety Report," 2025; a second edition followed in 2026, and the confidence-marking discipline it models carried over.) When you read a vendor card, you are reading toward this standard and measuring the distance the card falls short of it.
Example 2: The Google model card proposal (Mitchell and colleagues, 2019). The original "Model Cards for Model Reporting" paper introduced the structure operators now read: model details, intended use and out-of-scope use, factors, metrics, evaluation and training data, disaggregated quantitative analysis, and caveats (Mitchell et al., ACM Conference on Fairness, Accountability, and Transparency, 2019). Its motivating example was a face-detection model whose aggregate accuracy hid large performance gaps across gender and skin tone. The lesson for a reader: the disaggregated breakdown is the load-bearing section, and its absence in a real card is a silence, not a simplification.
Example 3: The OpenAI GPT-4 System Card (2023). When OpenAI published the GPT-4 System Card in March 2023, it documented not just capabilities but red-team findings and residual risks after mitigation across areas including bias, disinformation, over-reliance, privacy, cybersecurity, and proliferation. It is a genuinely substantive system card, and reading it teaches both halves of the skill: the claims are specific and useful, and the silences (for example, limited detail on training-data composition) are real and worth listing. A strong card still has silences; the reader's job does not end because the document is good.
Example 4: Anthropic frontier system cards (2025 onward). By late 2025 and into 2026, frontier developers were publishing lengthy system cards running to hundreds of pages, for example the Claude Opus 4.5 system card (Anthropic, November 2025), documenting safety evaluations, misalignment testing, and residual behaviors in unusual detail. The reading lesson flips: when a card is very long and very detailed, the risk is not a thin document but reader fatigue, that the one operator-relevant limitation is buried in a hundred pages of evaluation. Close reading scales with the document; a long card demands a checklist, not more patience.
Example 5: A claimed metric is not a reproduced metric. The gap between a metric a vendor states in a card and the same metric measured independently is not academic; it is the gap regulators now litigate and the gap external validation keeps exposing. This program treats two such cases in depth: a state enforcement action over a healthcare vendor's misrepresented accuracy claims (see Topic 4.2), and an external validation of a widely deployed clinical model whose real-world performance was far worse than its reputation (see Topic 4.6). For a reader here, the shared lesson is narrow and durable: read every metric in a card as a claim to be reproduced on your own use, not a fact to be believed.
Example 6: The EU AI Act GPAI documentation duty (in force 2 August 2025). The Act requires GPAI model providers to give downstream providers the information they need to understand a model's capabilities and limitations and to publish a summary of training content (Article 53, Regulation (EU) 2024/1689). This is the legal system formalizing exactly the reading you are learning: the law now presumes a downstream operator will receive and use vendor documentation to meet their own obligations. The card is no longer optional to read; a whole compliance posture now assumes you read it.
Example 7: Policy documents read the same way. The skill transfers past vendor cards to every document an operator meets. A soft-law declaration on frontier-AI safety, rich in shared intention and deliberately thin on binding commitment, is treated as its own topic in this program (see Topic 6.5); reading one teaches the same discrimination you apply to a model card, separating what was agreed and enforceable from what was merely affirmed. A regulator's guidance, a standards draft, a benchmark paper: each yields to the claim, hedge, and silence sort, so the habit you build on a model card pays off across your whole reading life.
Example 8: The GPT-4o System Card (2024), a card that grew with the capability. When OpenAI published the GPT-4o System Card in August 2024, the model was multimodal, including a voice capability, and the card added risk categories the earlier text-only cards did not need, covering areas such as voice-related risks and the potential for users to form emotional reliance on a human-sounding system. The reading lesson is that a card's structure tracks the capability: as a model gains a new modality, new sections appear, and the absence of a section that a new capability clearly demands is itself a silence. When you read the card for a capability the vendor just shipped, check that the card actually grew to cover it rather than reusing last version's template.
Example 9: The model-card ecosystem on public model repositories. Model cards have become a standard artifact on public model-hosting platforms, where thousands of models each ship with a card template. This is the good news (documentation is now expected) and the trap (a present card is not a complete card). Many community cards are auto-generated stubs with the honest-bad-news sections left blank, so the existence of a file named "model card" tells you nothing about whether it answers an operator's questions. The reading discipline is identical whether the card is a two-hundred-page frontier document or a three-line stub: apply the same checklist and let the silences fall out. A blank limitations section on a stub is exactly the same finding as a missing one on a polished card.
Example 10: When there is no truthful card at all, the AI-washing pattern. The sharpest lesson about reading a primary source comes from the cases where the primary source was a lie or did not exist. In a run of enforcement actions, regulators found companies whose public claims about an "AI" product did not match a system that was, in reality, largely run by people, and the paper trail of overstated capability was central to the case (the AI-washing pattern treated in Module 4 and the vendor-interrogation case in Module 3). (see Topic 4.1) (see Topic 3.3) For a reader, the takeaway is that the absence of a truthful, specific card is not a neutral gap; it is the same condition that precedes the enforcement actions, because a company that will not document what its system actually does is often a company whose marketing has outrun its engineering. A missing card is a maximal silence, and a silence that regulators have repeatedly found sitting on top of a misrepresentation. The retail sector produced its own version of the pattern in the widely reported case of an automated checkout system marketed as computer vision that turned out to depend heavily on remote human review (see Topic 0.2); no card or claim survives contact with the actual headcount behind the "automation," which is exactly why the reading discipline in this topic treats a vendor's own description as a claim to verify, never a fact to accept.
Where people go wrong
- "The card is the vendor's official statement, so I can rely on it." A card is a document written by the party that most wants you to say yes. Relying on it is not the same as reading it. The International AI Safety Report was written by independent experts mandated to mark uncertainty; a vendor card is written by the seller. Read the vendor card knowing who wrote it and why, and treat every metric as a claim to verify, not a fact to accept.
- "A confident, polished card is a trustworthy card." Polish measures the writing budget, not the honesty. A thin or generic limitations section ("as with all AI systems, review outputs") is a warning sign; a specific, uncomfortable limitations section ("fabricates policy details for niche queries") is a sign of a card you can actually use. Trustworthiness lives in specificity and in the willingness to state bad news, not in fluency.
- "No news is good news, if the card does not mention a risk, there is not one." Silence is the most dangerous register, not the safest. A card that says nothing about non-English performance has not told you it is safe in other languages; it has told you nothing, which for your decision is the same as a no. Absence of a claim is never evidence of safety. Keep a silence list precisely because silences do not announce themselves.
- "A hedge is basically a claim." "Designed to be safe," "strong performance," and "significant mitigations" are engineered to sound like claims while committing to nothing a test could falsify. Reading a hedge as a claim is how organizations deploy on documents that never actually said what they needed. When you meet a hedge, rewrite it in the margin as the question it is dodging.
- "A model card is enough to approve a deployment." A model card describes the model in the lab. Your risk lives in the deployed system: the prompt, the filters, the human steps, your data, your users. A model card's silence on system-level behavior is not covered by reading it harder; it is closed only by a system card (see Topic 10.3) or by your own testing. Never let a document about the model stand in for evidence about the product.
- "If the card cites a benchmark score, the capability is established." A benchmark can be unrepresentative of your use, contaminated by test data leaking into training, or funded by an interested party, all live problems at the frontier (see Topic 12.2). A cited score is a claim about a test, not proof of fitness for your purpose. Reproduce the claims that matter on your own data.
- "The version does not matter, it is the same product I evaluated last quarter." Vendors update models behind stable product names, and behavior can shift sharply with no changelog, a documented phenomenon (see Topic 10.3). A card without a build version, or a card older than the last update, may describe a model you no longer call. Check the date and version first, not last.
- "Out-of-scope use is legal boilerplate I can ignore." The out-of-scope section is a quiet transfer of liability: it puts on record the uses the vendor refuses to stand behind, which means if your use is out of scope and it fails, the vendor already warned you. Ignoring it does not make your use covered; it makes you the party holding the risk. Read it to find out whether your actual use is inside the bounds someone validated.
- "An AI tool summarized the card for me, so I have read it." A tool's summary of a card is a secondary source built on top of a secondary source, and it will smooth over exactly the hedges and silences you are trying to catch, because summarizers reward fluent generalities. Use a tool to speed the first sort if you like, but verify its output against the actual document, and never let the summary become the thing you file. The prompts in this topic are aids to your reading, not replacements for it.
- "The card cites a certification, ISO/IEC 42001 or a NIST AI RMF profile, so it has been independently checked." A management-system certification like ISO/IEC 42001 (established 2023) certifies that an organization runs a governance process, not that a specific model behaves as the card claims; it is a voluntary, certifiable standard, not a mandate, and not a test of your use case. A NIST AI RMF profile is a structured self-assessment against a framework, not an outside party grading the card's claims. Read a cited certification as evidence the vendor has a process, and keep reproducing the specific claims that matter to your deployment regardless.
- "The card is honest, so I can deploy without my own testing." Honesty is not sufficiency. Even a scrupulously honest card describes the vendor's tests, not your data, your users, or your edge cases, and it cannot know your deployment context. A close reading of an honest card still ends in a list of claims to reproduce; the honesty just means the list is shorter and the reproduction more likely to confirm. Reading grades the document; testing grades the fit.
- "If I ask the vendor a question and they do not answer, I have learned nothing." A declined or evasive answer to a specific, reasonable written question is itself evidence, and you record it as such. A vendor who will not put the training cutoff or the evaluated languages in writing has told you the fact is either unknown to them or uncomfortable to state, and either way that shapes how much weight your deployment can put on the model. Silence from a vendor is read the same way as silence in a card.
Questions people ask
- What is model card?
- A short, structured document that travels with a trained model and states what it is, what it is for, where it works, where it fails, and how the vendor knows. Introduced formally by Mitchell and colleagues (2019). For a reader, the primary source that describes the model in the lab, distinct from the deployed system. More on Model card
- What is system card?
- A document describing a deployed system built around one or more models, including guardrails, filters, human-oversight steps, and residual risks after mitigation. Where a model card describes the model, a system card describes the product a user actually touches. (see Topic 10.3) More on System card
- What is primary source?
- The original document a claim comes from, as opposed to a summary, ranking, or commentary built on top of it. In applied AI the primary source is usually the vendor's model card or system card, or a foundational report; frontier discipline means going to it directly. More on Primary source
- What is claim?
- A firm assertion backed by stated evidence that could in principle be checked or reproduced, for example a named benchmark score under a named condition. The parts of a card an operator can actually use, once traced to evidence. More on Claim
- What is masked claim?
- A sentence that names a number but strips out the condition (the test, the split, the population) that would make the number checkable, for example "91.2 percent accuracy" with no benchmark named. Not a hedge, because it commits to a figure, but unusable until the missing condition is found or requested; treat it as a hedge until then.
Keep going
This lesson builds Content provenance and authenticity, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.