The conformity file: assembling the evidence for the system you shipped in Module 3
The short answer
The file is assembled, not authored
You produced almost every page of the conformity file while building and testing the system in Modules 1 through 4. This topic is about putting those artifacts into the nine Annex IV slots in the order a reviewer reads, not generating new evidence. Treating the file as an end-of-project writing task is the mistake that leaves you reconstructing facts from memory.
What you will be able to do
- Assemble a conformity file for a high-risk AI system by mapping the artifacts you already built (the trained model and its failure explanation, the data provenance file, the evaluation report, the incident log) into the technical documentation structure the EU AI Act requires in Annex IV.
- Distinguish your role as provider or deployer of the system you shipped, and identify which evidence pack each role owns under Articles 16 and 26 of the EU AI Act (Regulation (EU) 2024/1689).
- Name the nine elements of Annex IV technical documentation and place each of your existing artifacts into the correct slot.
- Explain the difference between an internal self-assessment (Annex VI) and a notified-body assessment (Annex VII), and identify why a biometric system may trigger the notified-body route under Article 43.
- Draft the EU declaration of conformity (the short signed statement required by Article 47 and Annex V) that sits at the front of the file, and state how long it must be retained (ten years, Article 18).
- Attack your own file the way a regulator will: find the missing necessity justification, the undated version, the untraceable claim, and close each gap before it is found for you.
- State the current application dates so you build to the real deadline, not a stale one: prohibitions and AI literacy apply since February 2025, general-purpose AI model rules since August 2025, and stand-alone Annex III high-risk obligations are deferred to 2 December 2027 after the 2026 Digital Omnibus amendments.
The lesson
Between 2018 and 2021, the Australian hardware retailer Bunnings ran facial recognition software across 63 of its stores. Cameras scanned every visitor entering the building, turning human faces into numeric strings to check against a threat watchlist to protect staff from violence. The technology functioned exactly as intended.
It successfully caught watchlisted offenders. Non-matching data was deleted almost instantly. From an operational standpoint, the deployment solved the company's security problem.
Then the Australian privacy regulator intervened. The regulator did not challenge the software's accuracy or the company's intent to protect its workers. Instead, they asked the company to hand over the documented evidence proving the system was necessary, that visitors were properly notified, and that the risks were weighed before turning the cameras on.
Bunnings could not produce that evidence in the required form. In November 2024, Privacy Commissioner Carly Kind determined that the retailer breached the Privacy Act. The problem was the absence of assembled, pre-existing proof that the company had conducted the risk assessments the law demands.
If you cannot hand a regulator a single, complete evidence pack justifying your AI system's existence, the technical success of your model offers zero legal protection. Many organizations treat compliance documentation like a college essay written the night before a deadline. We call this the binder fallacy.
The assumption that you can build a system first, then draft a report from memory to satisfy the regulator. The standard required by modern AI law is a chain of custody. In this model, every claim about how your AI performs points directly to a version-controlled, time-stamped engineering artifact created at the moment the system was tested.
When a hostile examiner audits a backdabin report, they pull one thread, like a claimed accuracy metric, and ask to see the test data. If the cited test predates the live model version, the chain breaks. Regulators treat broken chains as evidence of concealment.
A defensible conformity file for a high-risk AI system is never authored from scratch at the end of a project. It is systematically assembled from the outputs your engineering team already generated. For organizations deploying high-risk AI, the European Union's AI Act is the definitive regulatory standard.
The law strictly divides liability between two roles. If you build an AI system, rebrand an existing one, or substantially modify how a model works, you are a provider under Articles 3 and 16. Providers own the heavy legal burden of compiling the Annex 4 technical documentation.
If you simply purchase a system from a vendor and use it exactly as instructed, you are a deployer under Article 26. Your compliance burden is lighter. You must maintain human oversight, retain logs, and run defined impact assessments.
Companies routinely buy a vendor model, retrain it on their own customer data, and assume they are still protected as a deployer. They are not. Modifying the system or changing its intended purpose legally converts you into a provider, and you immediately inherit the full Annex 4 documentation liability.
Knowing your precise legal role is the absolute prerequisite to compliance. Your role dictates exactly which evidence file you are required to build. If you are the provider of a high-risk system, the EU AI Act removes all guesswork about your documentation.
Annex 4 provides a strict nine-slot framework that serves as your non-negotiable table of contents. You already possess the materials to fill most of the structure. Slot 1 demands a strict boundary of the system's intended purpose.
Slot 2 houses your data provenance files, detailing exactly where your training data originated. Slot 3 requires an unvarnished account of your model's capabilities and limitations. Regulators immediately distrust any file claiming zero failure modes.
This slot is where you place your honest failure logs. Slot 4 tracks performance metrics. As this segmented bar chart illustrates, headline averages are often used to hide high false positive rates on specific demographic groups, the exact failure pattern seen in the Federal Trade Commission's action against Rite Aid.
This slot requires your disaggregated accuracy data to prove your model does not systematically fail vulnerable populations. The structure demands continuous traceability. Slot 6 logs time-stamped version histories.
Slot 7 lists the exact technical standards you applied, such as ISO 42005. Slot 9 mandates a living post-market monitoring feed to capture field errors. Seven of these nine regulatory slots do not require drafting new legal arguments.
They require the disciplined shelving of artifacts your engineering team has already built. The entire conformity file rests on slot 5. This is the Article 9 risk management requirement, and it operates as a continuous four-part feedback loop rather than a static form. First, you must document a legitimate aim for the system and prove that less intrusive alternatives were genuinely evaluated and rejected.
You must show why a simple rule-based system or a manual review failed to solve the problem. Next, you must explicitly document the harm the system imposes on targets. If your cameras scan innocent shoppers to find one thief, you must name those shoppers and calculate their exposure.
Finally, you execute a proportionality judgment, a written argument demonstrating that the intended benefit outweighs the residual harm to those non-targets. Because data drifts, this loop must cycle continuously, monitoring real-world conditions to verify the system remains justified. This documented contemporaneous argument for why the system deserved to exist in the first place is exactly what Bunnings lacked.
A file that only explains how a system functions, without justifying its existence, fails under regulatory scrutiny. Once the file is assembled, someone has to read it. Under Article 43, most high-risk systems follow the default path of internal self-assessment.
The provider issues a self-declaration, backed by the threat of massive financial penalties if an audit later proves the file is false. There is a critical exception. Biometric systems that lack fully applied harmonized standards strip away the self-assessment option.
These systems trigger Annex 7, forcing your file in front of a notified body, an independent, accredited external expert who will aggressively scrutinize your risk assessments and test data before your product is legally permitted to the market. Because a product's regulatory routing can change during development, you must assemble every conformity file to the standard of a hostile external examiner. The evidence must be perfectly legible to a competent stranger.
The final component of your compliance pack sits at the very front of the folder. This is slot 8, the EU Declaration of Conformity. The rule here is simple.
Sign last. Under Article 47, this declaration asserts the sole responsibility of the provider. Adding your signature before physically verifying the underlying chain of evidence converts a documentation task into a massive personal and corporate liability.
If you discover a missing risk assessment during assembly, you write it, date it with today's date, and flag it as a remediation plan. An honest dated gap is legally defensible. Fabricating a backdated document turns a simple compliance failure into actionable fraud.
You must anchor this work to accurate timelines. The 2026 Digital Omnibus deferred standalone high-risk obligations to December 2, 2027. However, the AI Act's general prohibitions and AI literacy requirements went live in February 2025.
The regulatory gates are already closing, meaning the internal structures to categorize and govern your models must be active right now. A true conformity file is characterized by its precise adherence to current law and absolute honesty regarding its own gaps. This discipline separates genuine AI governance from empty compliance theater.
Start by mapping your existing engineering artifacts and operational controls directly into an obligation-to-control matrix. Enforce strict version control. Ensure every test result, failure log, and data provenance file carries an immutable timestamp and author attribution before it enters the compliance folder.
Establish a recurring quarterly review cycle exclusively for slot 5. As your model encounters real-world data drift and scale, your risk management case must update to reflect current conditions. Governance that falls apart the moment a regulator asks for proof is just an essay. The assembled conformity file provides the traceable, concrete evidence that your AI system is safe, legally justified, and ready to deploy.
The ideas, one by one
Role decides the file
Provider or deployer is the first question, because it decides who owns the technical documentation, who signs the declaration, and what the file must contain (Articles 16 and 26). You shipped the Module 3 system, so you are its provider and you own the full file. Buying a system does not outsource responsibility; a deployer still owns oversight, monitoring, and impact-assessment records.
Annex IV is your table of contents
Nine slots: general description, development and data, capabilities and limitations, metric justification, risk management, lifecycle changes, standards applied, the declaration, and post-market monitoring. Seven you fill from existing artifacts, one is fresh (the declaration), one is a living commitment (post-market monitoring).
Self-assessment is the default; biometrics is the exception
Most Annex III high-risk systems self-assess under Annex VI, with no outside body reading the file before market. Biometric systems can trigger the notified-body route under Annex VII and Article 43 when harmonized standards are not applied. Assemble every file to the standard of a competent stranger, because you rarely know which route you will end on.
The necessity case is the heart of the file
The risk-management slot (Article 9) is where you justify building the system at all: the aim, the less intrusive alternatives, the harm to non-targets, and why the benefit outweighs it. This is the exact evidence the Bunnings determination found missing. A file that documents only how the system works, not why it should exist, cannot defend the decision to run it.
Sign last, and honestly
The declaration of conformity (Article 47) is issued under your sole responsibility and retained ten years (Article 18). Sign only after reading the file. An honest, dated late document is defensible; a fabricated contemporaneous one turns a compliance gap into a fraud finding. Mark your own gaps rather than hiding them.
Build to the real deadline
Prohibitions and AI literacy apply since February 2025; GPAI rules since August 2025; stand-alone Annex III high-risk obligations are deferred to 2 December 2027 and product-embedded high-risk to 2 August 2028 after the 2026 Digital Omnibus. The file is a living folder to have complete before the obligation bites, not a scramble at the deadline.
A chain of custody, not a binder
The expert does not read your file to admire it; they read it to break it, by pulling one claim and seeing if the trail goes cold. A claim that traces to a dated, versioned artifact survives; a claim that floats free, or points to a test that predates the live model, is a broken chain. This is why assembly from dated artifacts beats writing a polished report at the end: the report asserts, the assembled file proves. Every dated version stamp you add while shelving is one link an attacker cannot break later.
The file is a living folder, not a signing-day snapshot
Two of the nine slots are not retrospective: the declaration you sign last, and post-market monitoring (Article 72), which keeps the file true after signing. A file whose performance claims describe a model that has since drifted is already false, even if it was accurate the day it was signed. (see Topic 4.5) Seed post-market monitoring from your Module 3 incident process, name what you watch and how a field signal reopens the risk assessment, and review the file on a fixed cadence whether or not anyone asks. The file is complete only in the sense that a garden is complete: it stays true only if it is tended.
Mark your gaps; do not hide them
A gap you flag with an owner and a remediation date reads as a governance process working; a gap you hide reads, the moment it is found, as concealment that taints the whole file. A regulator will find material gaps during an inspection regardless, so the only thing you control is whether they are found already marked and being fixed. Marking converts a weakness into evidence of diligence.
The file is built to be attacked
Your conformity file is opened by a hostile examiner in Module 11, rebuilt after, linked by the evidence annex in Module 10, and inspected by the board in the capstone. A file that survives contact is real governance. Find and close your own gaps now, because every gap you leave is one the attack will use.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 39 of the podcast.
Read the full conversation
Imagine, for just a moment, that you're running a massive, nationwide hardware retail chain. Right. You're the CEO, or maybe head of operations, and you have this very real, very pressing problem.
Yeah. Something you can't just ignore. Exactly.
We're talking about coordinated theft, escalating physical violence directed right at your floor staff. Wow. Yeah.
It's not some theoretical risk. It's happening daily, and your mandate, your main job, is to protect your people. Of course.
So, you turn to the most advanced tech available, right? You decide to roll out a facial recognition system across 63 of your largest stores. That's a massive deployment. Huge.
The system captures the faces of everyone walking through the front sliding doors, instantly turns those faces into mathematical vectors like strings of numbers, and cross-references them against a highly curated threat list of known offenders. If there's no match. If there's no match, the system just deletes the data in milliseconds.
If there is a match, an alert goes to a human security guard who, you know, makes the final call on whether to approach the person. It's pretty robust on paper. Right.
In your mind, you've done everything right. You've minimized data retention, kept a human in the loop, and you're protecting your staff. Yeah.
You feel like you've solved it. But then, months later, the privacy regulator knocks on your door. Oh.
Here we go. Yeah. And they don't ask to see the code.
They don't ask for a live demo of the cameras. They ask one deceptively simple question. And that question is? Where is the evidence that you weighed the harms against the benefits before you switched the system on? Right.
And that question is the exact moment where a company's operational reality just violently collides with its legal reality. Yeah. Because the company, in this exact scenario, Bunnings Group Limited, they could not produce that evidence.
Not at all. Not in the format the regulator demanded. Yeah.
So, in November 2024, the Australian Privacy Commissioner issued this landmark determination ruling that Bunnings had breached the Privacy Act. Wow. And what's absolutely crucial for you to understand about this ruling is that it wasn't a debate about whether the technology successfully caught offenders.
I mean, it likely did. Right. The tech worked.
Exactly. The ruling was entirely about whether the company could prove they had done the required rigorous thinking in advance, and this is key documented, in one centralized, defensible place. Because they were deploying sensitive biometrics.
Right. Without adequate notice or consent. And when asked to justify it, they just couldn't produce the unified evidence file.
They failed the documentation test. Okay. Let's unpack this, because that is a terrifying scenario for any engineering or compliance team.
It really is. So, welcome to today's Executive Education Deep Dive. We are tackling Topic 5.6, the conformity file, assembling the evidence for the system you shipped in Module 3. Yes.
For those of you listening and following along in the curriculum, Module 3 represents all those grueling previous phases, right? Building the model, auditing the training data, shipping the feature, and evaluating its performance in the real world. All the hard engineering work. Exactly.
So, today, our mission is singular and critical. We are going to learn how to survive a regulator's audit by building a rock-solid evidence folder. We're going to figure out how to avoid the Bunnings Trap.
We are. And the cornerstone of this entire defensive strategy is a concept called the conformity file. Let's define that.
Yeah. Let's define that clearly right up front. A conformity file is the single, structured evidence pack that a regulator, an auditor, or even an underwriter can open, from page one to the last page.
Okay. And they can see definitive proof that your AI system was built, tested, documented, and legally defended before anyone ever asked a question about it. So, it's essentially the ultimate show-your-work document for AI.
It is the embodiment of your work. I mean, governance that hasn't survived contact with a physical, structured folder is, frankly, just an essay. Just talk in a meeting.
Right. It's a nice idea discussed in a Slack channel. Yeah.
The conformity file is where compliance transitions from abstract theory into tangible, legal reality. It makes it real. It's a load-bearing artifact.
If the file holds up under hostile scrutiny, your governance is real. But if the file is just a scramble of papers hastily assembled the night before the regulator's meeting... Which happens a lot, I'm sure. ...all the time, but a trained examiner will break that defense in a matter of minutes.
Which brings us to a massive psychological hurdle I see engineering teams face all the time. People seem to think that once the AI system is shipped and running in production... Yeah. ...some poor compliance officer has to sit down, open a blank Word document, and just author this massive, sprawling legal file entirely from scratch, trying to remember what the engineers did six months ago.
That instinct is incredibly common. Yeah. And it is exactly what turns a survivable compliance gap into an indefensible, even fraudulent position.
How so? Well, the golden rule here is that a conformity file is not authored at the end of the project. It must be assembled from real contemporaneous evidence that your teams generated naturally while they were building and testing the system across those earlier development modules, modules one through four. Okay.
Here's where it gets really interesting, and I want to ground this in a way that makes sense outside of a tech lab. I like to think of this using a culinero analogy. Oh, I like that.
It's exactly like a professional chef's mise en place. When you're cooking a complex meal in a high-end restaurant, you don't wait until the customer orders to start figuring out where the ingredients are. You crash and burn.
Exactly. You prep your station, you chop the garlic, measure the spices, reduce the stock, and you organize them in little containers at your station. In the AI engineering world, those prepped ingredients are your data lineage logs, your red teaming reports, your failure metrics.
You produce them as a natural byproduct of building the system. Right. So when the regulator arrives, or when the dinner rush hits, you are simply assembling the final dish from beautifully prepared ingredients.
You're not frantically trying to chop onions while the pan is actively catching fire. That mise en place comparison holds up beautifully because it highlights the danger of memory decay. If you try to chop those onions at the very end, like if you force an engineer to sit down and reconstruct complex facts, like the specific legal basis for scraping a training data set, or what the accuracy metrics were three model versions ago, purely from memory.
They're going to get it wrong. They will inevitably create gaps. They'll guess, or they'll misremember.
A conformity file is a living folder that grows alongside the system's architecture. A pile, on the other hand, is what you get when executives panic and scramble the night before the audit. Just like Leonard's fictional retail safety audit at Groundpost across Belgium and the Netherlands, taking a pile of scattered documents and organizing them into a defensible folder is an exercise in disciplined assembly, not fresh authorship.
Perfectly said. I have to play devil's advocate here, though. Go for it.
Let's look at the reality of a modern engineering team. You're telling me we have to stop engineering to document every single data input and parameter tweak for the lawyers. Not stop, no.
But in a fast-paced environment where teams are shipping code daily, pushing updates constantly to remain competitive? That level of contemporaneous documentation isn't just hard. For a lot of startups, it feels like a death sentence for innovation. How is that actually feasible without grinding development to a halt? You're touching on a very real tension, but the premise that it stops innovation is, uh, it's a misconception.
Really? Yeah. The reality is that modern engineering teams are already creating this documentation. They just don't view it as legal evidence.
Oh, I see. The proof isn't in some beautifully formatted legal brief. The evidence lives in their JIRA tickets, their GitHub commit logs, their automated testing outputs, and their post-incident review channels on Slack or Teams.
So the work isn't writing a novel. The work is translation. Yes.
Translation and mapping. The discipline of building the conformity file is simply taking those existing raw pieces of engineering truth and mapping them into the legal structure that the regulator expects. So you're shelving the evidence the team has already produced across their various shared drives and code repositories and organizing it into a defensible folder.
Exactly that. And that highlights why assembling is infinitely safer than authoring. Because you aren't making things up.
Right. When you author a document late, you inadvertently invent timelines that didn't exist. You smooth over rough edges that regulators actually want to see.
When you assemble, you rely strictly on the time-stamped artifacts of your actual engineering process, which gives the file an undeniable ring of truth. Okay. That makes operational sense.
But before we even start opening JIRA and organizing this evidence, we have to figure out whose job it is to hold this massive folder, right? Because depending on where you or your company sits in the AI supply chain, your legal obligations change completely. They absolutely do. The regulatory framework draws a very sharp, unforgiving line based on your specific role in the ecosystem.
You must determine if you are the provider or the deployer before you attempt to assemble a single page of evidence. Let's define those. Let's define them clearly, as they are the bedrock of the AI Act.
A provider is defined as the natural or legal person, public authority, agency, or other body that develops an AI system, or has one developed for them, and places it on the market or puts it into service under its own name or trademark. So to simplify, if my engineering team writes the code, trains the machine learning model from scratch, and we launch it to the public as our flagship product, we are undeniably the provider. That is the clearest example, yes.
And as the provider, you own the heavy, absolute documentation burden. Which means? You are legally responsible for producing the Annex 4 technical documentation. Annex 4. Let's define that for everyone.
Annex 4 is the specific structural table of contents required by the EU AI Act, which we will break down in granular detail shortly. Beyond that Annex 4 file, the provider also owns the obligation to establish a comprehensive quality management system. A QMS, right.
Yeah, a QMS. They must also sign the formal declaration of conformity, and they must register the system in the official EU database under Article 16. The provider holds the massive load-bearing folder.
Okay, let's look at the other side of the coin. What if my company doesn't build AI? What if we just buy an off-the-shelf system from a massive vendor like, say, a facial recognition tool for our office security, and we just plug it in and use it? In that scenario, you are the deployer. Deployer.
Yes. A deployer is defined as the party that uses an AI system under its own authority in the course of its professional activity. If you fall into the deployer category, you own a much lighter pack of evidence, which is governed primarily by Article 26 of the AI Act.
Lighter but definitely not empty. I'm assuming we still have to prove we aren't using the tool recklessly. Precisely.
As a deployer, you're not expected to reverse-engineer the vendor's massive technical file or explain their neural network weights. Thank goodness. Right.
However, you absolutely must keep proof of human oversight. You must document that you are following the vendor's instructions for use, maintain system monitoring logs, and crucially, in certain high-risk use cases, you may be required to execute a fundamental rights impact assessment. Often abbreviated as an FRIA.
Right. Let's define that for the audience because it sounds heavy. A fundamental rights impact assessment is a formal, documented evaluation required for specific high-risk deployments.
It forces the deployer to assess how using the system in their specific context might negatively impact people's basic rights. Like what? Things like non-discrimination, privacy, or freedom of expression. The deployer owns that localized assessment because the vendor couldn't possibly predict every nuance of how their tool will be used in the real world.
Makes sense. But this feels like a massive danger zone for modern businesses. Let's call them value chain traps.
Because I can easily imagine an IT director sitting there thinking, oh, I'm just a deployer. I bought this software from a tech giant. Their army of lawyers handled the heavy Annex 4v file so I'm legally shielded.
But you can accidentally cross that line and become the provider, right? Yes, you can. And if you cross it without knowing, you're walking into a minefield blindfolded. It is perhaps the most dangerous trap in the entire regulatory landscape, governed by Article 25.
The trap triggers if the deployer makes what the law calls a substantial modification to the AI system. Okay, we need to define this carefully. Right.
A substantial modification is a change made to the system after it has been placed on the market that is significant enough to affect its compliance with the safety requirements or a change that modifies its intended purpose. It essentially restarts the conformity assessment clock. I want to make sure I understand the mechanics of this trap.
Let me speculate on how this plays out in the real world. Say a regional hospital network buys a diagnostic AI tool meant to read x-rays. They buy it from a massive vendor.
They are the deployer. But after a few months, the hospital's internal data science team realizes the vendor's model flags way too many false positives for their specific local demographic. I see where this is going.
Yeah. So the hospital IT team decides to tweak the sensitivity threshold, and maybe they retrain the model just a bit using a few thousand of their own local patient records to make it more accurate. While they're at it, they slap the hospital's logo on the dashboard so it looks nice for the doctors.
That scenario is a perfect storm. Why? Because the moment that hospital executes those actions, specifically retraining the model on materially new data that shifts its performance or even just putting their own brand name on the user interface, the law completely shifts its view. The hospital is no longer just a deployer.
By modifying the system and rebranding it, the law now views that hospital as the provider of a brand new AI system. And I'm guessing the vendor's original legal paperwork doesn't cover them anymore? The vendor's original conformity file stops protecting the hospital entirely. In that instant, the hospital just inherited the massive Annex 4 documentation burden, the requirement to build an internal quality management system, and all the associated legal liability.
That's terrifying for a hospital IT director. If they didn't budget the months of time and the legal resources required to build a provider-level conformity file, they are operating a high-risk AI system illegally. Knowing exactly which side of the provider-deployer line you sit on is step one, because the line is remarkably easy to cross by accident.
That is a staggering realization. For the sake of our deep dive today, let's assume we know who we are. We didn't just buy the system.
We built it. We shipped it ourselves in Module 3. We are firmly the provider. Now we are staring down the barrel of this massive requirement called Annex 4. If I look at the text of the law, it sounds incredibly intimidating.
It's a wall of legal text. How does an engineering team even begin to tackle it? You tackle it by demystifying it. You must treat Annex 4 exactly as what it is fundamentally designed to be.
A table of contents. Just a table of contents. Yes.
The EU AI Act is giving you a structural nine-part checklist. You assemble your file by mapping the artifacts your team has already created into these nine specific slots in a very specific order. So let's walk through this the way an auditor would.
But let's ditch the numbered list and think about the narrative of an investigation. If I'm an auditor walking into your office and I demand the conformity file, I don't just start reading raw Python code. I need to understand the boundaries of the system first.
Where do they look to figure out what this thing actually is? The auditor starts at the very beginning of the file, which aligns with slot 1, the general description of the system. This is your front matter. It includes basic hygiene, the version number, how it integrates with hardware or other software.
But its most critical function is defining the intended purpose. Let's define intended purpose because that's a loaded phrase. It is a legal concept.
It is a single, tightly bounded sentence describing exactly what the AI system is designed to do and, critically, what it is explicitly not designed to do. I see. It adds the fence line of the property.
Everything inside the fence is what you are legally defending. Yes. And that boundary is what your entire file is judged against.
If your intended purpose sentence in slot 1 states, This system is intended exclusively to flag potential shoplifters for human security review, then an internal use case, where the system automatically bans customers from the store with no human in the loop, is operating completely outside the file's protection. And what if I just keep it vague, like, this system optimizes retail operations? If you write a vague, marketing-speak intended purpose like that, you fail this slot instantly because an auditor cannot test a marketing slogan. Okay, so slot 1 is the fence.
If I know what the system is supposed to do, my next question as an auditor is probably, okay, where does it break? I want to know the weaknesses. And that deduction leads an auditor straight to slot 3, which is the failure explanation. The law requires you to rigorously document the system's capabilities, its known limitations, and any foreseeable unintended outcomes.
This is where your honest, unvarnished explanation of how the model breaks down belongs. I find this fascinating because it feels so completely counterintuitive to traditional corporate compliance. How so? Usually, the general counsel wants a document that says, our product is flawless, it is perfectly safe, there are no risks.
But here, claiming the AI system has no limitations seems like an instant red flag. It is a massive red flag. If an auditor opens a conformity file and reads a claim that a machine learning model has zero limitations and no foreseeable failure modes, it tells the regulator one thing.
You simply haven't studied the system. Machine learning is inherently probabilistic. It makes mistakes by design.
Honesty and transparency about those mistakes are what pass this slot. So we've set the fence line and we've documented the failure points. Now the auditor needs to look at what's actually powering this thing.
The fuel. The data. Exactly.
That brings us to slot two, data provenance. This is the section where you must detail your training methods, the system architecture, your data sources, and the explicit legal basis for collecting and using that data. So if we did our homework in the earlier modules.
Yes. If your team properly audited your data in the earlier development modules, this is simply the folder where that provenance report lives. The reviewer is going to look at this slot and ask a very binary question.
Can I trace every single training input back to a lawful documented source? And if you can't. If you have an untraceable data set, say a massive scrape of the internet where you can't verify copyright or consent, your file syncs right here. And I'm guessing they don't just take our word for it that the data produced a good model.
They want to see the scorecards. They want to see the proof. Which leads directly to slot four, metric justification.
This slot isn't just a place to dump your accuracy scores. It asks why the metrics you chose actually measure what matters in the real world. And crucially, it demands that your performance accuracy is broken down or disaggregated for specific affected groups.
So I can't just write a bold headline that says the model is 95% accurate and call it a day. You cannot. Because a single headline average flatters the model and hides localized failures.
Give me an example of that. Imagine a fraud detection system that is 95% accurate across the general population. But if you disaggregate the data, you find it has a 40% false positive rate for a specific minority demographic.
The headline average is actively hiding a massive discriminatory harm. The reviewer checks slot four specifically to see if you have disaggregated your metrics by the different types of people who will actually interact with the system. That makes total sense.
Now, I know from looking ahead at our outline that the next piece, slot five, the risk management loop, is a massive topic on its own. It's the core of the necessity case. So let's put a pin in slot five for just a moment because I want to give it the time it deserves.
Let's look at the rest of the file. What happens when the system inevitably changes? Because as we established with the hospital example, AI isn't static code. That operational reality is captured in slot six.
Life cycle changes. This is your version control history. It requires dated, detailed records of every version update and significant change made to the system.
Just a log of everything we've tweaked. Right. When you retrain a model, when you shift a parameter, you record it here with exact timestamps.
An undated file cannot be trusted by an auditor because they have no idea which version of the system your testing metrics actually describe. Traceability is everything. Traceability.
Got it. But as an engineering team is building all this, the data provenance, the metrics, the life cycle logs, is there a blueprint they should be following? Do they just invent their own best practices? They shouldn't invent their own. And that brings us to slot seven.
Standards applied. Here, you are required to list any harmonized standard you use during development. Harmonized standard.
Let's define that one. It is a pillar of European product safety law. A harmonized standard is a specific technical specification adopted by a recognized European standardization organization.
Crucially, if you apply a harmonized standard, it creates a legal presumption of conformity with the requirements of the AI Act. So basically, if I build my AI system following the exact step-by-step blueprint of this EU harmonized standard, the regulator legally has to assume I'm compliant until they can definitively prove otherwise. It's like a golden ticket for compliance.
That is a very apt way to describe the mechanism. It flips the burden of proof somewhat. However, there is a significant real-world catch.
There usually is. Right. As of our current timeline in 2026, many of the specific EU harmonized standards for AI are still being drafted and settling into place.
So in practical terms, you won't always have a perfect European standard to point to. So what do you do? Leave slot 7 blank? Never. A blank slot is a beacon for an auditor.
It reads as a negligent omission, not a legal exemption. So what's the alternative? If no harmonized standard exists yet for your specific technology, you don't stay silent. You explicitly state the gap.
You write, no fully harmonized standard currently applies to this specific architecture. Therefore, we applied globally recognized solutions such as ISO-IX standards and internal engineering protocols X and Y to meet the safety requirements. You must mark and explain your methodology.
Okay. So we've mapped the boundaries, the failures, the data, the metrics, the version history and the standards. We've built the case.
Who actually signs off on this? That is slot 8, the EU Declaration of Conformity. Let's define that. This is a short, highly formal, signed legal statement required by Article 47 of the Act.
We will explore the immense weight of this signature later, but essentially it is the document where the provider claims sole legal responsibility that the system meets all regulatory requirements. And the final piece, I assume the regulator doesn't just want you to sign a paper and then walk away while the AI runs wild for a decade. They do not.
The final piece is slot 9, postmarket monitoring. This is the formalized system required under Article 72. Meaning we have to keep an eye on it.
Exactly. It is a forward-looking commitment to watch the AI's performance in the field after launch, gather real-world data, and feed any emerging problems or performance drift back into your risk management loop. If we step back and look at this entire table of contents, these nine slots, what's amazing to me is that if a development team has been doing their engineering work properly, adhering to basic MLOPS best practices, they already have seven of these nine slots filled.
Yes, they do. Model explanations, data lineage, evaluation reports, they literally just need to drag and drop them into slots 1 through 7. Only slot 8, the declaration, is a fresh legal document you have to author from scratch. And slot 9 is just a documented plan to keep watching the system.
That is the absolute essence of the assemble, not author philosophy. The heavy lifting is done during engineering, not during compliance review. Okay, so our table of contents is full.
The physical folder is assembled and sitting on the desk. Who actually grades this homework? Does every single high-risk AI system get audited by a government agency before it is allowed to launch? This is where we need to contrast the two primary assessment rights established by the AI Act. The default route for the vast majority of high-risk systems, those listed under Annex 3, is a process called internal self-assessment.
Break down the mechanics of self-assessment for us. What does that mean? Internal self-assessment, governed by Annex 6, means exactly what it sounds like. You as the provider check your own system against the legal requirements.
You assemble the conformity file, you sign the declaration of conformity, you affix the CE marking. And CE marking is? It's the physical or digital mark indicating European conformity. Once you affix that, you place the system on the market.
Crucially, no outside regulatory body is legally required to open your file and bless it before launch. So it's like doing your own corporate taxes. You calculate your own liability, you organize your own deductions, you sign the return, you send it in, and you just launch.
You just have to hope you don't get audited. But you better have an immaculate filing cabinet full of receipts if the IRS comes knocking. It is like doing your own taxes, but with one major difference.
The penalty for fraud or gross negligence here isn't just a financial fine. It can mean pulling your product off the market entirely and devastating your corporate reputation. It runs on self-declaration, backed by the severe existential threat of market enforcement if they find out you lied or cut corners.
However, there is a major critical exception to this self-assessment default route, and it involves biometrics. Explain the exception. Under Article 43.1, biometric systems, such as facial recognition, emotion recognition, or biometric categorization, trigger a completely different, much more rigorous route if harmonized standards do not fully cover them.
They trigger the notified body route. And what exactly is a notified body? Is it a government police force? No. A notified body is typically a private, independent, highly accredited organization designated by a national authority to assess conformity.
They act under Annex 7th of the Act. They are specialized technical auditors. To stick with the tax analogy, this is like being legally required to hire an independent forensic accounting firm to review your entire corporate tax return.
Audit your books and formally bless your math before the government even allows you to file it. That is precisely how it functions. An outside expert team is going to open your conformity file, read every page, challenge your metrics, and demand answers before the system ever sees the market.
And this raises an incredibly important piece of strategic advice for anyone listening who is building a conformity file today. Which is? Always build your file to the rigorous, legible standard required by a notified body, regardless of which route you think you are currently on. Why though? If I know for a fact I'm building a system that qualifies for internal self-assessment, why not just use our internal engineering shorthand? Why spend the extra hours polishing it for a stranger who might never read it? Because you rarely know, at the beginning of a multi-year build cycle, if a system might eventually trigger external review.
Scope creep is real. Maybe the system's intended purpose expands into biometrics. Maybe the regulatory law gets updated.
True. Things change. But even if you definitively stay on the self-assessment route, your file is never truly safe from outside eyes.
It can be subpoenaed by a court during civil litigation. It can be demanded by a massive enterprise customer doing vendor due diligence. It can be required by an insurance underwriter before they issue a policy.
Shorthand, internal jargon, and unwritten tribal knowledge shared only among your core developers will not survive external scrutiny. You must write the file so that a competent stranger, picking it up cold, can follow your logic from start to finish. Survive scrutiny.
That seems to be the operative phrase for this entire module. So speaking of surviving scrutiny, let's pull that pin out of slot 5. We skipped over the risk management loop earlier. If a regulator wants to take you down, if they want to dismantle your defense, where exactly are they looking first? They are looking straight at slot 5, the risk management loop, and specifically the necessity and proportionality analysis required under Article 9. This is the heart of the file.
It is the absolute load-bearing pillar of the entire conformity file. We talked earlier about slot 1 being the fence and slot 4 being the metrics. Those slots explain how the system works technically.
Slot 5 is entirely different. Slot 5 is the philosophical, ethical, and legal justification for why the system is justified to exist in the first place. Let's bring us all the way back to the hook we started this deep dive with, the Bunnings Group limited facial recognition case from November 2024.
Let's dissect exactly what went wrong in their boardroom. That case is the perfect anchor proof for why slot 5 matters above all else. Bunnings had 63 massive stores running facial recognition from roughly 2018 to 2021.
And let's be intellectually honest about the technology. It functioned. It did what they wanted it to do.
Yes. It successfully captured faces, vectorized them, matched them against a database of individuals known for theft or violence, and deleted non-matches. It likely did flag individuals who posed a genuine safety threat to the staff.
But the tech working flawlessly wasn't enough to save them from a massive regulatory failure. Not even close. Because when the Australian Privacy Commissioner investigated, the fatal flaw wasn't in the Python code or the camera hardware.
It was in the justification. Bunnings could not produce an assembled necessity case proving they had seriously, rigorously considered less intrusive options before deciding to scan the sensitive biometric data of every single ordinary customer who walked through their doors to buy a hammer or a can of paint. They lacked the pre-assembled proof that the harm inflicted on the public was actually weighed against the benefit to the store.
Yes. And this proves irrevocably that slot 5 is load-bearing. If you cannot justify the fundamental necessity of the system, the fact that your accuracy metrics in slot 4 are fantastic doesn't matter.
You could have a model that is 99.9% accurate, but if you shouldn't be running it in the first place, the whole conformity file collapses. So how does an engineering or compliance team avoid the Bunnings trap? What exactly goes into a valid defensible necessity case? Let's walk through the mechanics of it. A defensible necessity case is not a vague paragraph about synergy or security.
It has five distinct, rigorous elements that you must document honestly. Let's break them down. First, legitimate aim.
You must state the specific evidence-based problem you are trying to solve. Not a vague generalization like we want to increase store security, but a hard, data-backed statement. We have a documented 40% year-over-year increase in physical assaults on floor staff resulting in medical leave.
So you need concrete evidence of the problem before you propose the AI solution. What is the second element? Second, less intrusive alternatives considered. This is where you must document what else you looked at before resorting to AI.
Did you consider hiring more physical security guards? Did you consider requiring ID checks at the door? And crucially, you must explain precisely why those alternatives failed or were deemed insufficient to meet the legitimate aim. You can't just jump to a massive facial recognition rollout as the first and only option. Absolutely not.
Deploying high-risk AI should ideally be the last option, justified by the documented failure of less intrusive methods. Okay, third element. Third element, harm to non-targets.
You have to explicitly acknowledge the collateral impact of your system. In the Bunnings case, the core of the issue was the privacy impact on the hundreds of thousands of ordinary, innocent shoppers who were not threats, but whose faces were scanned and vectorized anyway, just to find the few bad actors. Acknowledging the harm.
It feels weird for a corporation to write down exactly how their product harms people. Usually, PR teams fight to hide that stuff. But here, it is a legal requirement to map it out.
What's the fourth element? Fourth, controls in place. What specific engineering and operational measures have you designed to minimize that harm to non-targets? For instance, implementing an ultra-short data retention policy where non-matching vectors are deleted in milliseconds, or ensuring a highly-trained human in the loop makes the final critical call on any biometric match before action is taken. And the final element that ties it all together.
Fifth, is the proportionality judgment. This is the synthesis of the entire argument. You most explicitly logically argue why the residual harm, the harm that remains even after all your technical controls are applied, is legally, ethically, and operationally justified by the overwhelming importance of achieving the legitimate aim.
I am listening to you outline this, and I can hear engineers pulling their hair out. Proportionality! Harm! These sound incredibly subjective. Proportionality is such a gray area.
If I'm a machine learning engineer used to optimizing for clear mathematical loss functions, how do I map this? Is there a cheat sheet for this? A framework engineering teams can use so they aren't just guessing what a regulator's subjective opinion might be. Yes, there is. And it brings us to a vital real-world integration that bridges the gap between legal theory and engineering practice.
Look at example 5 in our source material. ISO EC 42005, published in 2025. I'm assuming, based on the acronym, this is a massive international standard.
It is. ISO AES 42005 is an AI system impact assessment guidance standard. What makes it so brilliant for practitioners is that it perfectly structures this exact slot 5 necessity case.
It forces organizations to walk through the impacts on individuals and society in a repeatable, standardized, highly structured way. So it removes some of the guesswork. Exactly.
It takes the subjective concept of harm and turns it into a rigorous checklist. Furthermore, it maps neatly back to the broader ISO AES 42001 AI management system standard. Let me make sure I understand the legal weight of this.
If I use ISO 42005 to write my slot 5 necessity case, am I legally invincible? Is it a get-out-of-jail-free card? No. And we must be very careful not to overstate it. It is not an automatic legal shield.
It is a guidance standard, not a statutory legal mandate. But it is a globally recognized, highly respected engineering framework. If you use it to structure your impact assessment, and you cite it in slot 7 under standards applied, you transform a subjective, messy, highly vulnerable argument into a disciplined, standardized process.
Which looks great to an auditor. Serious enterprise buyers, sophisticated insurers, and most importantly, regulators respect that methodology immensely. It shows you didn't just guess.
You followed global best practices. Okay, so let's visualize where we are. We've justified the system's existence using the five elements.
We've used ISO 42005 to structure it so we don't sound crazy. Our physical folder is fully stocked. Slots 1 through 7 are looking beautiful.
The data is traced. The metrics are disaggregated. Now it's time for slot 8. Ah, yes.
It is time to sign on the dotted line. The Declaration of Conformity. But I know from our notes that the timing of the signature is everything, isn't it? The timing of the signature is the literal difference between routine corporate compliance and severe personal liability.
Let's look very closely at the mechanics of slot 8. The Article 47 Declaration of Conformity is a formal legal statement issued under the provider's sole responsibility. Sole responsibility. Yes.
Furthermore, Article 18 of the AI Act requires you to keep this signed declaration, along with the technical documentation, at the disposal of national competent authorities for 10 years after the AI system is placed on the market. 10 years. And the phrase sole responsibility carries a lot of weight.
If I'm the executive signing this, I'm putting my personal name or my company's legal liability on a document that effectively says, I promise to the European Union that this AI complies with every relevant law. Exactly right. Which is why the ironclad, non-negotiable rule of AI governance is, you must sign last.
Because if you sign it early, just to get the paperwork out of the way so the dev team can launch. If you sign a Declaration of Conformity when the conformity file is incomplete, or if you sign it without having personally fully read and verified the evidence within it, you are taking sole legal responsibility for sweeping claims you do not actually know are true. That is how a survivable corporate position becomes a completely indefensible personal liability nightmare.
The signature is not a bureaucratic hurdle to be checked off. It is the deliberate, informed acceptance of the evidence contained in the folder. But let's deal with the messy reality of corporate life.
Because it rarely works out perfectly. Say I'm Leonard, our fictional auditor at GroundPost. I'm assembling my massive file.
I'm getting ready to sign. And I realize with a sinking feeling, oh no, we never actually wrote down the necessity case for slot 5. Very common realization. Right.
And the system has already been running in test stores for eight months. The instinct, the overwhelming corporate reflex, is to quickly type up a necessity case, backdate it to eight months ago so the timeline looks perfectly tidy, slip it into the folder, and pretend it was there all along. And acting on that reflex is incredibly dangerous.
It is the fastest way to turn a slap on the wrist into a criminal referral. Here is the hard, unyielding rule that modern regulators operate by. An honest, late document is a defensible governance process.
A backdated, contemporaneous document is fraud. Wow. Let's really unpack that difference because it feels completely unnatural to most corporate cultures to admit a mistake on paper.
It does feel unnatural, but you have to understand how regulators view these files. If you find a missing document, a gap in your file, the correct action is to write the document today, date it today, and explicitly state in the file itself that it was formalized in response to a prelaunch compliance review. So you just lay it bare.
Yes. If you have old calendar invites or messy meeting notes proving that you actually did discuss alternative security measures eight months ago, you cite them. You write, we discussed less intrusive alternatives in key one.
Here are the meeting traces in the appendix, but we are formally documenting the proportionality analysis in this final format today. But that feels so exposed. You're explicitly admitting a flaw in your process to the regulator.
You are admitting a flaw, but regulators view a marked gap with a clear remediation date as evidence of a functioning, mature governance process. It shows self-awareness. Regulators know that perfect, seamless engineering development is exceptionally rare.
That's true. But if you invent a document and backdate it to make it look like it existed perfectly at launch, regulators and their forensic IT investigators are not easily fooled. They look at document metadata.
They look at email server logs. They can spot fabricated timelines with ease. And once they catch even one backdated document, they view the entire file and the entire organization as deceptive.
It poisons the well. If you lied about the date on the necessity case, they assume you lied about the accuracy metrics in slot four. Precisely.
It converts a simple, routine documentation gap into a formal fraud finding, which carries vastly harsher financial penalties and public relations damage. You must mark your gaps, bait them honestly rather than hiding them. OK, so knowing what to build and how to honestly sign it is half the battle.
The other half is knowing when this massive, intimidating file is actually due. Because if you are listening to this and looking at a compliance calendar from early last year, you are already lost. The ground shifted beneath us.
The timeline is absolutely critical. And as you noted, it shifted significantly very recently. We need to introduce a major legislative update, the digital omnibus on AI.
Let's define that. This was a comprehensive simplification package adopted by the Council of the EU on June 29, 2026. It completely altered the application dates for the AI Act to give the industry breathing room.
Give us the new real deadlines. When do these conformity files actually have to be finished? The omnibus deferred the heavy lifting. For standalone Annex III high-risk systems, things like biometrics or critical infrastructure AI, the obligations were pushed back to December 2027.
OK, December 2027. Right. For high-risk AI that is embedded into products already regulated under Annex I, such as medical devices, aviation tech or heavy machinery, the obligations were deferred even further out to 2 August 2028.
OK, I guarantee people listening to this are breathing a massive sigh of relief right now. December 2027, August 2028. We've got plenty of time.
Let's worry about it next year. Do not relax. That is a highly dangerous takeaway.
While the heavy high-risk conformity file obligations were shifted, you must remember that the Article V prohibitions, the explicit list of AI practices that you absolutely cannot do under any circumstances, and the Article IV AI literacy duties have been actively enforced since 2 February 2025. So just to clarify, if your system uses subliminal manipulation or social scoring or untargeted facial scraping, which are strictly prohibited under Article V, you're already breaking the law today. Completely.
If your system violates Article V, no perfectly formatted conformity file will save you. That gate is already closed and enforcement is active. The conformity file we are discussing is solely for systems that are legally permitted to exist but are classified as high-risk.
But I'll push back again. If the conformity file deadline for my permitted high-risk system isn't until Q4 2027, why on earth am I listening to a deep dive about billing it right now in 2026? Why not just wait until the summer of 2027 to start gathering all these logs? Because of the mise-en-place principle we discussed in Segment 1, if you wait until Q4 2027 to start gathering evidence for an AI system you shipped in 2024 or 2025, you are guaranteeing that you will be building a scrambled pile, not a defensible file. You will be trying to reconstruct complex data provenance and initial evaluation metrics from memory, or worse, from engineers who have since quit and left the company.
Right, the logs might not even exist anymore. You must assemble it now. You must build the documentation habit into your engineering workflow today so that it naturally grows into a living, truthful reflection of the system by the time the massive legal obligation officially bites in 2027.
We've talked a lot about what goes into the file, the honest way to date it, and when it's due. Let's shift our perspective entirely. Let's talk about how an auditor actually reads this folder once they demand it.
Because they don't sit down with a cup of coffee and read it like a novel, starting at page 1 and admiring the beautiful corporate prose. They read it like a forensic detective looking for a lie. They absolutely do.
And to survive that kind of hostile reading, you need to adopt what we call the expert mental model. We refer to this framework as the chain of custody. Let's define chain of custody in this context.
A chain of custody is a rigorous mindset where every single claim, metric, or assertion made in your conformity file traces directly back to a dated, versioned, underlying artifact with absolutely no unexplained gaps in the logic. Contrast that for me with how a beginner or a junior compliance officer builds a file. What does the wrong way look like? A beginner builds what we call a static binder.
They write sweeping claims that flow completely free of evidence. A beginner's file might boldly state in slot 4, our facial recognition model is 94% accurate. It sounds great on the page.
It makes the executives happy, but the claim is completely unmoored. There is no proof attached to it. And the expert.
An expert, however, builds a chain of custody. Their file says model version 2.3 demonstrated 94% overall accuracy on held out test set version C dated March 14th, which is detailed exhaustively in the evaluation report located in appendix 4B. It points directly to their seats.
There's nowhere to hide. Yes, because examiners test files by pulling on threads. If they pull on the 94% accurate thread in the beginner's binder, the trail goes cold immediately and the auditor instantly suspects fraud or incompetence in an expert's chain of custody.
The auditor pulls the thread and it leads straight to a time stamped verifiable engineering artifact. Let us look at a very serious real world example of where a broken chain of custody leads to operational disaster. I want to bring in example 2 from our sources.
The U.S. Federal Trade Commission's 2023 ban on Rite Aid's use of facial recognition. That's a vital case study. Now, I want to state this impartially as a factual regulatory case study based on the FTC's findings.
The FTC found that Rite Aid deployed facial recognition technology without reasonable safeguards. This lack of oversight led directly to false positive matches, which then led store staff to improperly stop, search and accuse innocent shoppers. Yes.
Crucially, the FTC found these errors disproportionately affected people of color. The ultimate result. The FTC slapped them with a devastating five-year ban on using the technology entirely.
It is a profound and sobering example of governance failure. If we analyze why Rite Aid failed through the specific lens of the conformity file framework we are discussing today, they categorically failed the chain of custody test, specifically in what would be slot 4, the performance metrics. If they had a structured file at all prior to deployment, their metrics almost certainly relied on a single flattering headline average.
Something like, the vendor says the system is 95% accurate. But that single average completely hid the localized, uneven harm happening in the real world. Exactly.
A rigorous chain of custody approach in slot 4 explicitly forces an organization to justify the metrics and disaggregate the testing data by demographic group. If they had built that chain of custody, the uneven error rates across different demographics would have been glaringly visible in the documentation before full deployment. They could have seen it coming.
It would have allowed them, or forced them, to implement stricter human oversight safeguards, retrain the model, or halt the rollout entirely until the unacceptable risks were mitigated. The chain of custody isn't just paperwork. It is an early warning system for catastrophic risk.
OK, here is a major edge case that I know a lot of engineering teams are wrestling with right now, and it completely complicates this chain of custody idea. What if I didn't actually build the core AI model? What if my system is really just a software wrapper built around a massive foundation model, a general purpose AI, or GPI from a giant vendor like OpenAI, Google, or Amtropic? Doesn't their massive, publicly available model card replace the need for my own conformity file? They've already done all the testing, right? This is perhaps the single most common trap for modern developers building on APIs. The short, definitive answer is no.
Their model card does not replace your file. The upstream GPI provider absolutely has their own stringent obligations under the AI Act. They must provide technical documentation, ensure copyright compliance, and publish summaries of their training data.
But their model card is merely one input for your slot 2 data provenance. It does not replace your burden. Because I'm the one defining the specific high-risk intended purpose in slot 1. They just built a general tool.
Precisely. You're taking a general tool and applying it to a specific high-risk environment. You still must document exactly how you integrated their model, how you constrained its outputs using prompts or fine-tuning, and how you rigorously tested it for your specific intended purpose in slot 4. And critically, you must honestly document the gaps in the chain of custody that are outside your control.
What do you mean by gaps outside your control? If the giant upstream vendor refuses to disclose exactly what copyrighted data they trained their model on, you do not simply leave slot 2 blank and shrug your shoulders. You state the gap explicitly. Upstream vendor did not disclose underlying training data.
Therefore, we implemented highly rigorous red teaming and boundary testing in slot 4 to compensate for the unknown vulnerabilities. You prove you engineered around the blind spot. You take responsibility for what you can control.
Okay, let's say we've done it all perfectly. We have a pristine chain of custody on launch day. The file is immaculate.
The CEO signs it. We throw a launch party. Fantastic.
But AI isn't static. It learns. It interacts with new populations.
It drifts over time. How do we keep this beautiful file from decaying into a work of fiction six months later? That operational challenge brings us to the final ongoing piece of the puzzle. Slot 9, postmarket monitoring, which is legally required under Article 72.
This is the structural mechanism that prevents the file's truth from decaying due to model drift or changes in the operating environment. I think about it like a passport photo. On the day the photo is taken, it is a perfectly accurate, verifiable representation of what you look like.
But five, ten years later, if you haven't updated it, you age, you change your hair, and that photo no longer matches the reality of the person standing at border control. A conformity file on the day of launch is just a passport photo. That is a brilliant analogy.
To stop the conformity file from becoming an outdated, useless passport photo, the risk management analysis you did in slot 5 must be treated as a living, breathing loop, intimately tied to the real-world field data you are continuously collecting in slot 9. You need a fixed, mandatory review cadence. Give me a concrete example of how this loop actually works in practice. How does field data change the file? Let's go back to our initial hardware store example.
You justified a biometric surveillance system for 12 specific stores that had documented exceptionally high rates of violent incidents. That proportionality judgment in slot 5 was completely sound at that specific scale and location. But a year later, the system is working well, and corporate decides to expand the rollout to 200 stores nationwide, most of which are in quiet suburbs with absolutely no history of violence.
The facts on the ground just change completely. The legitimate aim doesn't match the new locations. Completely.
The harm to non-targets scanning hundreds of thousands of perfectly safe shoppers in quiet suburbs has massively increased. But the legitimate aim-stopping severe violence doesn't apply to those new stores. If your conformity file is just a static snapshot from launch day, you are now operating illegally because your proportionality judgment is completely stale.
So what happens? If your file is a living loop, the field signal of expanding the deployment footprint triggers a mandatory reopening of the risk assessment. You have to re-justify this system for the new reality. So this massive file, it's literally never done.
It is complete only in the sense that a botanical garden is complete. The structure is there. The pathways are laid out.
The initial plants are in the ground. But it stays true, healthy, and defensible only if it is continuously tended. Field signals like disaggregated error rates creeping up over time or human operators constantly overriding the AI's recommendations because they don't trust it must feed directly back into your risk management loop to keep the file honest and compliant.
OK, let's step back, take a breath, and summarize the massive journey we've just taken through Topic 5.6. We've moved from a scattered, frantic pile of Jira tickets and GitHub logs to a highly structured, defensible Annex 4 folder. We've learned how to identify our specific legal role as either a provider carrying the heavy file or a deployer holding our own vital oversight pack, and we mapped the trapdoors between them. We've drafted the necessity case step by step, learning how to avoid the Bunnings disaster, and understood the absolute non-negotiable necessity of signing the declaration last and signing it honestly.
We've anchored ourselves to the post-omnibus deadlines of 2027 and 2028, and finally, we've set up a living loop to ensure our chain of custody stays intact against the inevitable forces of model drift. If we connect all of these tactical steps to the bigger picture, assembling this conformity file is ultimately the difference between having good operational reasons in your head and having admissible legal evidence in your hand. When the regulator's letter inevitably arrives, robust governance ensures that an investigation becomes a simple, confident document handover rather than a frantic, doomed scramble to invent history.
And that leaves us with one final, provocative thought for you to mull over as you go back to your teams. Think about the AI systems running in your organization right now, the ones actively making decisions, flagging risks, or interacting with your users. If a regulator or an auditor knocked on your door tomorrow morning, would you be offering them an essay of excuses? Or would you be handing over a living, breathing chain of custody? Governance isn't about avoiding risk entirely.
It's about making your boldest decisions survivable. Thank you for joining us on this Deep Dive.
Real cases
These are real, documented cases. Each shows what a conformity file is for by showing what happens when the evidence is not assembled in advance.
Example 1: Bunnings facial recognition (Australia, OAIC determination 2024). This is the anchor. Bunnings ran facial recognition across 63 stores in Victoria and New South Wales from 2018 to 2021. The system captured every visitor's face, converted it to a mathematical vector, and matched it against a database of persons of interest; non-matches were deleted almost immediately and matches were escalated to a trained team for a human decision (OAIC, 19 November 2024). Privacy Commissioner Carly Kind found breaches of the Privacy Act 1988: collecting sensitive information without consent, failing to take reasonable steps to notify people, failing to have practices and procedures to ensure compliance, and insufficient transparency in the privacy policy. The Commissioner did not impose a fine, citing Bunnings' cooperation, but ordered it to cease the practice and publish a statement of its failures. The lesson for the conformity file is exact: the company may have had operational reasons and even a genuine safety problem, but when the regulator asked for the evidence of necessity, proportionality, notice, and assessment, in one place, it did not exist in the required form. A conformity file is that evidence, assembled before the question is asked. (Note for currency: the Administrative Review Tribunal later reviewed aspects of the matter; the teaching point about assembling evidence in advance stands regardless of the review's outcome.)
Example 2: Rite Aid facial recognition ban (United States, FTC 2023). The US Federal Trade Commission banned the pharmacy chain Rite Aid from using facial recognition for five years after finding it deployed the technology without reasonable safeguards, generating false-positive matches that led staff to stop, search, and accuse shoppers, disproportionately affecting people of color (FTC, December 2023). The failure was not only the technology; it was the absence of documented testing, accuracy assessment, and human-oversight design that a conformity file would have forced into existence before deployment. This case is the owned anchor of another topic and is referenced here only to show the pattern. (see Topic 4.4)
Example 3: The EU AI Act's own text as the specification. The clearest "example" of what belongs in the file is the law itself. Annex IV (nine documentation elements), Article 9 (risk management), Article 47 and Annex V (declaration of conformity), and Article 43 (assessment route) are not commentary; they are the checklist. A provider who reads Annex IV as a table of contents and fills each slot from existing artifacts has a file. A provider who treats the law as background reading has a pile. The globally relevant point: even outside the EU, this structure has become the de facto template that serious buyers, insurers, and courts recognize, which is why assembling it is worth doing even where it is not yet legally compelled.
Example 4: A deployer's parallel pack (illustrative of Article 26). Consider a hospital in Ireland that buys a diagnostic support model from a vendor. The hospital is a deployer. Its file is not the vendor's Annex IV documentation; it is the vendor's declaration of conformity plus the hospital's own records: proof that clinicians using it were trained, that human oversight is real (a doctor signs the decision, the model advises), that the system is used only for its stated purpose, that logs are retained, and that a fundamental rights impact assessment was completed where required (Article 27). (see Topic 10.4) The lesson: the file follows the role. A deployer who assumes the vendor's file covers everything discovers, when a patient is harmed, that the oversight and impact-assessment records were theirs to keep and are missing.
The pattern across these cases. Read together, the examples show one recurring shape, and naming it makes the file's purpose concrete:
- The technology often worked well enough that the failure was not primarily technical.
- What was missing was assembled evidence that the deployment was necessary, proportionate, and disclosed, produced before the harm, not after the question.
- The performance failures that mattered were uneven ones (false positives falling on a group), which a headline average hid and a disaggregated file would have surfaced.
- The role (provider or deployer) decided who owed what, and buying the system never outsourced the duty to prove it was thought through.
- In every case, a conformity file assembled in advance would have converted an investigation into a document handover.
That shape is why this topic exists: the file is the difference between "we had reasons" and "here is the evidence of our reasons, dated and traceable, before you asked."
Example 5: ISO/IEC 42005:2025, the international standard for the assessment inside your file. In 2025 the International Organization for Standardization and the International Electrotechnical Commission published ISO/IEC 42005:2025, "Information technology, Artificial intelligence, AI system impact assessment" (ISO, published 2025). It is a guidance standard, not a legal mandate: it gives organizations a repeatable structure for assessing an AI system's impact on individuals and society across the lifecycle, and its Annex A maps that assessment back to the AI management-system standard ISO/IEC 42001:2023 so a team already running a management system does not duplicate the work. The reason this matters for the conformity file is direct: the necessity and proportionality reasoning that lives in your Article 9 risk-management slot is exactly the kind of impact assessment ISO/IEC 42005 structures. A provider who conducts the impact assessment to a recognized standard, and cites it in slot 7 (standards applied), turns a slot that would otherwise say "no harmonized standard exists yet" into a slot that says "we applied ISO/IEC 42005 to structure the impact assessment." That is a stronger file. (Marker: emerging. ISO/IEC 42005 is newly published and is not yet a harmonized standard under the EU AI Act, so it does not by itself create a presumption of conformity; treat it as a credible engineering framework you chose, not as legal cover, and re-verify its status before relying on it.) The globally relevant point is that the "assemble the evidence" discipline this topic teaches is now codified in an international standard, which is why serious buyers and insurers increasingly expect it even where no law compels it. (see Topic 6.3)
Where people go wrong
- "The conformity file is paperwork you write at the end." Wrong, and expensively so. The file is the assembly of evidence you generated while building and testing the system. If you treat it as an end-of-project writing task, you will be reconstructing facts from memory, which is exactly the position that turns a compliance gap into an indefensible one. Assemble as you build; the file is a folder that grows, not a report you draft.
- "If we self-assess, no one will read the file, so it can be rough." Dangerous. Self-assessment under Annex VI means no outside body reads it before market, but a regulator, a court, an insurer, or a claimant's lawyer can demand it later, and a biometric system may force the notified-body route (Annex VII) where a stranger reads it up front. Assemble every file to the standard of a competent outsider, because you do not control who opens it or when.
- "We bought the system, so the vendor's file covers us." Only partly. A deployer inherits real obligations under Article 26: oversight, monitoring, logging, and in defined cases a fundamental rights impact assessment (Article 27). The Bunnings case shows a deployer of a third-party capability still carrying heavy duties of consent, notice, and assessment. The file follows the role; buying does not outsource responsibility.
- "A signed declaration is a formality." The declaration of conformity (Article 47) is issued under the provider's sole responsibility. Signing it is asserting the file behind it is true. Sign last, after reading the file, never first. A signature on an unread file is not a formality; it is a liability with your name on it.
- "Overall accuracy is enough for the performance slot." No. Annex IV asks for the appropriateness of the metrics and the accuracy for specific persons or groups. A headline accuracy figure that hides a high false-positive rate on one demographic is precisely the failure the Rite Aid and Bunnings cases turned on. The file must show disaggregated performance, not a flattering average.
- "An undated file is fine as long as the facts are right." An undated file cannot be trusted because a reviewer cannot tell which version of the system the evidence describes. Every artifact needs a version and a date, and the lifecycle-changes slot (Annex IV point 6) must record what changed and when. A substantial modification (Article 43(4)) can even require a fresh assessment; without dated change records you cannot prove you triggered or avoided that.
- "We will build the file to the August 2026 high-risk deadline." Stale. The 2026 Digital Omnibus deferred stand-alone Annex III high-risk obligations to 2 December 2027 and product-embedded high-risk to 2 August 2028 (Council of the EU, 29 June 2026). Building to the old date wastes effort on the wrong schedule. Note also that prohibitions and literacy already apply, so parts of your obligation are live now.
- "If we cannot prove necessity, we can add it later quietly." The necessity and proportionality reasoning is the heart of the risk-management slot and the exact thing Bunnings could not produce. If it does not exist, write it honestly and dated, and say when it was written. Backdating a fabricated version converts a documentation gap into a fraud, which regulators treat far more harshly than an honest late document.
- "The foundation model's documentation is our documentation." Wrong, and it is the most common gap in files built on a licensed large model. The upstream general-purpose AI provider owes its own technical documentation and training-content summary under the Act (applicable since 2 August 2025), but that describes the general model, not your specific high-risk system built on it. (see Topic 5.5) Your file must still document how you integrated, configured, constrained, and tested that model for your purpose, and must record honestly where the upstream provider would not disclose what you needed. The fix: treat the vendor model card as an input to slot 2, then write your own integration and testing evidence; where a disclosure gap exists, name it and show how your own testing compensated, rather than leaving a silent hole.
- "We scoped the file to how the system is supposed to be used." Incomplete. Article 9 requires the risk-management system to address reasonably foreseeable misuse, not only intended use. A watch-list tool designed for human-reviewed flags is foreseeably misused if a manager treats a flag as an automatic ban, and a file that ignores that path has a hole exactly where real harm happens. The fix: in the risk-management slot, list the two or three ways a rushed or incentivized human will actually misuse the system, and show the control (human-in-the-loop policy, override logging, reviewer training) you built against each. Scope to the path people take, not the path the design assumes.
- "Post-market monitoring is next year's problem, so slot 9 can be a placeholder." Dangerous, because the file's truth decays the day you sign it. Slot 9 (post-market monitoring, Article 72) is what carries a field problem back into the risk-management system and keeps the whole file current; a placeholder there means that by the time a regulator opens the file, its performance claims describe a model that has since drifted. (see Topic 4.5) The fix: seed slot 9 from your Module 3 incident log and post-incident review (see Topic 3.4), name concretely what you watch (error rates by group, override frequency, new failure reports), how often, and the route by which a field signal reopens the risk assessment. A living slot 9 is what makes the quarterly review in the scenario real rather than aspirational.
- "A high-risk classification means the harmonized standards will tell us what to do." Not yet. As of 2026 the harmonized standards for the AI Act are still being finalized, so most files cannot claim a presumption of conformity from a harmonized standard and must instead describe the other solutions applied (for example ISO/IEC 42001 for the management system, ISO/IEC 42005 to structure the impact assessment) and justify the engineering choices. The fix: in slot 7, state plainly that no fully harmonized standard yet applies, name the international standards and internal methods you used instead, and explain why they meet the requirement. Waiting for a harmonized standard to appear before assembling is how the deadline arrives with an empty file.
- "A filled slot is a passing slot." Not the same thing. A slot can contain a document that fails its own test: a performance slot with only a headline average, a data slot pointing to an untraceable source, a risk slot describing function but no necessity case. Completeness (something in every slot) is easy to fake; a passing file has the right, traceable, dated evidence in each slot, tested against the one question a reviewer asks of it. The fix: grade each slot against the reviewer's question in section 3H, not against whether it is empty, and treat a filled-but-failing slot as a gap.
Questions people ask
- What is conformity file?
- The single assembled evidence pack for a high-risk AI system, structured on the EU AI Act's Annex IV technical documentation, that lets a reviewer open it and follow the system from intended purpose through data, testing, risk decisions, and human oversight to the signed declaration. More on Conformity file
- What is provider?
- Under Article 3(3) of the EU AI Act, the party that develops an AI system (or has it developed) and places it on the market or puts it into service under its own name. The provider owns the Annex IV documentation, the declaration of conformity, the CE marking, and registration. More on Provider
- What is deployer?
- Under Article 3(4), the party that uses an AI system under its own authority in the course of its activity. The deployer owns the Article 26 obligations: use per instructions, human oversight, monitoring, log retention, incident reporting, and (in defined cases) a fundamental rights impact assessment. More on Deployer
- What is Annex IV?
- The list in the EU AI Act of the nine elements of technical documentation a provider of a high-risk system must draw up and keep current, referenced by Article 11. It functions as the table of contents for the conformity file. More on Annex IV
- What is article 9 risk management system?
- The continuous, iterative process of identifying, evaluating, and mitigating the risks a high-risk AI system poses across its life. This is where the necessity and proportionality reasoning for building the system lives.
Keep going
This lesson builds Evidence collection and audit-ready documentation, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.