Skip to main content

The evidence annex: wiring every artifact so the Module 13 audit finds a paper trail, not a scramble

The short answer

The side with the paper trail wins

Li v. Liu was decided not by talent or honesty but by provenance: the party who could document who made the AI-generated image, through prompts, parameters, iterations, and an identifying watermark, prevailed, and the party who could not show where their copy came from lost. Every serious challenge to an AI decision, from a court to a viva, resolves the same way, and the evidence annex is your standing paper trail prepared before the challenge arrives.

What you will be able to do

  • Assemble an evidence annex that wires every governance artifact you have built across the program into one indexed register, so any piece of evidence can be produced on demand rather than hunted for.
  • Structure each annex entry so it carries the six things an auditor needs: what the artifact is, what it proves, who made it, when, where it authoritatively lives, and which system and obligation it maps to.
  • Stamp provenance on every entry, recording who made what and when in a way that can be verified as unaltered, so the annex answers the Li v. Liu question ("who made this, and how do you know") for your whole AI estate.
  • Map evidence in two directions, tying each artifact to the system it concerns and to the obligation it discharges, so a request framed either way ("show me everything about this system" or "prove you meet this duty") lands on the right artifact.
  • Read your own gaps as findings, using the annex to expose which obligations have no evidence and which systems have no trail, and treating each blank cell as a task, not an embarrassment to hide.
  • Keep the annex live, maintaining it as artifacts change rather than assembling it in a panic at the end, so the register is always current when the audit arrives unannounced.
  • Defend the annex against the charge that it is a cosmetic index over a hollow file: prove that every entry points to a real, dated, provenance-stamped artifact an auditor can open and trust.

The lesson

On November 27, 2023, the Beijing Internet Court issued a ruling that set a new operational reality for artificial intelligence. In a case recorded as Li V. Liu, the court recognized copyright for an AI-generated image. The plaintiff did not win because he typed a single prompt and pressed a button.

Working in stable diffusion, he documented over 150 prompts, mapped his exact parameter adjustments, recorded his iterations, and applied an identifying watermark. When challenged, he produced that precise trail of intellectual work. The defendant took the image and removed the watermark.

When asked where she obtained it, she had no provenance. Against the plaintiff's documented inputs, she could not answer who made the file or how she knew it was hers. The court evaluated this dispute on originality, but decided it in practice on evidence.

They rewarded the party who could present verifiable provenance and penalized the party who relied on assertion. Whether you face a tribunal, a regulator, or an internal board audit, the dynamic is identical. The side that can immediately produce a documented paper trail prevails.

The side that scrambles for evidence does not. In a corporate setting, this failure state looks like a 40-minute hunt for a governance document while an auditor waits, ultimately producing a file that no one can confidently verify as the final version. An examiner watching this process draws an immediate conclusion.

If a control's evidence cannot be produced on demand, that control is aspirational, not operational. The underlying work is instantly rendered suspect. To eliminate this risk across your entire AI estate, you build an evidence annex.

This is not a narrative essay you write from scratch to summarize your compliance. It is a single, unassailable indexed register that wires together every governance artifact your organization has already built. It converts the abstract mandate to govern an AI system into a concrete data engineering defense, ensuring any required piece of evidence can be produced and verified in seconds.

When preparing for an audit, the most common reflex is to gather every compliance document and copy it into one massive centralized folder. This approach breaks the moment an engineer updates the original file in its native repository. The copied version sitting in the audit folder is instantly disconnected from reality.

Maintaining a warehouse of copies creates a museum of stale duplicates. Presenting these out-of-date files to an investigator actively misleads them about the true state of your systems. The annex avoids this by operating strictly as an index.

The authoritative files remain exactly where they live, under version control in their original environments. Instead of holding the files themselves, the register relies on digital pointers, clean vectors that connect the annex to external repositories without moving the original data. By leaving the data in place, a legally defensible annex acts as a live, accurate map to the truth, rather than a stagnant warehouse of liabilities.

To keep this map rigorous, every entry must adhere to a strict rule, populate exactly six specific fields, or be reject entirely. First, what it is and what it proves define the artifact's purpose. Next, who made it and when, lock in the author and version timestamp.

Finally, where it lives and what it maps to provide the exact pointer and regulatory duty involved. If a team cannot fill the what-it-proves field, that artifact is merely corporate clutter, it discharges no actual duty and does not belong in the register. Filling these six fields systematically preempts every hostile question an examiner plans to ask you under live pressure.

A document detailing how a model was tested is merely a claim. True evidence requires verifiable proof that the document is genuine, authored by a specific person, on a specific date, and completely unaltered since. Regulators are writing this exact standard into law.

China's 2025 labeling measures mandate that implicit metadata, a content identifier and provider record, must travel inside the AI-generated file itself. Because the stakes of AI systems vary, you must deploy a hierarchy of provenance strength. The security method proving a file's integrity must match the risk level of the system it documents.

A system card for a low-risk internal tool can rely on standard version control history, but forensic logs for a high-risk model demand mathematical lockdown, utilizing cryptographic hashes to prove byte-for-byte integrity. Adversarial counsel and forensic auditors know exactly where to look for weaknesses. They will attack any gap in your chain of custody to question whether a critical document was simply written last week.

Unverified files crumble under that kind of adversarial questioning. Hard mathematical provenance is the only reliable shield for your high-risk models. Auditors will never restrict themselves to a single line of questioning.

They demand evidence perpendicularly, first by filtering down to a specific AI system to review its entire lifecycle. Then, they pivot to verify a single obligation, like data governance, across every system in the entire enterprise simultaneously. The vertical, system-based view serves board-level inspections.

It provides a seamless, end-to-end trail for one model, from initial training data sourcing through to final deployment. The horizontal, obligation-based view serves regulatory frameworks. It allows you to prove compliance against specific laws, keeping in mind that provenance duties are strictly local to where the system actually operates.

For example, the European Union's AI Act imposes strict, localized record-keeping duties on high-risk standalone systems, requiring precise mapping to meet the 2027 deadline. By contrast, a U.S. system must map to a patchwork of state-level bias audit rules or frameworks like the NIST Risk Management Framework. An annex mapped exclusively to a single home country default leaves your organization legally exposed abroad.

The register must function as a multi-jurisdictional matrix. Applying this two-way mapping produces an immediate byproduct, the painful exposure of blank cells, representing missing, stale, or unfindable artifacts. The most common compliance mistake is attempting to quietly hide or delete these rows, to present a seamless facade to the examiner.

A concealed gap operates as a landmine. Once an auditor discovers a hidden deficiency, their trust in the integrity of the entire estate is instantly destroyed. You must reframe the blank cell as a governed finding.

Turn it into a deliberate disclosure by assigning a named owner, a remediation date, and a designated interim control. A disclosed, governed gap demonstrates mature oversight. A hidden gap demonstrates deception.

Consider a governance lead facing a hostile board VIVA, unexpectedly queried on a high-risk model's failure rates. With a proper annex, she pulls the model's conformity file instantly, supported by verified cryptographic hashes. That response is impossible if you rely on panic assembly.

Building the annex the week before an audit guarantees it will be obsolete upon arrival. The systemic solution is to wire the maintenance of the annex directly to live event triggers, entirely abandoning scheduled calendar reviews. Operational realities, like retraining a model or logging a critical incident, must automatically ping the annex to demand an immediate update for that specific row.

The annex must breathe with the AI estate, keeping your evidence perpetually audit-ready at any given second. To reach that state, you must lay out the six-field register, wire in your existing artifacts using clean pointers, and stamp verifiable provenance onto every entry. Once the index is assembled, aggressively red-team the register.

Hunt for dead pointers and un-evidenced obligations yourself, before an auditor does it for you. Under frameworks like the EU AI Act, the ultimate legal reality is clear. The burden of demonstration sits entirely on the organization.

Merely doing the governance work carries zero legal weight, if you cannot instantly produce and authenticate the records on demand. The evidence annex is the mechanism that guarantees readiness. When the examiner arrives, you do not offer assertions.

You present a massive, unassailable trail.

The ideas, one by one

The annex is an index with provenance, not a warehouse or an essay

It points to every artifact where it authoritatively lives and stamps each pointer with who made it, when, and proof it is unaltered. Copying artifacts into one folder creates a museum of stale duplicates; writing a fresh narrative creates an essay, which is a claim without a trail. The value is entirely in the wiring.

Six fields or no row

Every entry carries what it is, what it proves, who made it, when, where it authoritatively lives, and what it maps to. Each field answers a question the auditor will otherwise ask you live under pressure. A row missing any of them is a false promise of evidence, and an artifact whose "what it proves" is blank may be evidence of nothing.

Provenance turns a document into evidence

Anyone can produce a document that says a model was tested; evidence is a report you can show a named person produced on a real date and nobody has altered since. Version control gives you most of this for free; where artifacts live outside it, a dated signature or a stored file hash approximates it. Evidence that can be tampered with is decoration, not proof.

Map in two directions or fail half the questions

Auditors ask by system ("show me everything about this model") and by obligation ("prove you meet this duty"). Build both lookups over the same rows, using the AI systems inventory for the system axis and the binding frameworks for the obligation axis, and the annex answers either shape with equal ease.

Gaps are findings, and finding them yourself is a gift

The most valuable output of building the annex is the blank cells it exposes: obligations with no evidence, systems with no assessment, stale cards, unfindable reports. A gap you mark as a dated finding with an owner is a task and a signal that you govern yourself; a gap you hide is a landmine the auditor will step on, and then wonder what else you concealed.

Keep it live, and build it to be inherited

An annex assembled the week before the audit is a scramble with better handwriting, and one that lives in a single person's head leaves when they do. Wire its maintenance to the events that change evidence, give it a named owning role, and name each artifact's accountable maker, so any audit, ninety days out or two years out under a new lead, finds a paper trail.

The burden of demonstration is on you

Under the EU AI Act, demonstrating compliance means producing the evidence on request, not merely having done the work. An organization that did everything right but cannot produce and authenticate its trail has not demonstrated compliance, and the annex is precisely the mechanism that makes the whole program demonstrable, which is why it is the last thing built before the audit reads it.

One annex serves every regime that asks you to show your work

The same register, with an obligation axis that can group rows by the EU AI Act, ISO/IEC 42001, the NIST AI RMF, or a state law, discharges the demonstration expectation of all of them at once. Wiring evidence well once, with provenance and two-way mapping, is leverage across every framework and regulator, which makes the annex an investment rather than a compliance cost.

Match provenance strength to the stakes

A low-risk tool's card can live under version control; the evidence behind a decision that harmed a person, or a log that a forensic reconstruction will examine, deserves the strongest integrity you can give it. Record per entry how each artifact's integrity is shown, so a weak-provenance row on a high-stakes artifact stands out as its own kind of gap before an adversary points at it.

You read it. Now prove it.

Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.

The conversation

The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.

Listen to it as episode 81 of the podcast.

Read the full conversation

So, in November of 2023, a man walks into a Beijing courtroom, and he's there to claim copyright over an artificial intelligence-generated image. Which, you know, is usually a pretty tough sell in court. Exactly.

I mean, it's historically a gray area, but he didn't win with some brilliant, impassioned legal argument. No, not at all. He won with a paper trail, like 150 lines of documented paper trail.

And today, we're going to talk about exactly why that matters to you. Yeah, it matters a lot. So, welcome to the Deep Dive.

Today, we have a very specific mission, and we are pulling from a massive stack of sources to accomplish it. It is a really dense stack today. It is.

We are looking at the translated legal briefings from the Beijing Internet Court on that landmark case, Live-A-Loop. Fascinating case. And we are also pulling apart the full text of the European Union's AI Act, specifically Regulation 2024-1689, along with the recent 2026 Digital Omnibus Updates.

Plus, we've got the NIST AI Risk Management Framework 1.0 and ISO-IEC 42001.2023. Right. And on top of all that, a stack of highly confidential field notes from actual corporate auditors. I mean, it is a dense stack, but it all points to a single urgent problem that, well, that every organization deploying AI is currently facing.

Exactly. The mission of this Deep Dive is figuring out how you take the scattered pile of AI governance artifacts you've been building over the last year. You know, the model cards, the impact assessments, all of that.

Right. The incident logs. Yeah.

The data sourcing agreements. And wire them into one indexed provenance stamped register. Because the goal is to ensure that when a regulatory audit finally arrives, like a Module 13 audit, they find a pristine paper trail.

Yes. It's not an entire organization just scrambling in the dark. And that scramble is what we really want to eliminate today.

Because in the world of AI governance, I mean, a disorganized defense is essentially indistinguishable from no defense at all. It really is. Which brings us back to that Beijing courtroom.

So the case is recorded as Li v. Liu, and it completely changes how we should think about AI in evidence. The Beijing Internet Court did something no court had really done before. Yeah.

They ruled that an AI-generated image, in this case, it was created using the stable diffusion software, could actually carry copyright, and that the person who made it owned that copyright. But, you know, the ruling itself isn't even the most important part for you. As someone managing AI at an enterprise level, what's fascinating here is why the plaintiff won.

Yes. And why the defendant lost. Right.

The mechanics of the victory are everything here. The plaintiff, Mr. Li, he didn't just walk in and say, you know, I typed a prompt and got this cool picture, so it's mine. Right.

If he had done that, he would have definitely lost. Oh, instantly. Instead, he produced this documented step-by-step trail of his entire creative process.

He brought over 150 prompts to the court. He didn't just show the final output. He showed his specific parameters.

He showed the sequential iterations, like demonstrating exactly how he adjusted and readjusted the weights, the negative prompts, the artistic styling. Until the picture matched his vision. Exactly.

He proved his intellectual labor. And crucially, he had an identifying watermark embedded on the image. So he maintained a perfect chain of custody for his own intellectual achievement.

Now, if you look at the defendant, Mrs. Liu. Right. Her defense completely collapsed.

It did. And for the exact opposite reason. She had taken his image and used it.

She had removed the watermark that named its maker. And well, when the court asked her where she obtained the image, she just could not answer. She had zero provenance.

Zero. She couldn't answer the single question that decides so many disputes in the AI era, which is, who made this and how do you know? Two people. One image.

And the entire case turned on a paper trail, which this introduces our first core concept today. And we're going to name these principles clearly so you can anchor your enterprise strategy to them. The first spine of AI governance is this.

The side with the paper trail wins. That is the immutable law of governance, and it completely transcends international borders and legal jurisdictions. Yeah.

Now, you might be thinking, look, I am not fighting Chinese copyright law. I'm deploying an enterprise AI tool for logistics or human resources. Right.

Completely different context or so it seems. Exactly. But the underlying mechanism of accountability is identical.

I mean, you face regulators, you face cyber insurers, you face your own enterprise board of directors. Oh, absolutely. And when they challenge an AI decision, say, an algorithmic hiring tool that allegedly discriminated against a candidate or maybe a loan approval model that seems to favor certain demographics, they don't reward the most sincere organization.

Right. They don't care about your good intentions. Not at all.

They don't reward the team with the best intentions. They reward the one with verifiable provenance. The side that can produce the trail prevails.

The side that scrambles does not. I mean, think about it. Provenance didn't magically create the legal copyright in that Beijing courtroom.

Right. It didn't create it. It's what allowed the court to see that the right was real.

If you did the underlying work, like if your engineering team lawfully sourced your training data, if you spent weeks testing your model for bias. If you mapped your edge cases. Yeah.

Right. But if you can't produce the trail of that work when asked, you are basically in the exact same position as an organization that never did the work at all. If we connect this to the bigger picture of enterprise risk, an evidence annex is necessary in every defense, though it's sufficient in none.

What do you mean by that? Well, it wins you a fair hearing. It doesn't guarantee a favorable outcome if your actual AI system is fundamentally flawed or illegal. OK.

Right. But without the paper trail, you don't even get the chance to defend the substance of your engineering. You just lose by default on the administrative failure.

That makes total sense. OK. Let's unpack this.

So what does this actually mean for your organization? If the paper trail is what wins the case or what passes the regulatory inspection, what exactly is the mechanism we use to build it across an entire enterprise? Right. The how-to. Because we definitely know what doesn't work.

I mean, I've been in audits where an examiner asks for a document and you can literally hear the Slack notifications firing in the background. Oh, I know that sound. As three managers panic message each other just trying to find a PDF.

It's agonizing. We call that the scattered shoebox of receipts. It doesn't work.

It absolutely doesn't work. And that brings us to the need to define what the proper tool actually is and probably more importantly what it isn't. OK.

The rule, which is our second spine today, is this. The annex is an index with provenance, not a warehouse or an essay. OK.

Let's unpack this right up front. We are using the term evidence annex. Let's define the key term for everyone listening.

What are we actually talking about when we say evidence annex? The evidence annex is a single indexed register. That's it. It points to every governance artifact an organization has produced.

Conceptually, it is just a map. But to truly understand its power, we have to look at the common mistakes that highly intelligent people make when they try to build one. Right.

Let's talk about the instinct to hoard because it is so universal. When general counsel or the compliance team says, we have a major AI audit coming up in 90 days, the very first human instinct of the project manager is to create a giant folder on a shared drive. Yes.

The dreaded mega folder. Exactly. They name it Audit Materials 2026, and they just start copying and pasting every document, every spreadsheet, every Jupyter notebook and every PDF they can find in that one giant warehouse.

And that instinct is fatal to an audit defense. Completely fatal. Really? Just because of organization? No.

It's about versioning. If you treat the annex as a warehouse, you are actively destroying your own governance. We need to think about the mechanics of versioning here.

The moment you copy a document from its original location into your giant audit folder, you have created an unlinked duplicate. Oh, right. Now, what happens three days later when the engineering team updates the original model card? Because they say, retrain the algorithm with new seasonal data.

Right. So the original document in their Git repository gets updated, but the copy sitting in the audit folder stays exactly the same. Yeah.

It's instantly stale. Precisely. Exactly.

Within a matter of weeks, your beautifully organized audit folder is no longer a reflection of reality. It is a museum of stale versions. A museum of stale versions.

I like that phrase. And if you hand a museum of stale versions to an auditor, it actively misleads them. It asserts a false current state of your AI estate.

The Gangnam rule of governance is that artifacts must stay where they live, under version control, in their owning system or repository. The annex doesn't hold the files. It points to them.

It points to them. Okay. So it's not a warehouse.

But you also mentioned it's not an essay. And I see this happen all the time, too. Oh, constantly.

A governance team panics before an audit, and instead of indexing their evidence, they sit down and write a 50-page narrative. They write this beautiful, sweeping prose document explaining all the great things they've done to keep their AI safe, fair, and robust over the last year. Yeah.

If your instinct on hearing the word annex is to author a fresh narrative of everything you've done, you have to stop immediately. Stop typing! Right. An essay is a claim without a trail.

The annex contains almost no original prose. It is strictly pointers and provenance metadata. The value of the annex is entirely in the wiring, not in new writing.

That's a great distinction. An auditor doesn't want to read a story asking them to trust you. They want to look up a system, follow a pointer to the actual engineering artifact, and see for themselves.

I love analogies. So think of it like the index at the back of a massive reference book. The index doesn't rewrite the book.

It just tells you, you know, hey, you want to know about the hiring algorithms, data privacy controls? Go to page 412, paragraph 3. That is a perfect analogy. The index directs the inquiry. It doesn't try to answer the inquiry itself.

And there's a third thing the annex is not, which trips up very careful compliance professionals all the time. What's that? It is not the conformity file, and it is not the dossier. Right.

Because if you read the actual text of the EU AI Act or ISO 42001, you hear those terms thrown around constantly. Exactly. So let's delineate them.

A conformity file is a specific bundle of evidence for one particular high-risk system. It's a localized package. Just for one system.

Right. The dossier, on the other hand, is the final presentation or narrative you might assemble to contextualize the audit itself. The evidence annex sits above both of these things.

It's the ultimate enterprise-wide map. So it's the master index. Exactly.

The conformity file for one system is just a single mapped entry within your larger annex. If you only build conformity files in isolation, you still don't have a map of your entire enterprise estate, which means you can't answer systemic questions. Okay.

So if you don't build this indexed map, if you just rely on the warehouse folder or the beautiful essay or people's memories, you end up in a situation you referred to earlier as the scramble. Yes. And this is where the rubber meets the road in a high-stakes audit.

The paper trail versus the scramble. Let's paint the picture of what this actually looks like in a boardroom. An auditor sits across the table and asks for the data impact assessment for, let's say, your demand forecasting model.

Okay. In an organization that has built a proper paper trail, someone opens the annex, filters for the demand forecasting model, clicks the canonical pointer, and the assessment is on the screen in under 60 seconds. Boom.

Done. It clearly displays its author, its date, and its cryptographic version history. Clean, precise, and completely frictionless.

Exactly. Now, in an organization without an annex, the auditor asks the exact same question. Someone in the room says, oh, I think Priya led that assessment.

Let me check her SharePoint folder. But Priya is on paid time off, so they slack three other people on the engineering team. Forty minutes later, after frantic searching, someone produces a Word document titled Impact Assessment Final V3 Realy Final Reviewed dot docs.

Oh, man. I have seen that exact file name. We all have.

And nobody in the room can confidently swear to the auditor that this is the actual final version that the risk committee approved. Forty minutes of dead air in an audit room. That is excruciating.

But it's not just about the awkward silence or the embarrassment, is it? It's about what the auditor's thinking during those 40 minutes. Right. This raises an important question.

What is the silent cost of the scramble? It is entirely corrosive to your defense. Before that poorly named document even appears on the screen, the auditor has silently concluded three devastating things. OK.

What are they? First, if it takes you 40 minutes to find the evidence of a critical control, the auditor assumes the control might not be real. They assume it's aspirational, not operational. Right.

They're thinking they say they check their models for bias every quarter, but they can't actually find the check. So they probably aren't actually checking. Exactly.

It looks like a theoretical policy that nobody actually follows. Second, when you finally produce a document after 40 minutes, the auditor naturally wonders if it's genuine. Oh, because it took so long.

Right. Was this document finalized six months ago during a rigorous review process or was it quietly finished 20 minutes ago in the hallway to cover a gap? Yeah. Third, having caught one scramble, the auditor now suspects everything else you show them.

Even your solid, perfectly executed evidence is tainted by the assumption that it might be improvised. You just lose the benefit of the doubt entirely. And fourth, and to me, this is the scariest when the auditor realizes that your people are your actual risk.

Right. They conclude that your AI governance lives entirely in individuals' memories, not in an inspectable institutional system. And if governance only lives in Priya's head, well, the paper trail walks out the door the day Priya quits.

Yeah, that is an unacceptable risk profile for any enterprise deploying AI at scale. It really is. And this isn't just a matter of best practices or looking professional anymore.

There is severe binding legal weight bearing down on organizations to fix this. Let's talk about the regulatory landscape, specifically the European Union's Artificial Intelligence Act regulation, EU 2024-1689. Yes.

The AI Act is really the center of gravity for global AI regulation right now. Under the Act, providers and deployers of high-risk AI systems carry strict documentation and record-keeping duties. And the critical operative phrasing in the law is that the burden of demonstration sits squarely on the organization.

The burden of demonstrating. Demonstrate compliance doesn't mean saying you are compliant, it means producing the evidence on request. Now, I know there was a legislative deferral recently that gave people a bit of breathing room, right? The 2026 Digital Omnibus Simplification Package changed the timeline.

Can we break down what that actually means for our timeline? Yeah, the 2026 Digital Omnibus was a critical shift. It deferred when the standalone high-risk duties, those are the ones listed in Annex 3 of the AI Act, start to legally bind organizations. It pushed that deadline to December 2, 2027.

And just to clarify, when we say standalone high-risk, we mean things like a pure algorithmic HR hiring tool or an AI system used to evaluate creditworthiness for loans? Correct. Systems that are high-risk by their very application. Then, for high-risk AI that is embedded as a safety component in regulated products, which falls under Annex I, like an AI diagnostic tool inside a medical device or AI in aviation systems, that deadline was deferred to August 2, 2028.

Okay, so 2027 and 2028. But here's the critical takeaway. The deferral changes the calendar, not the principle.

The operative verb is still demonstrate. Precisely. You must be able to show compliance, not merely claim you achieved it.

Work you cannot produce is work you did not do in the eyes of the law. If you walk into that December 2027 deadline with your evidence already wired, indexed and provenance stamped, you meet the regulator with a pristine paper trail. If you wait until Q3 of 2027 to start figuring this out, you start your frantic scramble on the exact day the regulator is finally watching.

Okay, so we've established that the Annex is strictly an index, and we know we have to avoid the scramble to satisfy the law, the auditors, and our own board. Yes. So let's get into the mechanics of this.

If I am building this register, whether it's a highly customized database or just a really very disciplined spreadsheet, what exactly must each entry look like? What are the actual columns? This brings us to the third spine of our methodology. Six fields or no row? Six fields. No exceptions.

No exceptions. The architecture of a row in your evidence annex must be absolutely strict because every single field earns its place by answering a fundamental question the auditor is guaranteed to ask. Okay, let's break them down meticulously.

All right. Field number one is what it is. This is the plain, understandable name of the artifact.

You use plain language, not invented internal engineering codes. So no weird acronyms. Right.

You write conformity file for the CV screening system, not DOC14BV2. The auditor needs to know exactly what they are opening before they click the link. That makes sense.

It's just about legibility. What's field two? Field two is what it proves. This is the specific obligation, control, or claim that this artifact is evidence for.

For example, this artifact proves the training data used for the CV screening model was lawfully sourced and checked for demographic imbalances. Okay, let me push back here for a second. Six fields sounds great in a podcast studio, but in a chaotic enterprise where engineers are just shipping code and compliance is overwhelmed, what if I have a document and I can fill out five of the fields perfectly? Like I know the name, I know who made it, I have the link, but I'm not exactly sure what it proves.

Maybe it's just a general architecture diagram or some meeting notes from a brainstorming session. Can I just leave what it proves blank and keep it in the annex, you know, just in case? Absolutely not. No.

No. If you can't say what an artifact proves, it might be evidence of nothing. A row with a blank what it proves field is a false promise to the auditor.

It just creates noise. Oh, I see. If you can't tie an artifact to a specific obligation or control, it is just clutter and it doesn't belong in the annex at all.

The fields are self-auditing. They force you to justify why you were holding onto a document in the first place. I love that.

The structural discipline forces clarity. It prevents the annex from becoming a dumping ground. Okay, field three.

Field three is who made it. This is the accountable role or person who authored or approved the artifact. Then field four is when.

The exact date and the version number. Evidence has a very strict shelf life in AI. A model card written before a model was retrained on new data is stale evidence.

It's dangerous evidence, actually. Right, because it misrepresents the current state. Exactly.

The date lets the auditor instantly judge the artifact's freshness against the system's life cycle. Which means field five has to be the location. Yes.

Field five is where it authoritatively lives. This is the canonical pointer or URL. It must point to the version-controlled repository like GitHub, GitLab, or an enterprise document management system, not a copy sitting on someone's desktop.

And that link better work. Oh, that pointer. Must resolve.

The dead link in an annex is a scramble waiting to happen mid-audit. And finally, field six. Field six is what it maps to.

Which specific AI system on your enterprise inventory it concerns, and which regulatory framework obligation it satisfies. We'll dive deeper into how mapping works in a moment, but every piece of evidence must anchor to a known system. Okay, so that is the rule.

Six fields or no row. What it is, what it proves, who made it, when it was made, where it lives, and what it maps to. But I want to zero in on fields three and four for a second.

The who and the when. Sure. Because you said earlier that just writing a name and a date on a Word document isn't enough if it can be easily faked.

Which leads us to the fourth spine of our strategy. Provenance turns a document into evidence. This is the absolute core of defensibility.

Anyone can open a Word document, type, model was tested for bias on Tuesday by John, and hit save. That is just a claim. The essay trap all over again.

Exactly. Evidence requires provenance. Provenance is the verifiable, unquestionable account of an artifact's origin.

Who made it, when they made it, and proof that it has not been altered since that moment. And how do we do that? This is achieved through a chain of custody, an unbroken, accountable record of handling. So how do we actually prove that in the real world without turning our compliance department into a CSI forensics lab? Because if I have thousands of artifacts, I can't have a lawyer notarize every single one.

For most organizations, you actually get a significant amount of this for free if your artifacts live under strict version control. Systems like Git, or Dedicated Enterprise Document Management Software, record exactly who committed a file, at what exact timestamp, and what precisely changed between versions. Right.

You can't really fake that easily. Exactly. You can't silently back data commit in Git without leaving a trail.

It provides a foundational layer of provenance. But not all artifacts are created equal, right? A routine update to an internal IT help desk chatbot is one thing. But what if we are talking about the forensic logs for a high-risk medical diagnostic AI that gave a recommendation that ultimately harmed a patient? Yeah, that's completely different.

Or an autonomous driving system that was involved in a collision. What's fascinating here is how you must dynamically match the strength of your provenance to the stakes of the artifact. For low-risk tools, standard version control is entirely sufficient.

But for high-stakes artifacts, you need cryptographic certainty. You need a file hash, or a signed immutable log. Okay, you said file hash.

I know that's a cryptographic term. But for those of us who aren't security engineers, how does that actually work in this context? Because it sounds incredibly complex. It's actually a very elegant mathematical concept.

Imagine a mathematical blender. Mathematical blender. If you take a digital file, say a 40-page PDF detailing the safety testing of your medical AI, and you run it through a cryptographic hashing algorithm like SHA-256, it generates a unique, fixed-length string of characters.

And that string is the hash. That string is the hash, and it acts as a digital fingerprint for that exact file. And what happens if someone goes into that 40-page PDF and just changes one number? Say they change a safety failure rate from 5% to 2% to make it look better.

If they alter even a single comma in that file and you run it through the algorithm again, the resulting hash will be completely radically different. So in your evidence annex, you don't just store the link. For high-stakes items, you store the hash.

An auditor can look at the file you present, run the half themselves, compare it to the record in the annex, and know with absolute mathematical certainty that the document has not been tampered with since the day it was finalized. You can look an auditor or a judge in the eye and say, I can mathematically prove to you that this document is exactly what we say it is. That is incredibly powerful.

It is the gold standard of evidence. And we are seeing global regulators mandate this exact logic in law right now. Look at China's 2025 measures for labeling AI-generated and synthetic content, which went into force on September 1, 2025.

They require what the SAW calls implicit labels. Implicit labels, meaning it's not just a visual watermark you can see with your eyes, like what Mr. Li used in the copyright case. Correct.

It goes much deeper than a visual watermark. An implicit label means provenance metadata traveling inside the file itself, inextricably linked to the content. It records the provider, a content identifier, and the parameters of creation.

So the origin travels with it. The origin of the content travels with it mathematically. The regulator's instinct there is identical to what we are discussing for the enterprise.

Provenance should be attached and mathematically verifiable, not asserted from human memory, or a separate, easily faked spreadsheet. Here's where it gets really interesting to me. By building tamper-evident provenance, you aren't just protecting yourself from external hackers or skeptical auditors.

You are protecting your organization from its own future incentives. Yes. Human nature under pressure is a massive corporate risk.

Tamper-evident provenance prevents anyone internally from quietly backdating a document to cover a gap when they realize a Module 13 audit is looming next week. Right, because people panic. Provenance protects the honest version of events.

If your evidence can be tampered with by a panicked middle manager, it isn't evidence. It's just decoration that will collapse the second a serious examiner pushes on it. Okay, so we've established the architecture.

Every row has six ironclad fields. Every critical entry has mathematically verifiable provenance. But in a large Fortune 500 enterprise, you might have hundreds of AI systems and 10,000 rows of evidence.

Easily. How on earth does an auditor or even an internal governance lead actually navigate that without control effing for an hour? That organizational challenge requires our fifth spine today. Map in two directions or fail half the questions.

Map in two directions. What are the two directions? The system axis and the obligation axis. When an auditor begins an examination, they are going to ask you questions in two distinctly different shapes.

First, they might ask, by system, they'll sit down and say, show me everything you have on the CV screening model. Like a board inspection where they want to see the entire life cycle of one specific tool from cradle to grave. Exactly.

To answer that shape of question, you use the system axis, which is built on your enterprise AI systems inventory. You filter your annex by the CV screening model and instantly you get the whole trail for that tool. So the evaluation reports, the model cards, the incident logs? The data provenance files.

It gathers everything related to that specific technology. But they don't always ask by system, do they? Sometimes they ask by the rule book. Right.

The second shape of inquiry is by obligation. An auditor might sit down and say, prove to me that you meet the data governance duty across all your high-risk systems enterprise-wide. To answer that, you need the obligation axis.

This is built on the specific legal and voluntary frameworks that bind you. So we're talking about mapping to the articles of the EU AI Act, or the specific controls in ISO IC 42001.2023, or the functions in the NIST AI risk management framework. Exactly.

Think about the nightmare if you only indexed your annex by system. You would have to manually open every single systems folder, dig out the data governance document for each one, and hand-assemble a response while the auditor taps their pen, waiting. It's a mini-scramble.

It is a mini-scramble. But with the obligation axis, you just group the rest by the data governance duty, and it instantly outputs every artifact across the entire enterprise that discharges that specific duty. You know, it reminds me of reading a massive, complex novel, like Game of Thrones or something of that scale.

OK, I like where this is going. In those books, you usually have a chronological timeline of events, and you have a character map. If the reader asks, what happened in the year 298 AC, you use the timeline.

But if the reader asks, where was Jon Snow during all these events, you use the character map. Yes. If you only have one index, you can't answer questions about the other without manually flipping through every single page of a 1,000-page book.

That's a phenomenal way to visualize it. And this dual mapping is critical when you apply a global lens to your AI deployments. You cannot just map every system to your headquarters jurisdiction.

Right, because an AI tool deployed globally touches different laws in different places. Provenance duties are fiercely local. Let's say you are a multinational corporation headquartered in Paris.

If you only map your obligations to the EU AI Act, you're going to walk straight into a wall. Why is that? Because your AI algorithmic hiring tool deployed in the U.S. sits under specific state and local bias audit rules, like New York City's local law 144, and your content generation tool deployed in Shanghai needs that implicit metadata provenance we just talked about under Chinese law. The allegation axis must carry the local duties of where the system actually runs, not just where your legal team happens to sit.

It's a many-to-many relationship. One document say, a robust data sourcing log might simultaneously satisfy EU AI Act Article 10, NIST AI RMF Map 2.1, and a local New York bias law. Exactly.

The annex lets you wire one artifact to multiple obligations without duplicating the file. Map in two directions or fail half the questions. It's that simple.

But here is the terrifying part of building this. The moment you build that obligation axis, the moment you list out all the legal duties from the EU AI Act, NIST, and ISO, and you try to map your artifacts to them, you are going to start seeing empty spaces. Yes, you are.

And those empty spaces are the most valuable thing the annex will ever give you. Which brings us to the sixth spine of our strategy. Gaps are findings, and finding them yourself is a gift.

It doesn't feel like a gift, though. It feels like a massive panic attack when you see a blank row next to a mandatory legal requirement. It feels incredibly uncomfortable, yes.

But consider the alternative. Finding a gap yourself 90 days before an audit is a governed task you can manage. Having an auditor find that gap during a live adversarial examination is a wound to your compliance posture.

Let's break down the types of gaps because they aren't all the same and they require different responses, right? Yeah, there are three main types of gaps you will discover when you build the annex. First is the most obvious, missing. The artifact was simply never built.

You look for the Fundamental Rights Impact Assessment, and it doesn't exist because nobody ever did it. Okay, that's a pretty clear cut. What's the second? Second is a bit more insidious, stale.

The artifact exists, but the system changed out from under it. You have a model card, but the model was retrained on entirely new data three months ago, and the card wasn't updated. Stale evidence wearing a current face is a trap you set for yourself.

Because you think you're covered until the auditor actually reads the document and realizes it doesn't match the reality of the system today. And the third type. The third type is unfindable.

Someone made the document maybe years ago, but nobody can locate the authoritative copy. It's a dead pointer in your annex, a lost file on a defunct server. And to an auditor, an unfindable artifact is exactly the same as an artifact you don't have.

So human nature kicks in here. You build the annex, you see a row where the impact assessment is missing. The overwhelming temptation for a project manager is to just hit delete row.

So the spreadsheet looks perfectly seamless, perfectly green, and ready for management review. You must resist that temptation completely. A concealed gap is a landmine.

A disclosed gap is a governed task. I love that phrasing. When you find a gap, you do not delete the row to make things look pretty.

You create a defensible gap record. What goes into a defensible gap record? How do we document our failures safely? You record the gap type. You record the honest reason it exists.

Maybe the system was acquired in a merger and didn't come with paperwork. Or maybe the original owner left the company. You assign a named owner accountable for fixing it.

You log a concrete remediation plan with a hard date for completion. And crucially, you implement an interim control. An interim control.

Explain how that works in practice. An interim control is how you manage the risk until the paperwork catches up. It's like saying, we are currently missing the automated bias testing documentation for this loan approval system.

So until it is built and validated by Q3, we have instituted heightened human review on all of its outputs. You are showing the auditor how you manage the risk until the gap is closed. Honesty here is incredibly strategic.

If an auditor comes in and sees an annex where you are actively marking your own gaps, tracking them, assigning owners, and managing the risk with interim controls, they conclude that your organization governs itself. Exactly. You are in control of the reality, even if the reality isn't perfect.

But if you try to hide a gap to look seamless, and they catch you, and they will catch you, they instantly wonder what else are they hiding. The annex that shows its own holes is trusted. The one that pretends to be a seamless illusion and inevitably isn't loses all credibility.

Okay, so we have the theory, we have the six spines. Let's put this into practice. Let's walk through an immersive scenario to see how this actually plays out in the real world of an enterprise.

Let's talk about Evelyn. Let's do it. Evelyn is a newly appointed AI governance lead at a massive multinational logistics company.

Her board of directors has commissioned a Module 13 audit, a formal conformity assessment, and it's happening in exactly 90 days. High pressure. Very.

And the auditor they hired is known for being deeply adversarial, someone who will follow the paper trail all the way to the ground floor. Evelyn's company runs a few critical AI systems. They have a massive route optimization model that dictates driver schedules and fuel efficiency, which touches on labor union rules and safety regulations.

And they have an algorithmic hiring screen tool that processes thousands of resumes a day, which they rightfully treat as a standalone high-risk system. Evelyn knows her engineering and data science teams did the hard work over the last two years. She knows that somewhere in the company's vast digital footprint, there are data provenance files, rigorous evaluation reports, and detailed model cards.

But she also knows that somewhere is a massive liability. Exactly. So she stops her team from creating a giant shared audit folder.

She refuses to let them warehouse the documents. Instead, she begins building the evidence annex. She creates the register with the strict six fields.

She forces her engineers to find the authoritative links in Git and SharePoint. She ensures the version control provides the necessary provenance. And then she uses the two-way mapping.

She filters her new annex by the hiring tool on the system axis. And when she filters by the hiring tool, a robust 12-row trail appears. She sees the inventory entry, the data provenance logs, the evaluation report for bias, the joint impact assessments, the forensic logs, and the conformity file.

It's beautiful. It's a clean X-ray. But then she filters by the route optimization model.

And only three rows show up. Exactly. Before the auditor has even arrived, the annex has acted as a diagnostic tool.

Evelyn hasn't even intentionally looked for gaps yet, but the sheer thinness of the route optimization trail compared to the hiring tool tells her exactly where her program is dangerously under-evidenced. And because she groups by the obligation axis, mapping to the EU AI Act and ISO 42001, she spots the stale artifacts. She finds a model card for the route optimization model.

But the date field shows it's from last year, even though she knows for a fact the engineering team retuned that model in the spring to account for new electric vehicles in the fleet. It's a stale gap. She also finds that a customer service AI assistant had a bad incident a few months ago where it hallucinated policy to a customer.

But the post-incident review document is entirely missing from the annex. So what does Evelyn do? As we discussed, she doesn't hide them. She doesn't delete the rows for the missing incident review.

No. With 80 days left to go until the audit, she marks them as formal findings in the annex. She assigns owners.

She forces the engineering team to refresh the stale model card against the spring retune. She gets the post-incident review written and signed off by the product manager. She closes what gaps she can.

And for the gaps, she can't physically close in 90 days. Because sometimes engineering just doesn't have the spring capacity to rebuild a test suite in three months. She leaves them in the register as disclosed, governed deficiencies.

She attaches remediation dates for the following quarter and documents the interim controls they are using in the meantime. So 90 days later, the adversarial auditor walks in. They pick the hiring tool.

They ask for the date of provenance to ensure compliance with bias regulations. Evelyn doesn't scramble. Not at all.

She filters the annex, clicks the canonical link. And the file is on the screen in 45 seconds, complete with its version history and its cryptographic hash, proving it hasn't been touched since it was finalized. There's zero scramble.

The auditor then moves to the route optimization model and specifically points out a missing stress test document. Evelyn doesn't flinch. She points to the annex, shows it logged as a known finding, shows the interim human in the loop control, and shows the specific date it will be fixed.

And the auditor realizes they're dealing with an organization that rigorously governs itself. Evelyn wins the audit. Evelyn wins because she built the paper trail.

But you know, Evelyn can't just build this once and walk away, right? If she does this 90 days out, it's just a project. It's a sprint. Right.

How do we keep this alive so it's not just a frantic, one-time panic every time a regulator sends a letter? That is the crucial transition from project to program. An annex assembled the week before an audit is just a scramble with better handwriting. The actual value of the register is that it is current at any given moment.

Real audits, unexpected regulator letters, and sudden AI incidents that make the news do not schedule themselves for after you have conveniently tidied up your files. Yeah, life doesn't work that way. Maintenance must be live.

And live maintenance isn't about setting a calendar reminder for the first Friday of every month, is it? No. Calendar-based reviews are exactly how things rot. People ignore calendar invites.

Maintenance must be tied to events. We call these event-driven triggers. When a new AI system is onboarded into the enterprise, that is an event.

It enters the inventory and immediately opens its set of expected annex rows. Those rows are initially empty, which means its missing evidence is highly visible from day one. When a model is retrained by the data science team, that's an event.

It automatically triggers a card revision. When an incident occurs in production, that's an event. It mandates a new log entry to the annex.

The annex breathes with the lifecycle of the AI estate. And if you have a sophisticated engineering culture, you can automate much of this. If your models move through a standard CICD build pipeline, you can set up a webhook that automatically updates the version date and the file hash in the annex row the exact moment the pipeline fires and deploys the model.

That's amazing. But automation doesn't replace human ownership. Right.

Which brings us to the concept of inheritability. The annex must have a named owning role. Not a specific person.

Not just Evelyn or Priya, but director of data governance. If the institutional knowledge of how your AI is governed lives only in one hero's head, the paper trail walks out the door when they take a new job. You build the annex to be inherited.

Exactly. The audit that truly matters might happen two years from now under a governance lead who hasn't even been hired yet. And that future lead needs to be able to sit down on day one, open the annex, and instantly navigate the entire enterprise estate without having to interview 50 engineers to figure out where the bodies are buried.

Exactly. Add a light periodic backstop review, like a simple automated script that just pings every URL in the annex once a week to ensure the pointers still resolve and haven't become dead links and you have a living, breathing, defensible system. If we connect this to the bigger picture, the evidence annex is the vital junction between doing the hard work and proving the hard work.

So let's summarize the journey we've taken today. We started in a Beijing courtroom and established that the side with the paper trail wins. We learned that the annex is an index with provenance, not a warehouse of stale copies or a fictional essay.

We locked down the architecture. Six fields or no row. We learned that verifiable provenance, sometimes requiring cryptographic file hashes, is what turns a mere document into defensible evidence.

We mapped in two directions by system and by obligation so we can answer any question the auditor throws at us, whether it's about a specific model or a global regulatory duty. And we embraced the terrifying truth that gaps are findings. And finding them ourselves is a strategic gift.

It is the ultimate diagnostic tool for your AI governance program. It forces reality to the surface. So we always like to leave you with a concrete action.

The Monday morning move, the one thing you can do to start turning this theory into practice. It's simple, but it's incredibly revealing. Open your AI systems inventory on Monday morning, pick your highest risk system, just one, and try to fill out the six fields for its core evidence trail.

What it is, what it proves, who made it, when they made it, where it authoritatively lives, and what it maps to. If you can't fill all six fields for the core artifacts, or if the pointer is dead, or if you can't definitively prove provenance well, you know exactly what your team is working on this week. You find the gap before the auditor does.

Before we sign off from this deep dive into the evidence annex, and we wish you all the best of luck on building your paper trails, I want to leave you with one final provocative thought. It ties all the way back to the Lai-Vilou case in Beijing that we started with. Yes.

If we look back at that case, Mr. Lai's provenance didn't just prevent a regulatory penalty. It didn't just satisfy an auditor's checklist. Provenance proved ownership.

In the AI era, if you cannot definitively map the provenance of your system's data, its testing, its weights, and its outputs, you don't just risk failing a module 13 audit. You risk discovering that you don't actually own the intellectual capital you think you do. That is a chilling thought.

It is the ultimate stakes of evidence engineering. If you can't prove you made it, and exactly how you made it do, you really own it.

Real cases

These are documented cases used to illustrate the annex, not to predict your organization. Each is cited and used for a specific point.

Example 1: Li v. Liu and the trail that won the case (this topic's anchor). On 27 November 2023 the Beijing Internet Court ruled that an AI-generated image was copyrightable and belonged to the person who prompted it, because he could document his intellectual inputs: more than one hundred and fifty prompts, their arrangement, specific parameters, and iterative adjustments, plus a watermark carrying his identity (Kluwer Copyright Blog, 2023; China IP Law Update translation, 2024). The value for the annex is exact and load-bearing. The plaintiff prevailed because he had provenance he could produce; the defendant, who had removed the watermark and could not say where she obtained the image, had none. The case is the evidence annex in miniature: the party with the documented trail of who made what wins, and the party who scrambles for provenance loses. Your annex is the standing version of the plaintiff's trail for your whole AI estate.

Example 2: When the records, not the algorithm, decided the outcome. In the Dutch childcare-benefits scandal, the collapse came not only from a discriminatory risk system but from the state's inability to produce clean records of who was flagged, why, and on what basis, which turned a policy failure into a governance catastrophe (see Topic 10.1). Owned by Topic 10.1, it is referenced here for one point: the scandal shows that when evidence is scattered and unreconstructable, an organization cannot defend even the parts of its conduct that were defensible. The annex exists so that the records are never the reason you lose.

Example 3: Provenance had to be extracted because it was not offered. In the Rotterdam welfare-fraud case, investigators could only prove the scoring discriminated after obtaining the source code, model, and training data, evidence the organization had not organized for inspection (see Topic 10.2). Owned by Topic 10.2, the point for the annex is the inverse lesson: an organization that had wired its own evidence into a register could have produced that trail itself, on its own terms, rather than having it pried loose. The annex is the difference between disclosing your evidence and having it extracted from you.

Example 4: The EU AI Act puts the burden of demonstration on you. The AI Act's documentation and record-keeping obligations for high-risk systems, deferred by the 2026 Digital Omnibus to 2 December 2027 for standalone (Annex III) systems and 2 August 2028 for high-risk AI embedded in regulated products (Annex I), require the organization to be able to demonstrate compliance once they apply, not merely to have complied (Regulation (EU) 2024/1689; established). The example matters because it converts the annex from good practice into the mechanism of a legal duty: "demonstrate" means "produce the evidence on request," and an organization that did the work but cannot produce it has not demonstrated anything. Building the annex before the deadline binds, rather than after, is what turns the deferral into an advantage instead of a false sense of extra time. The annex is how the burden of demonstration is met.

Example 5: Provenance written into the content itself. China's 2025 labeling measures require AI-generated content to carry implicit provenance metadata (provider and content identifier) inside the file, so origin travels with the artifact (Cyberspace Administration of China, in force 1 September 2025, with GB 45438-2025; established). The example shows regulators independently reaching the annex's core instinct: provenance should be attached and verifiable, not asserted from memory. An organization operating in China must map that duty in its annex's obligation axis, and the deeper lesson generalizes everywhere.

Example 6: A management standard's evidence expectation. ISO/IEC 42001:2023, the first certifiable AI management system standard, expects an organization to maintain documented information and to make evidence of its controls available for audit (ISO/IEC 42001:2023; established as a voluntary, certifiable standard). For an organization pursuing certification, the evidence annex is close to the exact artifact an auditor of that standard will ask to walk, which means the same annex that serves the Module 13 audit also serves a real external certification. The point for the annex: build it once, and it discharges the demonstration duty across several regimes at once.

Example 7: The framework that assumes you can produce your trail. The NIST AI Risk Management Framework organizes governance around Govern, Map, Measure, and Manage, and every one of those functions produces artifacts a mature program is expected to be able to show (NIST AI RMF 1.0; established as voluntary guidance). The example illustrates the obligation axis: the annex can group its rows by RMF function to show, at a glance, that each function of the framework is backed by real evidence rather than by a claim of adherence. Frameworks assume the trail exists; the annex is where the assumption is made true.

Example 8: The generic scramble that precedes most failed audits. Across regulated industries long before AI, the recurring pattern in failed inspections is not that the work was undone but that the organization could not produce the records fast enough or prove they were genuine, so the inspector concluded the controls were not real (the pattern is visible in audit findings across financial, safety, and data-protection domains). The point for the annex is structural rather than tied to one named case: the scramble itself is the finding. An inspector who watches an organization hunt for its own evidence has already learned the most important thing about how that organization actually operates, regardless of what the eventually-produced document says.

Example 9: Records that could not be reconstructed after the fact. In several documented public-sector algorithm failures, the harm was compounded because the deciding organization could not later reconstruct which version of a system made which decision, on what data, under whose sign-off, so affected people could not even be told why they were flagged (a pattern the Dutch and Rotterdam cases both show, owned by Topics 10.1 and 10.2). The point for the annex is about timing: evidence captured as the work happens, with provenance attached, is producible; evidence reconstructed under legal pressure years later is weak, contested, and sometimes impossible to build honestly. The annex exists precisely so that the trail is captured while it is cheap and true, not excavated when it is expensive and doubtful.

Example 10: One annex serving several regimes at once. An organization certified to ISO/IEC 42001, subject to EU AI Act documentation duties, and voluntarily aligned to the NIST AI RMF does not need three separate evidence systems; the same well-built annex, with its obligation axis able to group rows by any of the three, discharges the demonstration expectation of all of them (ISO/IEC 42001:2023; Regulation (EU) 2024/1689; NIST AI RMF 1.0; all established). The point for the annex is leverage: the effort of wiring evidence once, with provenance and two-way mapping, pays off across every framework and regulator that asks the organization to show its work, which is why the annex is an investment rather than a cost.

Where people go wrong

  • Copying artifacts into one folder instead of pointing to them. The most common structural error. A folder of copies creates duplicates that drift from the originals the moment anyone updates a file, and within weeks the annex is a museum of stale versions that misleads the auditor. Keep every artifact where it authoritatively lives under version control and make the annex point to it. The annex is a map, not a warehouse.
  • Writing the annex as a fresh narrative essay. If your instinct on hearing "annex" is to author a new prose account of everything you did, stop; that is an essay, and an essay is a claim without a trail, which this program refuses. The annex is pointers and provenance, almost no original prose. Its value is in the wiring, not in new writing.
  • Leaving out the provenance and just listing documents. A list of document names is not evidence; it is a table of contents for claims. Without who made each artifact, when, and a way to show it is unaltered, every entry is a "trust us." Stamp verifiable provenance on every row, because provenance is what turns a document into evidence, and its absence is exactly what lost Li v. Liu for the defendant.
  • Building only one lookup. An annex indexed only by system fails "prove you meet this obligation," and one indexed only by obligation fails "show me everything about this system." The auditor asks both ways. Build both axes over the same rows, using the AI systems inventory for the system axis and the frameworks that bind you for the obligation axis.
  • Hiding the gaps to look seamless. The strongest temptation and the worst move. A gap you conceal is a landmine the auditor will step on; a gap you mark is a task that shows you govern yourself. Write every blank cell into the annex as a dated finding with an owner and a plan. The annex that shows its own holes is trusted; the one that pretends to be seamless and is caught is not.
  • Treating a stale artifact as current because it exists. A model card written before a retrain, an assessment done before a system changed, is stale evidence wearing a current face, which is a trap you set for yourself. The date and version fields exist so you and the auditor can judge freshness; tie refreshes to the changes that age artifacts, as the module's card and log logic requires (see Topic 10.3).
  • Letting a pointer dead-end. An artifact you cannot locate is, to an auditor, an artifact you do not have. A "where it lives" field that resolves to a moved file or a dead link is a scramble waiting to happen mid-audit. Every pointer must resolve to the authoritative copy, and checking that they do is part of keeping the annex live.
  • Assembling the annex the week before the audit. An annex built in a panic is a scramble with better handwriting, and real audits, regulator letters, and incidents do not wait for you to tidy up. Maintain the annex continuously, wired to the events that change evidence, so it is always ready rather than being made ready.
  • Giving the annex no owner, so it rots when a person leaves. An annex that lives in one person's diligence dies when that person moves on, taking the trail with them. Give the annex a named owning role and name each artifact's accountable maker, so the register is inheritable, which is the whole point of a trail that must outlast individuals (see Topic 12.5).
  • Mapping every system to your home jurisdiction only. A multinational annex that ties all systems to European duties and ignores that a system in China must carry content-provenance metadata, or that a United States hiring system sits under state bias-audit rules, is confident and wrong outside its home country. Map each system to the provenance duties of the jurisdiction where it actually runs.
  • Confusing the conformity file with the annex. The conformity file is a bundle of evidence for one high-risk system (see Topic 5.6); the annex is the register that points to every artifact across every system, including that conformity file as one entry. Do not mistake having assembled one system's file for having wired the whole estate. The annex is the layer above the files.
  • Believing that having done the work is the same as being able to prove it. Under the EU AI Act the burden of demonstration sits on you: "demonstrate compliance" means "produce the evidence on request." An organization that did everything right but cannot produce and authenticate the trail has not demonstrated compliance, and in an enforcement posture that is close to not having complied. The annex is how doing becomes demonstrable.
  • Giving every artifact the same weak provenance regardless of stakes. A model card for a low-risk internal tool may be fine under version control, but the log behind a decision that harmed a person, or evidence that will be examined in a forensic reconstruction, deserves the strongest integrity you can give it (a stored hash, a signed immutable log). Match provenance strength to the stakes, and note per entry how integrity is shown, so a weak-provenance row on a high-stakes artifact is visible as its own kind of gap.
  • Adding a row without confirming the pointer opens. It is easy to record a plausible-looking location and move on, but a pointer you have not actually followed is a pointer you cannot trust; the file may have moved, the repository may have been archived, the link may resolve to an old copy. Open every pointer as you add it, and re-check them in the backstop review, because a dead pointer discovered mid-audit is indistinguishable from a missing artifact.

Questions people ask

What is evidence annex?
A single indexed register that points to every governance artifact an organization has produced and, for each, records what it is, what it proves, who made it, when, where it authoritatively lives, and which system and obligation it maps to. It is a map to the evidence, not a warehouse of copies or a narrative essay, and it is the register the Module 13 audit reads from.
What is paper trail (versus scramble)?
The state of being able to follow a register straight to authenticated evidence on demand, as opposed to hunting through folders and hoping to find a document you cannot confirm is genuine. The paper trail is provenance you can follow; the scramble is provenance you are reconstructing while someone watches, and the scramble itself is a finding.
What is provenance?
The account of an artifact's origin and life: who made it, when, what version this is, what changed between versions, and a way to show it has not been altered since. Provenance is what turns a document into evidence, because it lets you answer "who made this, and how do you know it is genuine" without saying "trust us." More on Provenance
What is chain of custody?
The unbroken, accountable record of who created and handled an artifact and of its integrity over time, so that its authenticity can be shown rather than asserted. Version control provides most of it for free; outside version control it is approximated with dated signatures or stored file hashes. More on Chain of custody
What is file hash?
A short cryptographic fingerprint of a file that changes if a single character of the file changes, letting anyone check that a stored artifact is the same one that was recorded, without trusting anyone's word. One practical way to make an artifact's integrity verifiable.

Keep going