Skip to main content

Attacking to learn: you red-team a governance file and discover how thin most are

The short answer

Execute the file, do not debate it

The strongest attack takes a claim at its word, runs the claim's own logic on the claim's own kind of data, and shows what it actually does. Lum and Isaac did not argue that PredPol was biased; they ran it and it produced a feedback loop. Behavior beats rhetoric, and a file is strong on rhetoric and thin on behavior.

What you will be able to do

  • Decompose any governance file into its three structural parts (the claims it makes, the evidence it offers, and the assumptions it never states) so that the load-bearing joint becomes visible.
  • Distinguish a claim that is supported by diagnostic evidence from a claim that is merely asserted, restated, or backed by evidence that would look identical whether the claim were true or false.
  • Apply the re-run move (execute the file's own logic on its own data and observe the result) to a real governance claim, following the method Lum and Isaac used against PredPol in 2016.
  • Detect a feedback loop in a governed system: a place where the system's own output becomes a future input, so that the evidence for the system is manufactured by the system.
  • Expose a decorative control: a gate cited in the file as evidence that has never once blocked or changed an outcome, using the four questions that separate a control which governs from one which decorates.
  • Rank the weaknesses you find by how much weight they carry, and name the single attack that, if it lands, collapses the file's central promise.
  • Produce a red-team teardown of an external governance file: a short, evidence-anchored memo that states each claim, the attack against it, and the one weakness that matters most.
  • Rank a set of findings by weight rather than count, so your teardown leads with the one weakness that collapses the file's central promise and cuts the cosmetic ones.
  • Recognize what a governance file that survives the attack looks like, using the six marks of a strong file, so you know the standard your own file must meet.
  • Judge the ethics of the attack: steelman before you strike, attack the file and not the author, and handle a live system responsibly, including telling a genuinely diagnostic test from one a vendor built to pass.

The lesson

When organizations deploy artificial intelligence, they produce hundreds of pages of governance documentation to prove the system is fair, accurate, and safe. Yet, after these massive binders slam on a desk and earn approval, the exact systems they protect continue to fail catastrophically in the real world. The disconnect exists because physical size does not equal structural integrity.

Beneath the sheer volume of legal text and corporate reassurance, the vast majority of governance files are predictably and dangerously thin. An amateur reads these documents from front to back, reacting to the persuasive prose. A professional does not read the story at all.

They perform the attacker's read. We do not debate the vendor's brochure. We isolate the system's claims, apply its own logic, and execute it on paper.

By the end of this, you will know how to dismantle an external governance file. You will learn to locate the single load-bearing joint that carries all the risk and break it before the system ever goes live. Exposing compliance theater before a contract is signed is the only reliable way to stop algorithmic harm from operating at scale.

The first step of the attacker's read is stripping away the narrative. We force the persuasive document into a rigid three-column structure. We start with the claims column.

We hunt through the document for the load-bearing promises, stripping away all marketing adjectives and reassurances. What remains are the bare, testable propositions. Next is the evidence column.

For every claim, we locate the specific measurements, logs, or test results the file offers, and we align them directly next to the promises they support supposedly. Finally, we build the assumptions column. These are the ghostly, unstated premises about the real world that must be true for the evidence to actually support the claim.

This final column is where the fatal flaw of the system hides. Because these premises were assumed as fact, they are the only part of the architecture the authors never bothered to defend. Once the structure is built, we check for vulnerabilities.

Governance files fail at a predictable set of locations known as the six-week joints. You might find joint one, where a design intention is offered as proof of a result. Or it might break at joint four, resting on an unstated assumption, like presuming that historical hiring records reflect objective merit rather than historical bias.

You could spot joint five, scope laundering, where a vendor tests a highly controlled pilot but ships a sprawling, scaled system based on that safe data. Or joint six, a completely missing adversarial section, proving the red teaming was never actually performed. But the fastest tool to expose hollow data targets joint two, non-diagnostic evidence.

To find it, we apply the counterfactual test. Imagine a vendor selling a resume screening AI. Their file claims the system is unbiased and offers as evidence the fact that they removed the gender field from the input data.

We ask the counterfactual question, if this model were secretly discriminating based on proxy data, like a candidate's hobbies, zip code, or patterns of word choice, would the statement, we removed the gender field, look any different? The honest answer is no. The model can easily reconstruct gender from those proxies, and the vendor's evidence would read exactly the same. Evidence that reads identically whether a system is perfectly fair or deeply biased is non-diagnostic.

It carries zero mathematical weight. The deadliest vulnerability of all is joint three, the feedback loop. To see how it operates, we examine a 2016 investigation into PredPol, a patented algorithmic system heavily utilized by the Los Angeles Police Department.

The vendor claimed the model was highly accurate and completely blind to race, analyzing only location and time data adapted from academic earthquake prediction models. Without the source code, researchers refused to argue intentions. Instead, they ran the system's logic on paper, using public drug crime data from Oakland.

This animated map illustrates the logic they traced. Step one, the prediction output sends police officers to historically over-policed minority neighborhoods. Step two, officers patrol and make arrests in those specific locations, purely because that is where the model sent them to look.

Step three, those exact arrests become the fresh training data, cementing the neighborhood's risk score for the next day's prediction. Running this mechanism forward on paper revealed that Black residents were targeted at roughly twice the rate of White residents. This occurred within a model that explicitly claimed to be blind to race.

A feedback loop allows a system to grade its own homework. The vendor's massive volume of highly accurate data did not prove the system worked. It was the exact point of collapse, manufactured by the system itself.

This kind of hollow validation extends beyond algorithms. We also see it in human oversight, specifically in the human-in-the-loop defense. We classify this as a decorative control.

It is a review board or a human gatekeeper that exists entirely on paper, cited as evidence in the governance file, but which never actually alters an algorithmic decision. The test is simple. If a human control processes hundreds of decisions an hour with a near 100% pass rate, it is not governing risk.

It is rubber stamping the machine's output to launder the liability. A decorative, unexercised safety control is infinitely more dangerous than having no control at all. It provides the organization with the false security of oversight, allowing harm to scale without friction.

Returning to our three-column methodology, we have identified multiple vulnerabilities. Now we must prioritize them. A professional teardown operates on the principle of weight over count.

We ignore the missing version numbers and the typos. We isolate the single structural flaw that matters. We are hunting the load-bearing joint.

This is the singular assumption that, if proven false, collapses the central promise of the entire system. To execute this attack effectively, we apply strict tactical discipline. We steelman the vendor's claim by granting them their absolute best intentions and assuming they built the system honestly.

We attack the logic, not the author. If a system fails, even when assuming total good faith, the vendor cannot dismiss the critique by simply claiming they aren't lying. Granting a system its best-case scenario and breaking it anyway makes the structural attack mathematically undeniable.

A professional teardown never ends with a mere rejection. It concludes with a constructive operational solution. We apply the Harborview rule, demanding a cheap, random-sample diagnostic test that the system can actually fail.

Proposing a falsifiable test, like inspecting random city blocks to verify if a predictive algorithm is tracking real violations rather than chasing its own previous predictions, shifts the burden of proof back to the creators. We do not red-team governance files to act as destructive critics. We execute the attacker's read to internalize the base rate of thinness standard across the industry.

Breaking these files on paper today is the only way to ensure the systems we build survive reality tomorrow.

The ideas, one by one

Most governance files are thin, and predictably so

Across domains and jurisdictions, files fail at the same six joints: assertion in place of evidence, non-diagnostic evidence, circular evidence, unstated assumptions carried as facts, scope laundering, and a missing adversarial section. Knowing the joints tells you where to push before you have read the specific file.

Read the file as a structure, not a story

Take it apart into claims, evidence, and assumptions. The load-bearing joint, the one claim everything depends on resting on one weak assumption, becomes visible the moment the file stops being prose and becomes three columns. Break that joint, not the margins.

The counterfactual test is the fastest attack you have

For any claim and its evidence, ask what the evidence would look like if the claim were false. If the answer is "the same," the evidence is non-diagnostic and the claim is unsupported no matter how much data sits behind it. You can run this test on paper, in a meeting, in seconds.

A feedback loop turns a file's strength into its collapse

When a system produces the data used to validate it, the validating evidence agrees with the system because the system caused it. The mountain of confirming data becomes your finding rather than the file's defense. Trace one decision from output to input and see whether the loop closes.

Fairness makes the attack stronger, not softer

Steelman the claim before you strike, attack the file and not the author, and handle live systems responsibly. These are not concessions; they remove every avenue by which a defender could dismiss your finding as unfair, and they are what let your teardown survive the question "were you being fair?"

You attack others' files to harden your own

The purpose of this topic is not to become a critic. Every attack you catalog against someone else's file is an attack your own conformity file must now survive. The catalog you build here is the input to the rebuild in Topic 11.6, so the next red team that comes for your file finds less. (see Topic 11.6)

End the attack with the cheap test

A strong teardown does not stop at rejection; it names the specific, low-cost, diagnostic test that would settle whether the claim holds. This turns an attack into a path forward and demonstrates that you were attacking to learn, not to win.

Hunt decorative controls, not only missing ones

A control that has never once blocked anything deserves scrutiny. Ask of every gate the file cites: how many times has it been exercised, how many times did it change an outcome, what happened the one time it did, and if it never fired, is that because the risk never arose or because the threshold sits where nothing can reach it. A control that is present, cited as proof, and inert is thinner than an obvious hole, because the hole is visible and the decoration is not.

"Has a launch ever been delayed on a governance finding?" is the fastest read on whether governance is real

It is one question, it needs no data access, and the answer is legible: an organization that can name the launch it stopped has governance that changes outcomes, and one that cannot, while holding a thick file, has a file rather than a function. Ask it of the file you are attacking, then ask it of your own organization.

Weight beats count

One finding that collapses the central promise is worth more than twenty findings that do not, because a long list invites triage while a broken load-bearing joint forces a decision. Rank every weakness by whether the file's promise survives its failure, lead with the heaviest, and cut the cosmetic ones.

Most of the attack needs no system access

Four of the six weak joints (assertion, non-diagnostic evidence, scope laundering, and the missing adversarial section) are visible from the file alone. The counterfactual test runs on a printed page in a meeting. You can make a decisive finding before anyone gives you data or code.

You read it. Now prove it.

Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.

The conversation

The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.

Listen to it as episode 86 of the podcast.

Read the full conversation

Imagine you are sitting in a boardroom, right, and an organization has just spent millions of dollars, maybe several years, building this advanced automated system. Yeah, a massive capital investment. Exactly.

And the explicit mandate from day one, the thing everyone agreed on, is that this system must be entirely neutral. Right, they want zero bias. Zero bias.

So the engineers strip out every demographic variable, they ban the model from ever looking at race, gender, income, all of it, and the vendor hands you this dense, perfectly formatted governance file. I know the exact kind of file you mean. High gloss, very professional.

Oh, absolutely. And it claims it is mathematically impossible for the system to be biased. So the executives sign off, the system goes live, and immediately, like day one, it begins targeting minority populations at twice the rate of anyone else.

It sounds like this crazy, catastrophic edge case, right? But in the landscape of algorithmic governance, it's, well, it is actually the norm. Which is terrifying. It is.

What you're describing is this fundamental disconnect between a system's rhetorical defense, you know, the story it tells about itself, and its actual operational reality. And that's exactly what we are getting into on today's Deep Dive. We are looking at this purely from the perspective of the reviewer.

Right, the person sitting across that mahogany table. Yes, you are the busy professional. A vendor has just handed you that 50-page, high-gloss document that confidently assures you their new AI model is accurate, it's safe, and it's fair.

But your job isn't to just read that document and nod. No, your job is no longer to read that document. Your job is to break it.

I love that framing. Because today's focus is on adversarial governance. Specifically, how to red-team a governance file, right? Yeah, to discover its critical, load-bearing weaknesses before that piece of technology is ever deployed.

And to do this, we really have to recognize that the skills required to defend a system are entirely different from the skills required to attack one on paper. Completely different skill set, yeah. You aren't being a cynic just to be annoying.

You're learning the anatomy of a strong file by breaking a thin one on paper before it breaks in the real world and hurts people. Right. We're going to treat this as an executive-level teardown.

We are looking at the exact frameworks used to dismantle these confident vendor claims. And we have to start with the baseline principle of this entire discipline. Execute the file.

Do not debate it. Exactly. Execute the file.

Do not debate it. That principle, it's the dividing line between amateur reviewers and professional attackers. Because amateurs argue with the adjectives, right? They read the vendor's brochure and they argue about the wording.

Right. They get caught up in the philosophy of it. But professionals look at the underlying mechanics.

And to see what this looks like in practice, we have to look at this absolute watershed moment in algorithmic auditing. This is the 2016 paper, right? Yeah. Published by researchers Christian Loom and William Isaac.

It was in Significance, which is the magazine of the Royal Statistical Society. Great title too. To Predict and Serve.

It's an incredible title. And the target of that paper was a widely deployed predictive policing product called PredPol. Which, I mean, the origins of PredPol are wild.

It's deeply rooted in academia and municipal budgets. Yeah. It emerged from a collaboration between the Los Angeles Police Department and some researchers at UCLA, led by Professor Jeff Brantingham.

And they took the underlying mathematics from a completely different field, right? They completely repurposed it. They did. They took an algorithm originally designed to model HAWQS processes.

Right. Which seismologists used to predict earthquake aftershocks. And they applied it to crime.

Which, I mean, the theory makes a certain kind of intuitive sense if you are looking at very specific types of property crime. Like burglaries. Right.

In seismology, an earthquake relieves tectonic stress, but it transfers that stress to the surrounding area, which causes aftershocks. So the UCLA team theorized that crimes like burglary work similarly. Exactly.

A broken window on a street signals a vulnerability. That invites more burglaries in that immediate vicinity over the next few days. So the model learns those spatial and temporal clusters.

But the vendor didn't just sell it for burglaries, did they? No. They expanded its scope massively. Yeah.

And they wrapped it in this incredibly robust governance story. The core claim of the PredPol file was that the model was entirely objective. They literally said, math does not have a race.

Math does not have a race. That was the claim. Their primary defense was that the algorithm only ingests three data points, right? The type of crime, the location, and the time.

So the vendor stated explicitly that because the model never sees demographic data, it cannot be biased. It is mathematically impossible. But Lum and Isaac didn't write an op-ed debating the philosophy of blind algorithms.

They didn't argue with the brochure. No. They did something much more dangerous to the vendor.

They took the governance claim entirely at its word and they executed it. The rerun move. Yes.

The rerun move. They fed the PredPol model real drug crime data from Oakland, California. But the critical methodological choice they made, and this is so important, they introduced an external independent ground truth.

Right. To compare against the police department's internal data. Exactly.

They used the 2011 National Survey on Drug Use and Health. The NSDUH. Yeah.

It's a massive, rigorous public health survey. And the data from that survey revealed a reality that literally anyone working in public health already knew, which is that illicit drug use is ubiquitous. Yeah.

It's everywhere. It was spread almost perfectly evenly across the entire geography of Oakland, across all demographics. But when Lum and Isaac plotted the recorded drug arrests made by the Oakland police, the map looked entirely different.

Completely different. The arrests were hyper-concentrated in a small number of lower-income, predominantly minority neighborhoods. And this gets to the core issue.

This is the difference between victim-reported crime and police-discovered crime. Right. Because if your house is burglarized, you call the police, regardless of what neighborhood you live in.

Exactly. So, the police data for burglaries closely mirrors the actual occurrence of the crime. But drug possession is, well, it's a victimless crime in the reporting sense.

No one calls the police to report their own drug use. Obviously not. So, a drug crime is only recorded if a police officer happens to be standing there to witness it.

The arrest data doesn't map where people are using drugs. It maps where police officers are geographically deployed. Bingo.

It maps the historical patrol patterns. So when Lum and Isaac ran the earthquake algorithm on this specific dataset, the mathematical results were severe. The model looked at the historical arrest clusters and predicted that future drug crimes would happen in those exact same neighborhoods.

Which means it sent the police right back there. Yeah. In the simulation, the system directed police to target black residents at roughly twice the rate of white residents and other non-white residents at about 1.5 times the rate of white residents.

So the vendors claimed that the math didn't see race. It was technically true, right? But operationally meaningless. Exactly.

By relying on police-discovered crime data, the system was simply laundering historical deployment patterns through this complex mathematical earthquake algorithm. The model was learning the behavior of the police department, not the behavior of the community. It is the equivalent of a chef presenting you with this beautifully written 20-page menu detailing their complex culinary philosophy.

Right. All the adjectives. All the adjectives.

And instead of debating their choice of words, you simply ask them to sit down and actually eat their own food. Just eat the recipe. I love that analogy.

It shifts the fight from the ground of rhetoric, where the vendor's file is strong, to the ground of behavior, where the file is incredibly thin. And when Lum and Isaac did this, they exposed a terrifying reality for anyone sitting on a governance review board. Because the PredPol case wasn't just a one-off bad apple.

No. Not at all. It reveals a terrifying base-rate truth.

Most governance files, even from serious organizations with real lawyers, are fundamentally thin, and they fail in the exact same ways. Yes. When we say a file is thin, we are talking about a very specific phenomenon where a document reads as highly comprehensive, authoritative, complete.

It looks great. It looks amazing. Yeah.

But its structural integrity just collapses the moment it is subjected to mechanical testing. And this thinness is not random. It gathers reliably at six specific weak joints.

Let's go through those six joints, because this is where the actual attack happens. Right. So joint number one is the assertion in place of evidence.

This happens when a governance file states a design intention, as if the intention itself were a proven operational result. So you'll see phrases like, the system is built to avoid bias, or the model is designed to prioritize human safety. Exactly.

Now, in traditional deterministic software engineering, if you write code to execute a specific function, the intention usually matches the output. It does what you coded it to do. But in probabilistic systems, like machine learning, intention is entirely decoupled from the system's eventual behavior in the wild.

Right. An attacker reading that file recognizes that an intention is merely a starting condition. If the vendor cannot produce a specific empirical measurement, like a log, a behavioral test, an audit trail, that proves the intention survived the training process and actually exists in the deployed model.

Then the claim is treated as an empty assertion. The document is essentially offering the blueprint as proof that the building didn't collapse. Okay, joint number two is non-diagnostic evidence.

And we should define what diagnostic evidence actually is, right? Yeah, diagnostic evidence behaves like a functioning medical test. It yields one specific measurable result if the system is healthy, and a distinctly different measurable result if the system is failing. It has the capacity to fail.

Right. Non-diagnostic evidence, by contrast, looks absolutely identical, regardless of whether the system is functioning perfectly or causing massive harm. And the most common way we see this is the misuse of accuracy metrics.

Oh, all the time. A vendor will proudly display a 95% accuracy rate. But if you look at the methodology, that 95% was achieved by testing the model on a holdout set of the exact same data distribution it was trained on.

They're testing it in a sterile laboratory environment. Right. But the claim they are trying to support is that the model will be safe and accurate when deployed in the chaotic, shifting environment of the real world.

So that 95% lab score is real data, but it is non-diagnostic for the claim being made. Because the model could be failing catastrophically in deployment, and that historical lab metric would still read 95%. It doesn't change.

Exactly. It's like pointing to a thermometer that has been painted to always read 72 degrees and using it to claim the house isn't freezing. That brings us to the third joint, which is circular evidence.

We're going to dive really deep into this one in Section 5, but just broadly. What is it? Broadly, circular evidence occurs when a system is validated using outputs that the system itself generated. It's a feedback loop, and it's deadly because it looks like quantitative strength.

Okay, holding on to that one. The fourth joint is unstated assumptions carried as facts. Right.

Every governance file rests on a foundation of load-bearing premises that the authors rarely even realize they are making. Because data scientists and engineers are trained to treat data as ground truth, they often fail to recognize that data is merely a low-resolution shadow of human behavior. Perfect way to put it.

In the Pritiple case, the unstated assumption was reported drug arrests represent actual drug usage rates. And in algorithmic hiring tools, we see this all the time. The unstated assumption is often historical promotion rates are a pure reflection of employee merit.

Which completely ignores historical systemic biases against women or minorities in that company. The reviewer's job is to extract these invisible assumptions, put them under bright lights, and ask, do these actually map to physical reality? Because if the assumption fails, the entire mathematical structure built on top of it collapses. Absolutely.

Moving to the fifth joint, scope laundering. Scope laundering. I love this term.

What does it mean? It's a highly specific maneuver where the vendor swaps the parameters of the system that was tested for the parameters of the system that is actually being sold. You see a lot of this in medical AI, right? A vendor builds a diagnostic algorithm to detect, say, early onset sepsis. They test it in a heavily funded Tier 1 research hospital utilizing perfectly calibrated, state-of-the-art monitoring equipment.

With nurses who are explicitly trained to format the data perfectly for the AI? Right. And in that specific narrow scope, the tool achieves a 99% success rate. But then the vendor takes that 99% accuracy claim and uses it to sell the software to an underfunded rural clinic network.

Where the environment's completely different? Completely. You've got 15-year-old hardware, incompatible data standards, and overworked staff that cannot meticulously format inputs. The vendor is laundering the success metric from the highly controlled pilot scope into the chaotic deployment scope.

So the test for this is simple. You draw a hard mental boundary around the exact conditions of the test. Then you draw a boundary around the exact conditions of the proposed deployment.

And if those two shapes are not perfectly identical, the evidence is being laundered. Which brings us to the sixth and final joint, which might be the most obvious when you know to look for it. The missing adversarial section.

Right. If you have a 50-page governance document and it contains literally no section detailing how the system fails, which edge cases break it, or what stress tests the engineering team attempted that resulted in poor outcomes, that absence is your primary finding. Because complex automated systems do not operate flawlessly across all parameters.

It's impossible. Exactly. If the failure modes are absent from the file, it guarantees one of two realities.

Either the vendor lacked the competence to look for them, or they found them and made a conscious decision to hide them from the review board. Let me push back on this a bit, though, because as a reviewer, you get handed this glossy, beautifully formatted file. Isn't a highly polished, confident tone usually a sign that an organization actually has its act together? Like they spent the money on good tech writers and compliance people.

I get that instinct. But there is a psychological trap here for the reviewer. Human nature equates that polish with engineering competence.

But in adversarial governance, you must recognize that tone is free. Tone is free. Right.

Confident prose costs nothing. Diagnostic adversarial testing, on the other hand, is incredibly expensive. It's time consuming, and it often yields really uncomfortable results for the vendor.

Therefore, confidence and thinness often travel together. A hyper-confident tone is frequently a camouflage mechanism for a thin file. So if these files are built to look strong through confident prose, how do you, as a reviewer, physically read them so you don't get fooled? You have to stop reading them like a book.

Because if you read it front to back, you're just absorbing a story. Exactly. Corporate prose is explicitly designed to soothe the reader.

It maintains a smooth narrative flow, guiding your brain past logical gaps, using industry jargon and reassurances. Amateurs read front to back and react to the prose. Attackers read the file as a structure.

So how do we physically do that? You dismantle the file into three distinct columns. Just strip away the narrative entirely. Okay, so what's the first column? Pile one is the claims.

You comb through the document and extract only the load-bearing promises. You remove every adjective, every marketing buzzword, every statement of corporate philosophy. You reduce the text to bare, testable propositions.

Things like, the model does not utilize protected attributes. Or a human operator reviews every flagged decision before execution. Right.

Or the system processes queries with a 99% uptime. You are isolating the bare metal of the contract. Okay, what's pile two? Pile two is the evidence.

For every single bare claim you extracted in pile one, you search the document for the specific empirical evidence offered to support it. So you have to ruthlessly separate actual system logs from just future promises. Exactly.

Sort out the design descriptions from actual rigorous statistical measurements. And what you're really looking for here are the empty spaces. The moments where a massive load-bearing claim in pile one has absolutely nothing mapped to it in pile two.

Just a completely unsupported promise. Yeah. And then you build pile three, which is the assumptions.

This requires the most analytical rigor. Because the assumptions aren't usually written down, right? Exactly. For every paired claim and piece of evidence, you ask a singular question.

What specific conditions must be true about the real world for this piece of evidence to legitimately validate this claim? And you write these down as testable claims. So it's basically reverse engineering a magic trick, isn't it? Oh, that's exactly what it is. You stop looking at the magician's hands, the beautiful prose, the formatting, and you just look for the trapdoor.

And the trapdoor is usually in that assumptions pile because nobody bothered to defend it. Right. The vendor almost never defends the assumptions pile.

They usually don't even know it exists. So once you have the file sorted into this structure, these three piles, your objective becomes highly targeted. You aren't trying to find 20 minor typos.

You are hunting for the load-bearing joint. The single claim or the one weak assumption beneath it that carries the entire central promise. Because you don't need to break every claim in the file.

You only need to break this joint. If you break that one joint, the entire file collapses. So once we have the file sorted and we found our target, we need a weapon to attack the evidence pile.

Yes. And as a reviewer sitting in a boardroom, you don't always have the source code or the data to execute a live rerun like Lum and Isaac did, right? Right. You rarely have access to the API or the training data in that moment.

So what's the weapon? What is the paper equivalent of the rerun that we can use right there in the meeting? The counterfactual test. It is the fastest attack you have. It forces the reviewer to invert their perspective.

Instead of asking, does this evidence prove the claim is true? You ask, if this claim were entirely false, would this evidence look any different? Let's apply this to some real world examples because I think this is where the framework really clicks. Let's talk about the internal review defense. So say a vendor is selling a human resources platform.

They claim the algorithm's sorting mechanism is entirely neutral and fair. The evidence they provide in pile 2 is a statement reading, our internal compliance team conducted a rigorous fairness review and found zero issues. Okay, so we run the counterfactual test.

We ask, if this claim were false, if the algorithm was actually heavily biased, what would an internal review run by the vendor with vendor incentives have found? It would have found the exact same clean bill of health because the internal review is structurally designed to find compliance, especially when they are trying to close an enterprise contract. So the evidence is non-diagnostic. It looks identical whether the system is fair or biased.

Exactly. Consider another highly prevalent defense, the deleted proxy. Let's stay with a resume screening tool.

The vendor claims fairness because, and this is their evidence, we explicitly deleted the gender field from the training data. The model does not know the applicant's gender. Apply the counterfactual.

If the model were still actively discriminating against women, would the fact that they deleted the gender field look any different? No, it wouldn't. Because the model can still reconstruct gender from proxies, right? Exactly. High dimensional models are pattern matching engines.

They don't need explicit labels. They look at the ZIP code, the specific university attended, the extracurricular activities. Like playing women's field hockey.

Right. Or even the semantic vocabulary choices in the cover letter. The model synthesizes these proxies to reconstruct the applicant's gender with terrifying accuracy.

So the model discriminates based on the reconstructed proxy completely bypassing the deleted field. Yep. The vendor's evidence, we deleted the field, is factually true but operationally non-diagnostic.

It proves nothing about behavior. We see this same failure mode with endorsement pages, don't we? You flip to the back of the vendor's file and it's just a page full of impressive logos. ISA certifications, university partnerships.

Right. Do logos of certifiers prove safety? Apply the test. If the specific localized deployment of this algorithm was operating unsafely and causing harm today, would those logos still be printed in the brochure? Yes, of course they would.

Because a process certification only proves an organization has procedures, not that a specific model behaves well. But let me ask you this, doesn't this mean no evidence is ever good enough? Because a vendor could say, you can always play the what-if game, you're demanding an impossible standard of perfect evidence. And that is a fundamental misunderstanding of the framework.

A fair counterfactual doesn't demand perfect evidence, it demands diagnostic evidence. Evidence that moves at least somewhat when the truth moves. Exactly.

If a vendor claims a facial recognition system is highly accurate and their evidence is an audit performed by an independent adversarial third party utilizing a demographically balanced data set in the exact lighting conditions of the proposed deployment, that is diagnostic evidence. Because if the system were biased, that specific test would return a high error rate. Right.

The test has the capacity to fail. Evidence that cannot possibly fail isn't evidence, it is marketing. Now when you run this counterfactual test, especially on accuracy numbers, you will frequently uncover the most dangerous weakness of all.

Yes, joint number three, circular evidence. A feedback loop that turns a file's strength into its collapse. Why is this so deadly? Because it looks like overwhelming quantitative strength.

The system's evidence agrees with the system perfectly. But they agree because the system produced the evidence in the first place. So how do you trace a loop? How do you actually find it in the file? You have to trace the life of one single data point.

You just ask, does an output ever become a future input? We saw this in PredPol, right? The output is where to send the police. And the police go there, they make arrests because they're present, and the next day that new arrest data is ingested as fresh training data. Output becomes input with a one day delay.

It also happens with welfare or benefits risk scoring tools. All the time. Jurisdictions train models on decades of historical fraud data to predict future fraud.

But historical fraud data doesn't represent all fraud. It only reflects historical investigation patterns. Right.

Investigated groups generate more findings, which trains the model to score them as higher risk, which generates more investigations. Which validates the model perfectly. Let's walk through an immersive scenario to really solidify this.

Let's take the fictional city of Harborview. Okay, let's do it. So you are an executive in Harborview.

You are reviewing a 40-page glossy file for a system called Resident Risk. It's an objective code enforcement tool. And the vendor claims it has a 91% accuracy rate for finding building code violations.

Okay, so as the reviewer, I'm bypassing the UI mockups. I'm looking right at that 91% accuracy claim, and I apply the loop trace. So the model analyzes historical data and outputs a high risk score for the east side neighborhood, which happens to be older and lower income.

Right, so what do these inspectors do? They follow the model. They go to the east side. And building codes are dense.

If inspectors spend a week looking anywhere, they find violations. Peeling paint. Unpermitted water heaters.

So they log all those violations on the east side, and that becomes the training data. The 91% accuracy holds solid. Meanwhile, the west side neighborhood, the affluent area, gets a low risk score, so no inspectors go there.

So the algorithm sees zero newly recorded violations on the west side, and assumes it's in perfect compliance. The 91% accuracy is just the loop congratulating itself. It is the ultimate case of grading your own homework, but worse, you're also writing the answer key while taking the test.

The mountain evidence is actually the murder weapon. I love that, yes, exactly. But when you find a loop this devastating, the temptation is to write a memo calling the vendor dishonest.

You want to expose them. And that is the worst mistake you can make. Why? Shouldn't we point out that they're lying? Because in corporate governance, fairness makes the attack stronger, not softer.

Fairness is tactical armor, not politeness. If you launch an emotional attack or you are unfair, the defender will change the subject to your unfairness. And they'll win the room.

Exactly. The executives will side with the vendor because they already spent political capital securing the budget. You have to follow the three rules of a fair attack.

What's rule number one? Steelman before you strike. You must state the absolute strongest honest version of the claim. Grant PredPol, it's best case.

Assume it is a sincere demographic free model. Because it fails anyway under the best possible assumptions, the vendor has no escape. Precisely.

The intentions become irrelevant to the mathematical reality of the failure. Rule number two, attack the file, not the author. Don't assume bad faith.

Right. Some of the thinnest files are written by capable people who just assumed the false premise was true, like reported crime equals actual crime. Give them a finding they can act on, an objective mechanical finding, not an insult they will fight.

And rule number three applies to live systems. Yes, handle live systems responsibly. If the system is live and affecting real people, you must separate the governance finding from the operational attack detail.

So the governance finding is safe to share with the board. But the operational detail, the exact sequence that triggers the harm, must be handled with extreme care and routed straight to the operator to fix. Attacking to learn must never become harming the governed population to prove a point.

Let me ask you this though. What if I do all this? I steel man their claim, I run the counterfactual, I trace the loop, and I can't break it. Does that mean I failed the red team exercise? No, absolutely not.

That means the claim is sound. And you have to endorse it. You just tell the board it passed? Yes.

A red team that invents weaknesses where none exist destroys its own credibility. If the evidence is diagnostic and the assumptions map to reality, you have a good system. But assuming you did break it, you found the structural flaw, now it's Monday morning, and you have to synthesize these findings into a document that actually forces a decision from your executive.

The Monday morning move. The biggest mistake here is giving executives a list of 20 flaws. Because a long list invites triage, right? Exactly.

They'll patch the typos, ask for a few new charts, and move on. You have to rank your findings by weight, not count. How do you do the weight test? You ask, if this were the only thing wrong, would the central promise still stand? You discard the cosmetic footnotes, and you lead with the single load-bearing weakness that collapses the promise.

You also advise reviewers to hunt for decorative controls, right? These are gates or review boards that exist on paper, but have never changed in outcome. So what's the question you ask the vendor to spot these? You just ask, has a launch here ever actually been delayed or stopped on a governance finding? If they can't name one, the governance is decorative. It's just there for show.

That's brilliant. Now, how do you end the teardown memo? You say we shouldn't just offer a flat rejection. Never end with a flat rejection.

It's too easy for an executive board to overrule. You must end constructively by proposing the cheap diagnostic test. A specific, low-cost test that would settle the truth? Exactly.

So for resident risk, the cheap test is, inspect a random sample of low-scoring blocks. Go to the Westside. And for a credit model, it would be approving a random sample of denied applicants.

Right. If they run it and pass, great, you get a good tool. But if they insist on running a rigged test-like, they choose the sample.

Or if they refuse entirely, that refusal is your final damning finding. It's the ultimate superpower for anyone sitting on a review board. From finding the load-bearing joint to running the counterfactual to proposing the cheap test.

It changes the entire balance of power in the room. So as we wrap up, what is the single most valuable move our listener should make Monday morning? Pick up one governance file, compliance document, or vendor pitch currently sitting on your desk. Ignore the prose.

Do not read the adjectives. Extract its central claim and run the counterfactual test right there on the paper. Ask yourself, if this claim were completely false, would this evidence look any different? You are attacking these files to learn.

Because tomorrow you will have to write your own governance file and now you know exactly what it has to survive. You'll know how to build a file that isn't thin. Exactly.

Thanks for joining us on the Deep Dive.

Real cases

These examples show the attacker's read applied to real governed systems. The PredPol case is treated in depth because it is the anchor for this topic; the others show that the same weak joints recur across domains and jurisdictions. In each, the point is the method, not the verdict.

Example 1: PredPol and the re-run that broke it (Lum and Isaac, Significance, 2016). PredPol grew from a Los Angeles Police Department and University of California, Los Angeles project led by Professor Jeff Brantingham, running a patented algorithm adapted from the epidemic-type aftershock sequence model used to predict earthquake aftershocks. Its governance story rested on a single load-bearing claim: because the model uses only locations and times of reported crime, it cannot encode racial bias. Kristian Lum and William Isaac took that claim at its word and ran the model's logic on Oakland drug-crime data. They brought in one outside measurement the model never sees: the 2011 National Survey on Drug Use and Health, which indicated illicit drug use was spread fairly evenly across Oakland, while recorded drug arrests were heavily concentrated in a small number of poor and minority neighborhoods. Running the mechanism forward produced a feedback loop: predictions sent police to already over-policed neighborhoods, the resulting arrests raised the recorded crime there, and the next prediction sent them back. In their simulation, Black residents were targeted at roughly twice the rate of white residents, and other non-white residents at about one and a half times the rate, from a model that claimed never to see race. The attack broke every one of the file's defenses at once: it granted the honest best case (steelman), it ran the model's own logic on the model's own kind of data (re-run), and it named the loop where arrests, the model's output, became recorded crime, the model's input (follow the loop). PredPol was later renamed Geolitica in March 2021, shortly after the LAPD discontinued its use in April 2020; subsequent independent analysis of the tool's predictions (reported by The Markup and Gizmodo in 2021) continued to find the disparity the 2016 paper had predicted on paper years earlier. The paper is the model for this entire topic: it did not debate the file, it executed it.

Example 2: A hospital vendor's "critical hallucination rate." When a healthcare AI vendor represents a specific, reassuring error metric for a clinical tool, the attacker's first move is the scope-and-diagnosticity test. What population was the metric measured on, and is it the population where the tool is deployed? Would the number look different if the tool were in fact dangerous in the field? A metric measured on curated cases, reported by the party that profits from a low number, invites the counterfactual test from Section 3D: if the true field behavior were bad, would this particular measurement have shown it? The deep treatment of the eval-suite discipline that catches exactly this, and the enforcement action that followed one such misrepresentation, belongs to Topic 4.2; here it is an illustration of Joint 2 and Joint 5. (see Topic 4.2)

Example 3: Welfare and benefits risk-scoring files (multiple jurisdictions). Governance files for automated fraud-risk and benefits systems in several countries share a recurring load-bearing assumption: that the historical records used to train the risk score reflect actual fraud rather than actual investigation. If a group was investigated more in the past, it generated more recorded findings, and a model trained on those findings will score that group as higher risk, sending more investigation its way, which generates more findings. It is the PredPol loop in a different domain. The attacker does not need the file's internal numbers to make the finding; the loop is visible from the structure alone. The deep, sourced treatment of specific welfare-scoring cases and the fundamental-rights assessment they require is owned by Module 10. (see Topic 10.4) The transferable lesson here is that Joint 3 recurs wherever a system's enforcement generates the data that trains the next round of enforcement.

Example 4: A resume-screening tool's "we removed the biased feature" file. A common governance claim states that a hiring model was made fair by deleting protected attributes such as gender from its inputs. The attacker's read finds the unstated assumption (Joint 4): that the remaining features carry no information about the protected attribute. In practice, other features (schools, hobbies, gaps in employment, even patterns of word choice) can act as proxies, so the model reconstructs the protected attribute it was never given and discriminates through the proxy. The counterfactual test is quick: if the model were still biased through proxies, would "we removed the gender field" detect it? No. So the evidence is non-diagnostic for the claim. The deep case and enforcement history of hiring-AI bias is owned elsewhere in the program; here it is an example of how deleting an input is asserted as a fix and does not survive the assumption pile.

Example 5: A frontier vendor's model card as a governance file. Not every governance file is thin, and reading a strong one teaches as much as breaking a weak one. A serious model or system card names the evaluations that were run, the populations they covered, the failure modes the authors found, and the conditions under which the model should not be used. When you apply the attacker's read to a card like that, you find that the claims have diagnostic evidence attached, the scope of testing is drawn honestly around the deployed system, and there is an adversarial section describing how the authors made the model misbehave. That is what a file that survives looks like: it red-teamed itself before you arrived. Reading it, you learn the standard your own file in Topic 11.6 has to meet. (see Topic 11.6) The discipline of reading the card itself, rather than a press summary of it, is Topic 12.1's owned scope. (see Topic 12.1)

Example 6: The absent adversarial section as the whole finding. Across many real files, the strongest finding is not a broken claim but a missing one. A file that describes a deployed AI system and contains no section on how it fails, no red-team result, no "here is what we tried that broke it," has told you that the adversarial work was not done or not disclosed. You do not need to run anything to make this finding; the table of contents makes it for you. This is Joint 6, and in the base-rate sense it is the most common serious weakness in real governance files, because writing down how your own system fails is the part everyone is tempted to skip.

Example 7: The pilot that laundered its scope. A recurring pattern in vendor files is a glowing result from a controlled pilot, presented as evidence for a system that will run at scale in messier conditions. The pilot ran in one site, with attentive staff who knew they were being watched, on a curated population, for a short window. The deployed system runs across many sites, with ordinary staff, on the full population, indefinitely. The attacker draws one line around exactly what the pilot measured and a second around exactly what will ship, and the two lines are different shapes. Every number inside the pilot line is real; none of it supports a claim about the deployment, because the tested thing and the shipped thing are not the same thing. This is Joint 5, scope laundering, and it is easy to miss because the evidence is genuine, recent, and quantitative; only the substitution is silent.

Example 8: The benchmark that tested a different model. A vendor's file supports a safety claim about the shipped product with strong scores from a named public benchmark. The attacker checks one thing before reading the scores: was the benchmark run on the exact model that ships, in the exact configuration customers get? Often it was run on a larger, more expensive internal version, or an earlier checkpoint, or with settings no customer will use. The scores are real and the model that earned them is not the model on sale. This is Joint 5 again, and it recurs so often with AI benchmarks that "which model, which configuration, run by whom" is the reflexive first question, before any number is discussed. The deeper discipline of reproducing a vendor's benchmark yourself, and a case where a benchmark's funding was undisclosed, is owned by Module 12. (see Topic 12.1)

Example 9: The endorsement page as non-evidence. Many files include a page of endorsements: named experts, advisory board members, or a certification logo. The attacker treats the endorsement page with the counterfactual test like any other evidence. If the system were unsafe, would these endorsers have known, and would they have withdrawn their names? Usually the endorsers reviewed an early version, or lent their name to the organization rather than the specific system, or hold no ongoing responsibility for the system's behavior. An endorsement that would remain on the page whether or not the system is safe is non-diagnostic. This does not mean endorsers are dishonest; it means a name is not a measurement, and a governance decision cannot rest on one.

Example 10: The certification that certified the process, not the system. A file cites a management-system certification (of the kind that attests an organization runs a defined process for governing AI) as evidence that a specific model is safe or fair. The attacker separates the two claims. A process certification attests that the organization has procedures, documents decisions, and reviews risks; it does not measure whether any particular system behaves well, and it would remain valid whether or not the specific model discriminates. Using a process certification to support a system-behavior claim is a scope swap (Joint 5) dressed in an authoritative logo. The certification is real and valuable for what it attests; it is non-diagnostic for the behavior claim it is being borrowed to support. The deep treatment of what such management-system standards do and do not certify is owned by Module 6; here it is an example of authority standing in for evidence.

Example 11: The oversight that only looks like oversight. Many files claim a human is in the loop: "every automated decision is reviewed by a staff member before it takes effect." The attacker does not dispute that a human is present; the attacker asks what the human actually does. If the reviewer sees hundreds of decisions an hour, has no independent information beyond what the model shows, faces no consequence for agreeing and friction for disagreeing, and approves at a rate near 100 percent, then the oversight is nominal: the human is a rubber stamp that launders the model's decision into a "human" one. The counterfactual test applies directly: if the model were making harmful decisions, would this oversight process catch them? When the honest answer is no, the "human in the loop" claim is non-diagnostic, and the finding is that the file describes the presence of a human, not the exercise of oversight. The deep treatment of where a human must genuinely sign, and the enforcement action behind it, is owned elsewhere in the program; here it is an example of a claim that reads as strong and tests as hollow.

Where people go wrong

  • "A strong-looking file is a strong file." The opposite is often true. The files with the most confident language and the most impressive numbers are frequently the thinnest, because polish is cheap and diagnostic evidence is expensive. A glossy file with a mountain of confirming data may be resting entirely on a feedback loop (Joint 3), and the mountain is the finding, not the defense.
  • "To attack a file I need its internal data or code." Many of the strongest findings need nothing but the file's own structure. The missing adversarial section (Joint 6), the scope swap between demo and deployment (Joint 5), and the unstated assumption carrying the whole claim (Joint 4) are all visible from the outside. Lum and Isaac needed public arrest data and a public survey, not PredPol's source code, to break its central claim.
  • "Attacking the file means proving the authors lied." A strong attack works even if the authors were completely honest. Most thin files are written sincerely by capable people who made an invisible assumption. An attack that requires bad faith is weak, because the authors can simply and truthfully deny bad faith. Grant them their best case and break the file anyway.
  • "Accuracy is evidence of fairness." Accuracy measured against a system's own outputs is evidence of a loop, not of fairness or even of accuracy in the sense that matters. When the system chooses which cases get recorded, high agreement between predictions and records is guaranteed and meaningless. Always ask what generated the ground truth the accuracy is measured against.
  • "If I cannot run the system, I cannot use the re-run move." The re-run has a paper cousin that needs no execution: the counterfactual test. For any claim and its evidence, ask what the evidence would look like if the claim were false. If it would look identical, the evidence is non-diagnostic and the claim is unsupported. You can run this test on a printed file in a meeting.
  • "Finding many small weaknesses is a strong attack." A long list of minor flaws is easy to dismiss and easy to fix around. A strong attack finds the one load-bearing joint whose failure collapses the central promise, and spends its force there. Rank your findings by how much weight each carries, and lead with the one that matters.
  • "Removing a sensitive input makes a model fair." This is the proxy misconception. A model denied a protected attribute can reconstruct it from correlated features and discriminate through the proxy. "We deleted the gender field" is non-diagnostic for the claim "the model does not discriminate on gender," because the model can discriminate through everything else it still sees.
  • "An academic paper-only attack does not apply to a real deployed system." The PredPol case is the direct refutation. Lum and Isaac's 2016 paper-only re-analysis predicted the exact disparity that independent investigations of the live, deployed tool documented years later. A rigorous attack on the logic of a system is an attack on the system, because the deployed system runs that logic.
  • "Being fair to the file makes the attack softer." Steelmanning, attacking the file rather than the author, and handling live systems responsibly make the attack stronger, not softer, because they remove every avenue by which the attack could be dismissed as unfair. Fairness is not a concession. It is what makes the finding stick.
  • "The goal of the teardown is to reject the system." The goal is to learn where files are thin and to make systems, including your own, stronger. A good teardown often ends with the cheap, specific test that would settle the question (as Gretchen's did), not with a verdict. You are attacking to learn, and the best outcome is a better file, sometimes theirs and always yours.
  • "A human in the loop means real oversight." The presence of a human is not the exercise of oversight. A reviewer processing hundreds of decisions an hour, with no independent information and every incentive to agree, is a rubber stamp. Ask what the human actually does and whether the process would catch a harmful decision; if it would not, the oversight claim is non-diagnostic no matter how many humans are named.
  • "The controls the file lists are the part I can take at face value." They are the part most worth attacking. Hunt decorative controls, not only missing ones: a gate that exists on paper, is cited as evidence, and has never once blocked or changed anything is an assertion whose evidence turns out to be another assertion. Ask how many times it has been exercised, how many times it changed an outcome, what happened the one time it did, and if it never fired, whether the risk never arose or the threshold sits where nothing can reach it. A perfect pass rate over years is a finding, not a reassurance.
  • "A missing control is a worse finding than a weak one." Usually the reverse. A missing control is visible to everyone, including the file's authors, and it gets fixed. A decorative control is invisible precisely because it is present, and it buys the organization the appearance of governance at no cost in outcomes. The related organizational tell is the fastest test you have: ask whether a launch here has ever actually been delayed or stopped on a governance finding, and ask them to name it. An organization that can name one has real governance; one that cannot, with a thick file, has decoration.
  • "I should attack the file in the order it is written." Reading front to back keeps you reacting to prose, which is where the file is strongest. Sort it into claims, evidence, and assumptions first, find the load-bearing joint, and start there. The order of your attack is set by weight, not by the file's table of contents.
  • "A confident tone signals a strong file." Tone is free and evidence is expensive, so confidence and thinness often travel together. Bold, reassuring language is a reason to look harder at the evidence beneath it, not a reason to relax. The files that most need attacking are frequently the ones that sound most certain.
  • "The deployment population is different, so the training-data results do not apply, and that is a fair excuse." This is scope laundering wearing a disguise. A file's central claim is about the deployed system; if the vendor's own defense is that the deployment population differs from the population the evidence covers, the vendor has just conceded that the evidence does not test the claim. The correct response is not to accept the excuse but to name it as Joint 5 and require evidence measured on the population actually being deployed to.
  • "If the vendor refuses my cheap test, there is nothing more I can do." A refusal to run a cheap, fair, diagnostic test is itself a finding, not a dead end. Write the refusal into the teardown explicitly, and if the stakes justify it, escalate past the vendor relationship: request an independent third-party audit, invoke a right-to-audit or evaluation clause in the procurement contract, or route the finding to the relevant oversight body. A vendor who will not let their central claim be tested has told you something a passed test never could.

Questions people ask

What is governance file?
The assembled documentation that is supposed to show a system is fair, accurate, safe, and overseen: claims about the system plus the evidence and records offered to support them. In this program the conformity file assembled in Topic 5.6 is the canonical example. A file is strong when its claims survive being tested and thin when they collapse under an attacker's read.
What is red-team teardown?
The artifact this topic produces: a short, evidence-anchored memo that decomposes an external governance file into claims, evidence, and assumptions, names the single load-bearing weakness, states the one attack that breaks it, and ends with the cheap diagnostic test that would settle the question.
What is the attacker's read?
The method of reading a governance file as a structure rather than a story, by sorting it into three piles (claims, evidence, assumptions) and locating the load-bearing joint, instead of reacting to the file's prose.
What is load-bearing joint?
The single claim, or the single assumption beneath it, that everything else in the file depends on. Breaking the load-bearing joint collapses the file's central promise regardless of how sound the remaining claims are. The attacker's goal is to find and break this joint, not to list every minor flaw.
What is claim?
A load-bearing promise a file makes about a system, stated as a plain sentence that could be true or false, with the adjectives and reassurance stripped away. Claims can be tested; adjectives cannot. More on Claim

Keep going