Evidence Standards for Policy Claims
The short answer
Grade the claim, not the confidence of the sentence carrying it
A claim's evidence grade (measured, modeled projection, interested estimate, unsupported assertion) depends on how the underlying figure was produced, never on how declaratively it is stated.
What you will be able to do
- Grade any policy claim you encounter on a four-level evidence ladder (measured, modeled projection, interested estimate, unsupported assertion), based on how the claim was actually produced, not on how confidently it is stated.
- Trace a policy statistic from a headline or secondary retelling back to its primary source, distinguishing the original researcher's own words from every layer of paraphrase added on the way to you.
- Check a claim's date and its population before restating it: when the underlying measurement was taken, whether it has since been revised or superseded, and exactly who or what it describes, not the broader group a headline implies.
- Distinguish the verbs a primary source actually uses (measured, found, estimates, models, could, projects, may) from the stronger verbs a secondary retelling substitutes (will, is, causes), because the verb is where a claim's evidence grade quietly changes.
- Restate a graded claim at the exact strength its evidence supports, producing a sentence a skeptical reader could not fairly accuse of overclaiming.
- Assemble an evidence ledger: a working document that records, for every claim in a policy argument, its source, its grade, and its honestly restated form, ready to defend line by line.
- Apply the grading method evenly, to claims that support a preferred conclusion and to claims that complicate it alike, and set a named review trigger so a graded claim's currency is checked rather than assumed.
The lesson
Governance professionals process a chaotic flood of policy documents, white papers, and briefings daily. You must make split-second decisions on what data to trust. Often, the result of that incoming flood is a single, concentrated bullet point that looks like this.
It is bold, it is terrifying, and it cites a highly credible institution. But if we examine this graphic by tracing backwards through the layers of secondary aggregators and first-round media coverage, we see citation compression in action. The trail terminates at the primary source.
Reading the actual paper reveals a very different story. The economists didn't write that AI will replace jobs. They wrote it is likely to affect roughly 40% of jobs, with sharp variations by country income tier.
The original meaning shattered somewhere along the chain of citation. No one in that chain intentionally fabricated the headline. When the IMF's managing director discussed the paper, her framing was perfectly accurate.
But as the finding was passed from summary to summary, the vital nuance, the hedges, the population limits, was quietly stripped away. By the time a consequential AI policy claim reaches your inbox, it is rarely intact. To survive in this field, you have to systematically verify these degraded inputs before putting your own name behind them.
Policy professionals deploy these statistics in corporate boardrooms to set risk appetite, or in regulatory filings to argue for specific exemptions. They form the core evidence of formal legislative submissions. If you insert a single overstated, unchecked claim into a policy memo, you give every skeptic in the room a valid reason to dismiss everything else you wrote.
The defense against this is a strict verification discipline. You must grade a claim based entirely on how the underlying figure was produced, never on the confident tone of the sentence delivering it. A policy professional is ultimately selling their credibility.
Leaning on the authoritative name of an institution without examining its methodology is a professional hazard you cannot afford. To categorize any statistic before we trust it, we use the four-level evidence ladder. At the top is grade M for measured.
This is a direct count or empirical observation of something that has already happened. Next is grade P for modeled projection. These are figures produced by applying a transparent methodology to estimate a future event or an unobservable metric.
Grade I stands for interested estimate. This covers claims produced or funded by a party with a direct financial or reputational stake in the outcome. Finally, grade U represents an unsupported assertion.
This is a statistic completely detached from any traceable origin or disclosed methodology. We log and defend these evidence grades using a structured five-polym framework called the evidence ledger. This structure provides an objective standard for every figure in a memo.
It forces the analyst to justify a claim based on the source's actual production method rather than how well the number fits a desired narrative. Working through the ledger requires five mechanical steps. Step one is locating the actual document that produced the finding.
You must bypass the summaries and read the primary source yourself. Step two is always checking the date. Presenting a 2017 economic impact projection as a live current measurement in a 2026 briefing guarantees your analysis is stale before it is even read.
Step three is checking the population to confirm exactly who or what was measured, ensuring the sample matches the broad claim you are attempting to make. For example, the World Economic Forum's 2023 future of jobs report projected massive labor shifts based on a survey of 803 companies. Restating that specific employer survey as an ironclad prediction for the global workforce vastly overstates the actual data.
This chart illustrates the related danger of collapsing a model's range. When an analyst takes a McKinsey estimate of 2.6 to 4.4 trillion dollars and isolates only the top number, they discard the model's genuine uncertainty for the sake of a punchier headline. Concluding the trace without verifying the date and population guarantees that the analyst is actively over claiming on behalf of the original author.
Step four is checking the original author's verb because the verb explicitly carries the evidence grade. Measured indicates a count. Estimates signals a projection.
When you read the primary source, you are looking for hedge deletion. This happens when a cautious phrase like could affect is quietly upgraded to an absolute certainty like will affect during a retelling. That is an unauthorized grade change.
If the verb belongs to an interested party, a grade I claim, it is not automatically false. But it mandates an explicit check for a disclosed sample size and methodology before you can safely cite it. Finally, step five is restating the claim.
You must write the sentence at the exact strength the underlying evidence supports, combining the verified population, the date, and the correct verb. Contrast a hedged projection with a grade M measurement. When the Stanford AI Index tallies 59 AI-related federal regulations enacted in a year, you can restate that finding with a flat, direct declarative verb.
The evidence earned it. Your honest restatement will be longer and less dramatic than the viral headline, but it is the only version that survives a hostile cross-examination. Let's put this into practice.
On the left is the inaccurate draft about job destruction. On the right, our evidence ledger processes the primary source. It transforms that flat job loss statistic into a highly specific, defensible case for reskilling programs.
You must apply this exact same rigor to claims that support your preferred conclusions. Waving through convenient statistics while aggressively fact-checking opposing data is a confirmation bias trap. When you encounter conflicting data from credible sources, present both studies honestly alongside their methodological differences.
Never invent a blended average that neither author actually published. Similarly, a single executive's off-the-cuff prediction at a conference must be labeled as exactly that. You cannot launder one person's opinion into a peer-reviewed consensus.
Applying rigor only when it is convenient undermines the entire objective of the ledger. A policy voice that selectively edits its evidence will eventually lose its credibility under the weight of an adversarial check. A correctly graded claim does not stay pristine forever.
Evidence ages, and a bulletproof statistic today becomes a liability the moment a newer report supersedes it. This requires a currency discipline. Every logged claim in your ledger must have a named review trigger, a specific future event like the release of a new annual index, that forces you to re-verify the data.
At scale, mature organizations build a standing ledger. This allows analysts to trace a heavy, complex statistic exactly once and reuse it safely across dozens of documents, rather than relying on flawed memory for every new document. The ultimate test of this methodology happens here, in a hostile, high-stakes defense where every single word of your submission is aggressively scrutinized by experts.
Institutionalizing the evidence ledger makes your policy voice completely impervious to the aggressive scrutiny that routinely dismantles overstated claims. By stating every fact at its true strength, you secure long-term influence through undeniable accuracy.
The ideas, one by one
The primary source is the actual document that produced the finding, not the retelling closest to it
Follow every citation back to its origin; a secondary source citing a primary one is not itself the primary source.
Date and population are checks, not decoration
A checked date tells you whether a figure is current; a checked population tells you who or what was actually measured, almost always narrower than a claim's plain-language framing implies.
The verb carries the grade
"Measured" and "found" signal Grade M; "estimates," "models," "could," and "projects" signal Grade P; a stated interest signals Grade I regardless of verb; no traceable methodology signals Grade U. Watch specifically for a hedge quietly deleted between the primary source and the version in front of you.
An interested estimate is not automatically wrong, and it is not automatically trustworthy
Grade it Grade I, weigh it against whatever independent corroboration exists, and never let a stake-holder's own confident framing substitute for your own graded restatement.
Restate every claim at the strength its evidence supports
The correct restatement is usually longer and more qualified than the compressed version that first reached you; it is also the only version you can defend, unprompted, to a reader who independently checked the same source.
The evidence ledger makes grading auditable and reusable
Five columns, one row per claim: the original wording, the primary source, the checked date and population, the grade with justification, and the restated claim, ready to defend line by line.
An honest restatement can be a stronger argument, not just a safer one
As the immersive scenario shows, the correctly graded version of a claim sometimes makes a more precise, more persuasive case than the overstated original, because it survives scrutiny the overstated version invites and fails.
A claim's grade is not permanent
A Grade M count gets revised, a Grade P projection gets superseded, a study can face a later credible critique. Every ledger row needs a named review trigger, the same discipline this module has taught for a legal citation and a crosswalk row.
This skill is the prerequisite every other evidence discipline in this program assumes
The conformity file, the crosswalk, the evidence annex all assume the underlying facts feeding them are already solid; grading a claim before it enters any of those documents is what makes that assumption true.
Apply the method evenly, and build a standing ledger once claims start repeating across documents
Grading a favorable claim as rigorously as an unfavorable one is what keeps the discipline a credibility asset rather than a selectively applied performance of rigor, and reusing a graded claim's trace across documents concentrates fresh effort on genuinely new claims.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 51 of the podcast.
Read the full conversation
Artificial intelligence will destroy 40 percent of jobs worldwide. I mean, it really is just a terrifying sentence to hear out loud. It sounds like an absolute certainty.
Right. And the thing is, if you're listening to this and you see a stat like this in a newsletter or, you know, a policy deck, it not coming from some random forum post or a doom scrolling social media account. Right.
It has that sheen of authority. Exactly. This specific claim is sourced directly to the International Monetary Fund, the IMF, January of 2024, 40 percent gone.
Gone. And if you are listening to this deep dive right now as a busy professional, you are probably already calculating your own odds. You're you're looking around your office, thinking about your teams, your industry, maybe your own daily tasks.
Oh, absolutely. Everyone does it. Because the statement sounds so authoritative.
It sounds final. But and here's where it gets incredibly interesting and honestly a little alarming for anyone whose job relies on accurate information. When you actually take the time to trace that viral world shaking headline back to its exact origin, it is missing critical context.
And we aren't just talking about a little bit of flavor text being left out either. No, it is missing context that fundamentally changes the entire reality of the situation. Exactly.
The narrative we just laid out, it's a perfect illusion of certainty. And breaking that illusion is exactly what we are going to do today. So welcome back to the deep dive.
Yes, welcome. Think of today's discussion as a masterclass in professional executive education. If you are a sharp professional, maybe you sit on an A.I. policy team, maybe you are in government relations, a regulator's office or an advisory practice.
Your core function, your entire professional value revolves around your ability to make arguments using data. You're tasked with persuading people and often very skeptical people with a lot of money or power on the line. Right.
So today, our shared mission is to build a specific, unbreakable discipline for you. The discipline of grading policy evidence. I really love that framing because the goal here is simple, but the execution of it is going to completely change how you work, how you read and honestly, how you write forever.
It has to. We are going to make sure that you never, ever put your professional reputation behind a claim before knowing for absolute certain whether that claim survived its trip from a researcher's original data set to your inbox, completely intact. Because let's be honest about the environment you operate in.
The stakes for you are incredibly high. They are completely existential for your credibility. Yeah.
Picture this. You are in a boardroom. You have 30 minutes to make a case for a massive budget shift or a major policy pivot.
And your opening slide has a single powerful anchor statistic. Like the 40 percent stat. Exactly.
And a hostile skeptic in the room, and there is always a hostile skeptic, raises their hand and points out that your citable fact is fundamentally wrong. Which is a nightmare scenario. It's the worst.
Because getting that fact wrong doesn't just embarrass you in the moment. It hands every single skeptic in that room a legitimate, rational reason to dismiss everything else you say for the rest of the meeting. Right.
If your anchor statistic is proven false or even just heavily overstated, they will mentally throw out the parts of your memo that were entirely true and well supported. Your survival as a trusted policy voice depends entirely on your statistics. So having that first hostile question.
Because once the trust is broken, it just doesn't matter how brilliant your strategic recommendations are. You are the person who brought bad data into the room. You're tainted.
Yeah. So I want to be clear with you listening right now. This isn't about academic pedantry.
We are not here to obsess over footnotes just for the sake of being right. No, no. We are talking about wielding information as a reliable, highly defensible weapon in the boardroom.
This is Harvard Business Review meets a trusted mentor. I like that. Defensible weapons.
Right. And to get there, we are going to build a system around a core spine of principles. So let's introduce the first element of that spine right now, which is grade the claim, not the confidence of the sentence carrying it.
That is the foundation of everything. Grade the claim, not the confidence, because the confidence of a sentence is completely manufactured by the writer. The grade of the claim belongs to the evidence itself.
OK, so to understand why we have to separate the two, we have to look at how claims degrade in the wild. Right. Yes.
And we have a term for this process. We call it citation compression. Citation compression.
Right. The gradual, almost invisible loss of a claim's hedge, its date and its specific population across several rounds of individually accurate retelling. So let's look at the exact timeline of that IMF anchor case we mentioned at the pot, because I think it provides the perfect anatomy of how citation compression just ruins data.
Let's do it. Trace it back for me. Take me to day one.
What actually happened? OK, so day one is January 14th, 2024. A team of eight economists at the IMF published a formal staff discussion note. Eight economists.
Yeah. The authors were Kasanika, Germont, Lee, Molina, Panton, Pizzinelli, Rockall and Tavares. And the paper was titled Gen AI, Artificial Intelligence and the Future of Work.
OK, so this is our primary source. Exactly. This document is the primary source.
This is where the concrete was poured. They didn't just guess, right? They used a highly specific task overlap model. What does that mean practically? Well, they apply this model to occupational data across roughly 125 to 190 countries.
They were looking at the exact tasks that make up different jobs. And then evaluating how many of those tasks overlapped with the capabilities of generative AI. OK, so they are looking at the granular components of a job, not just a broad job title like accountant, but the specific things an accountant does all day.
Precisely. And so what did they actually find? Their actual published finding was that AI is, and I quote here, likely to affect about 40 percent of jobs worldwide. Wait, likely to affect? Yeah.
That is a massive difference from Will Destroy. It is a completely different universe of meaning, and the nuance goes much deeper than that, too. The paper noted that this exposure varied sharply by a country's income level.
Oh, really? How so? They modeled about 60 percent exposure in advanced economies, but that drops all the way down to 26 percent in low income countries. OK, so a huge spread. Huge.
And here is the most crucial piece of context that simply evaporated from the public consciousness. The authors estimated that roughly half of that exposed work was actually positioned for a productivity gain rather than displacement. Wait, I want to pause and make sure I am hearing this correctly.
40 percent of jobs are exposed to AI, but half of those exposed jobs might actually see their day to day work get easier, faster or more productive because of AI. Yes, exactly. That is not an apocalypse.
That is a massive economic opportunity completely baked into the exact same statistic. Entirely. If you are an investor or a policymaker, that productivity gain is the most important part of the paper.
Absolutely. But watch the mechanics of how this gets stripped away. The very next day, January 15, the IMS managing director, Kristalina Shorjeva, writes a blog post.
It's designed to coincide with the World Economic Forum in Davos. OK, Davos. So she's writing for a very specific crowd.
Right. Now, her framing is accurate to the underlying paper. She writes that almost 40 percent of jobs worldwide would be affected.
And she notes the skew toward higher income economies. Which is fair. It is.
This is standard public relations. You take a dense staff discussion note and you write a readable blog post for global leaders. And I imagine she is writing for an audience of extremely busy politicians and executives at Davos.
So, you know, she has to keep it punchy. But she is still accurate. Yes, she maintains the integrity of the claim.
But within hours, the global media machine takes over. CNBC, CNN, Euronews and dozens of others aggregate her blog post. And they are largely citing her blog post, presumably, rather than digging into the actual mathematical model published the day prior.
Precisely. They're citing the retailing closest to them. But by the time this story circulates through a second and third round of media aggregation over the next few days, the compression heavily sets in.
How does the language change? The phrase affected by AI, which, remember, could mean your job gets a great new software tool, quietly shifts. It becomes hit by AI, which sounds more aggressive. Right.
Hit implies damage. Then it becomes impacted by AI. And eventually, in the version that traveled the furthest across social media platforms and into corporate pitch decks, you arrive at the unsourced, terrifying claim that AI will simply replace or destroy 40 percent of jobs worldwide.
Wow. Full stop. No caveat.
No breakdown of advanced versus low income economies. No mention of the massive productivity gain. It is wild how quickly that happens.
But let me push back for a second, because, I mean, this just sounds like the classic game of telephone. Yeah. You know, you have a circle of kids.
One whispers purple monkey into the ear of the next. And by the end of the circle, the last kid shouts, turtle donkey. And people just mishear things.
They have bad memories and they accidentally fabricate new words. Isn't that all this is? Just a big game of corporate telephone. That is a very common assumption, but it is fundamentally different from the game of telephone.
And understanding that distinction is vital for your job. OK, tell me why. In telephone, the chain breaks because of an accident.
Someone mishears a phonetic sound and they invent a new word entirely. In citation compression, no one is accidentally mishearing anything and no one is fabricating a single new statistic. Wait, really? Really.
The original core words like 40 percent and IMF are technically maintained through every single step of the chain. Right. The number 40 never changes to 60 or 20.
It stays 40. Exactly. What happens is that the surrounding context is deliberately dropped to make the headline punchier, to fit a character limit or, frankly, to maximize click-through rates in an attention economy.
Ah, I see. The hedges, the population splits, the nuanced verbs, they're intentionally pruned away like branches on a tree until only the thickest, scariest trunk remains. The news aggregators didn't set out to lie.
They were incentivized to tighten the phrasing. And the next person down the chain tightened it a bit more. I see the mechanics of it now.
And that is why it's so uniquely dangerous for a professional. It happens quietly out in the open. Very quietly.
It makes it incredibly easy for a well-meaning, highly educated professional, someone listening to us right now, to unwittingly grab that perfectly confident headline and drop a completely unsupported assertion right into a multi-million dollar boardroom presentation. It happens every day. You think you are quoting an authoritative IMF dataset, but you are actually quoting the compressed exhaust fumes of three rounds of social media aggregation.
Exhaust fumes is the perfect way to describe it. Understanding how claims degrade in the attention economy reveals exactly why we cannot simply trust the confident tone of the final sentence that lands on our desk. Because the tone is fake.
The psychological confidence of a headline is almost always inversely proportional to the nuance of the actual evidence behind it. The louder and more absolute a statistic sounds, the more suspicious you should be of its journey. Which necessitates a formal grading system.
Exactly. If we cannot trust the sentence, we need a way to grade the actual underlying evidence before we ever let it enter our own work. Okay.
Which brings us to the methodology. How do we actually weigh this stuff? If we are going to grade the claim, not the confidence, we need a rubric, right? A way to evaluate how a figure was actually produced in reality. We use a four level evidence ladder.
It is a hierarchy of reality. Let's start at the very top rung with the most direct, unassailable form of evidence. We call this grade M. That stands for measured.
Measured, meaning something we can physically touch, count or observe in the real world. Exactly. A grade M figure is counted or directly observed.
It is a tally of something that has already happened in history. Do you have a real world example? Yeah. Let's look at the Stanford Institute for Human Centered Artificial Intelligence.
Specifically, their AI index report for 2025. In that report, they state that there were 59 AI related regulations introduced across 42 United States federal agencies in the year 2024. Okay.
And they contrast this with the previous year, noting it was up from 25 regulations across 21 agencies in 2023. Okay, so that is a literal tally. Someone at Stanford actually sat down, looked at the Federal Register and physically counted 59 distinct pieces of regulation.
It is a fact of history. It is not a guess about what Congress might do next year. Yes, it is an empirical observation.
It is grade M. Now, the next rung down the ladder is incredibly common in policy work, but it requires a very different kind of handling. This is grade P model projection. Okay, grade P. How does a projection differ from a measurement in terms of evidence weight? A grade P figure is produced by applying an explicit published methodology to estimate something that has not yet happened or something that is simply too vast to be directly counted.
Like the IMS study. Exactly. The IMF's 40 percent occupational exposure finding is the perfect example of grade P. Because they didn't literally interview every single worker in 190 countries to ask if their job was affected.
Right. That would be impossible. So instead, they built a rigorous, transparent model.
They took known data about tasks, built an algorithm to project exposure and arrived at an estimate. Which is highly valuable. It's rigorously tested evidence, but we must rigorously respect its boundaries.
It must always be framed as a model projection, not a measured historical fact. Okay, let me formulate a theory here and you tell me if I am off base. Go for it.
Let's say I cite a grade P claim in a memo today. A modeled projection about market growth for next year. And let's say 12 months pass and we get the actual economic data.
And it turns out the model was 100 percent flawlessly accurate. It nailed the exact number. Okay.
Does my grade P claim retroactively transform into a grade M claim? Because it turned out to be a fact. That is a brilliant question. But the answer is a hard no.
Really? Why not? Because the original claim, the paper you cited, remains a projection for the specific moment in time when it was made. You do not go back and change its grade in your ledger. But the number was right.
The number was right, but the method was still a projection. What happens is that the new observed reality, the actual economic data collected 12 months later, that confirms the model that becomes a brand new separate grade M claim. You can now cite that new grade M fact alongside the original grade P projection.
The projection was an educated guess about the future based on past conditions. The new data is a tally of the present. They are structurally distinct pieces of evidence, even if the numbers happen to match.
That makes perfect sense. The model is just the tool. The outcome is the measurement.
Okay. What is the third rung on our ladder? The third rung is where policy analysts get into a lot of trouble. This is grade one.
Interested estimate. Interested estimate. Yes.
This is a figure produced by a party with a direct financial, political, or reputational state in the conclusion of the data. And this brings us to another crucial element of our spine, which is an interested estimate is not automatically wrong and it is not automatically trustworthy. Okay.
I see this constantly. If a cybersecurity vendor releases a report saying ransomware attacks are up 300% so you need to buy our software, people have a really hard time knowing what to do with that. Of course they do.
Obviously the vendor wants you to believe the number so you buy the product, but it doesn't mean ransomware attacks aren't actually up. Precisely. And dismissing an interested estimate entirely out of hand just because a vendor wrote it is its own kind of dangerous bias.
You might be blinding yourself to a very real trend just because you don't like the messenger. So how do we handle it? A grade I claim just requires a significantly heavier burden of independent corroboration before you allow it into a policy memo. You cannot just accept the headline.
You have to aggressively verify the methodology. Right. Did they ask three people or 3000? Did that vendor survey five handpicked happy clients or did they fund a rigorous blind study conducted by an independent third party? The existence of interest dictates the intensity of your scrutiny.
Which naturally brings us to the absolute bottom of the ladder. Yes. Grade U, unsupported assertion.
Unsupported assertion. This is a figure or a claim with absolutely no traceable methodology. This is the exact grade that the IMF's 40 percent figure slipped into once it hit social media as AI will destroy 40 percent of jobs.
Because there's no math for that. No matter how hard you look, there is no mathematical model, no survey and no data set that proves that specific statement. It is just a number floating completely free of its evidence.
Okay. So we have the four grades. Measured is our historical tally.
Modeled projection is our rigorous guess. Interested estimate requires heavy scrutiny. And unsupported assertion is dead on arrival.
That's the ladder. Knowing this ladder is fantastic theory, but if I am a professional sitting at my desk and I have 50 pages of dense policy reports to get through by Thursday, knowing the theory feels useless if I don't have a systematic, repeatable way to assign those grades to the claims landing on my desk. And that is exactly why we do not rely on gut feeling or theory.
We use a rigorous five step framework. You have to trace the claim to its origin. You have to run the diagnostics.
Let's get into the mechanics of that framework then, because this is where the real work happens. Step one tackles the next core element of our spine, which is the primary source is the actual document that produced the finding, not the retelling closest to it. Explain step one.
Step one is the physical act of tracing. Find the primary source. The primary source is never the LinkedIn post you saw on your commute.
It is never the glossy slide deck your marketing colleague made. It is certainly never the news article. So what is it? The primary source is the actual original document, the raw data set, or the peer reviewed study where the numbers were born.
If you were reading a secondary source and it cites a primary one, you have an absolute professional obligation to click that link and follow the citation. You cannot simply trust the secondary source's interpretation. I think of it like conducting a financial audit or maybe a building inspection.
Oh, I like that. If I am a city building inspector, I am not going to sign a legal document certifying a skyscraper structural integrity just by reading the real estate developer's glossy promotional brochure. Definitely not.
I have to physically go down into the basement, turn on a flashlight and look at the actual concrete foundation. That analogy is flawless. You have to look at the concrete.
You have to see the rebar. Now, when you are walking down into that basement to find the foundation, you have to be highly vigilant against a very specific trap. We call it the widely reported fallacy.
The widely reported fallacy. What does that look like in practice? It looks like an illusion of consensus. It is when you see a specific claim, say, a massive job loss number repeated across 10 different major, highly reputable news outlets on the exact same day.
Right. Because it is everywhere, your brain naturally assumes it must be true. It has the weight of universal agreement.
But independent verification is very different from a circular wire trail. Because of how modern media works, right? The newswire economic. Exactly.
Modern journalism is driven by aggregation and speed. Often those 10 outlets are simply echoing each other, rewriting the same wire copy to capture search engine traffic. And every single one of them eventually points back to one single original study.
Ah, so it looks like 10 solid pillars of independent evidence holding up a roof, but it is really just one single pillar surrounded by 10 mirrors. Beautifully put. And if that one central pillar is cracked, the whole building comes down, no matter how many mirrors were reflecting it.
The claims grade rests entirely and exclusively on that one underlying source document. What happens when you hit a wall, though? Let's say I trace a citation. I find the original pillar, but the underlying academic study is locked behind a 200 dollar paywall and my organization doesn't have a subscription.
This is a daily reality for policy professionals. And how you handle it determines your integrity. Do not ever treat a paywall as a license to silently borrow the primary source's authority.
Meaning I cannot just write according to a groundbreaking study in the journal Nature if I only actually read the CNN summary of the Nature study. Exactly. Doing that is academic fraud.
You are claiming to have verified a methodology you haven't seen. So what's the fix? What you should do first is look for the free methodology note or the abstract. Academic journals almost always make the abstract and the broad methodology available for free, even if the full data set is locked.
Read that. If the text is completely, totally inaccessible, you must explicitly cite the secondary source as a secondary source in your own writing. So I would write something like, as reported by CNN, citing a study in Nature.
Yes. That phrasing is aggressively honest about the limit of your own verification. It tells your reader exactly where your visibility stopped.
Honesty over borrowed authority. I really like that rule. And what if the trail just goes completely cold? What if the CNN article cites a recent report, but doesn't name the report and no amount of searching brings it up? Then that dead end is your finding.
You cap the claim immediately at grade U. Unsported assertion. You absolutely cannot use the statistic as stated. Period.
Period. You either cut it entirely from your memo or you soften the language so heavily into qualitative general terms that it no longer poses a risk. You never, under any circumstances, push an untraced orphan number into a boardroom.
OK, so let's assume we do get our hands on the actual primary document. We found the concrete foundation. We can't just skim the title of the study and call it a day.
We have to interrogate the constraints of the data itself. Which brings us to steps two and three of the framework, covering the next element of our spine. Date and population are checks, not decoration.
Exactly. These are not just formatting requirements for a bibliography. They are structural load-bearing walls.
Let's look at step two. Check the date. You have to ask yourself, when was this figure measured or modeled versus when was the paper actually published? Why does the gap between the measurement date and the publication date matter so much? I mean, a year is just a year, right? Because in a hyper-accelerated field like artificial intelligence or really any modern technology sector, a time gap changes the entire reality of the environment.
Let's say you are writing a critical policy memo in late 2026. If you cite an AI capability benchmark that was measured in early 2023, but you fail to explicitly note that date gap to your reader, you are risking drastically understating what current live systems can actually achieve. So it's misleading by omission.
Right. The data you are citing might have been flawlessly accurate for the reality of 2023, but it is dangerously stale for your 2026 memo. Right.
It's like trying to navigate a rapidly growing city today using a street map printed 10 years ago. The map isn't technically lying about where things used to be, but entire neighborhoods have been built since then. If you trust the old map, you are going to drive your car into a wall.
Perfect analogy. So the date check prevents you from using an obsolete map. What about step three, check the population? What exactly are we checking here? You are checking precisely who or what was measured by the researchers.
In the plain language retellings we see in the media, the framing almost always artificially broadens the true scope of the original study. People want things to sound universal. Give me an example of that.
Let's look at another major real world example to illustrate this. The World Economic Forum, the WEF and their future of jobs report 2023. OK, this is a massive, highly cited report.
I've seen it everywhere. What did they claim? The headline, finding that rocketed around the world, stated there will be a net 14 million jobs lost globally by the year 2027. They broke that down further, predicting 69 million new jobs created, but 83 million existing jobs eliminated.
14 million net jobs lost in just a few years. That sounds massive and terrifying. It does.
And the compression error that inevitably happened in the retellings was that global media and consequently thousands of corporate policy slides claimed this figure represented the world's workers or the global workforce. But it doesn't represent the global workforce. No, it absolutely does not.
The actual population check, if you take 10 minutes to read the WEF's methodology section, reveals that this data is derived from a survey of exactly 803 companies. 803 companies. Now, those companies collectively represent about 673 million jobs, which is a massive sample.
But the data itself is specifically measuring employer expectations. It is a survey of human resources executives guessing about their future hiring plans. It is not an empirical global labor market census of all workers.
OK, I have to play devil's advocate here and challenge this because I can hear executives screaming in my head. 673 million jobs represented by those 803 companies is a truly colossal number for a busy executive who is reading a brief between flights. Isn't global workforce a perfectly fine, harmless shorthand? Does that level of semantic pedantry really matter in the real world? It matters deeply.
Yeah. And it is absolutely not semantic pedantry. It fundamentally changes the epistemological nature of the evidence.
How so? When you change the phrase surveyed employers expect to eliminate into a global workforce will experience an elimination. You're changing the core claim from a measurement of corporate sentiment into an absolute physical global truth. That is a structural lie.
A structural lie. Wow. An HR director's expectation of job cuts over the next four years is a completely different data type than an economist senses of actual displaced workers.
Executives are notoriously bad at predicting their own workforce needs four years out. If you use the shorthand, you are wildly over claiming what the evidence actually proves. And if someone on the board checks your work and understands labor economics, your credibility goes up in flames.
I see it now. Shorthand is where the truth goes to die. It feels efficient, but it's actually just stripping the structural integrity from your argument.
So the date and the population set the context of reality. But the grade of the evidence itself is ultimately determined by the exact specific words the original researchers chose to describe their findings. Which brings us to step four of the framework and our fourth spine element.
The verb carries the grade. The verb is the absolute most critical checkpoint in this entire process. The primary source's own language dictates the grade you assign.
You don't guess the grade based on how formal the PDF looks or how famous the institution is. You look at the verbs. Walk me through the signals.
What are the specific verbs we are hunting for? If you are looking to assign a grade M, a measured claim, the signals are retrospective, definitive words, verbs like measured, counted, found, recorded, observed. These verbs grammatically indicate something that has already happened and was physically captured by the researchers. For a grade P, a model projection, the signals are forward looking or probabilistic.
Verbs like estimates, models, projects, forecasts, could is likely to. This makes me think of the dashboard on my car. Measured or recorded is the speedometer.
It is telling you exactly how fast you are physically traveling in this exact second. It is a historical, undeniable fact of your momentum. Models or projects is the GPS screen estimating your arrival time.
You would never, ever treat your GPS ETA as a historical fact. It's a highly educated, algorithmically rigorous guess based on current traffic conditions. But it is still just a guess.
If you hit a pothole, the ETA changes. That is a phenomenal analogy. Treating a GPS ETA as a historical fact is exactly what happens when people fall victim to a very specific error we call hedge deletion.
This is a massive, high value failure mode in corporate communications. Break down hedge deletion for me. How does it happen? It happens when a careful, accurate sentence like the IMF estimates AI could affect 40 percent of jobs goes through a round of editing by someone trying to make the text punchier.
And it becomes the IMF says AI will affect 40 percent of jobs. I see. They drop the words estimates and could and they replace them with the words says and will.
Exactly. And in doing so, deleting those two tiny words upgrades a grade P model projection directly into a grade U unsupported assertion. Just from two words.
The underlying evidence, the actual math the IMF did, never supported the absolute certainty of the word will. The hedge, the word estimates or could, is not just plate academic decoration, is the exact structural component carrying the claims true evidence grade. When you delete the hedge, you delete the truth.
And there is another incredibly sneaky shift that happens with verbs, too, especially in policy arguments. Right. Moving from correlation to causation.
Yes. The correlational versus causal shift. This ruins more arguments than almost anything else.
You have to be hyper vigilant when a study says two things are associated with or linked to each other. That is a correlation. It means they happen at the same time.
But a corporate retelling will almost always try to upgrade that gentle verb to an active causal one like causes or drives. Like taking a study that says AI adoption is associated with a decline in entry level postings and putting a slide up that says AI adoption causes a decline in entry level postings. Exactly.
Associated with simply means we observed both trends happening simultaneously in the data set. Causes implies a direct proven mechanical relationship that the study, if it was purely observational, never actually established. Maybe a broader economic downturn caused both the AI adoption and the hiring freeze.
Which is a huge distinction. This is a massive legal and reputational risk in policy memos. Analysts make this mistake because they desperately want a silver bullet argument for their policy.
But if you stand before a regulatory committee and claim a new rule will cause a market shift based purely on correlational verbs, a sharp opposing lawyer will dismantle your entire argument in 60 seconds. OK, so we have traced the source. We have checked the date and the population and we have guarded the verb to identify the true grade.
Now we have to actually do the hard part. We have to write the final sentence that goes into our presentation or memo. We have to put our name on it.
This is the payoff. Step five of the framework, which corresponds to our final spine element, restate every claim at the strength its evidence supports. Step five is where the rubber meets the road.
It is where you earn your paycheck. You must write a completely new sentence carrying the correct verb, the correct date and the correct population explicitly stated for the reader. To really understand how this works under the pressure of a real corporate environment, let's walk through a detailed immersive scenario.
Let's introduce a fictional policy analyst. We'll call her Geraldine. I love this.
Let's set the stage for Geraldine. She is a senior policy analyst at a fictional software vendor called Thornwick AI. Thornwick AI.
OK. And Thornwick sells enterprise workplace co-pilot tools software that helps employees write emails faster, summarize meetings, that sort of thing. It is a Tuesday afternoon and Geraldine gets a high priority assignment from her boss, the head of government relations.
She needs to urgently review a draft legislative submission that their company is going to submit to a Senate committee studying the labor effects of artificial intelligence. This is a highly visible, high stakes document. This is going to lawmakers and their staffs.
Extremely high stakes. It will be public record. Now, the draft was originally written by an expensive outside lobbying consultant, and it opens with this incredibly aggressive sentence.
It reads, the IMF has found that AI will destroy 40 percent of jobs worldwide, underscoring the urgent need for a massive reskilling tax credit. Wow. Her boss emails her and says, Geraldine, just verify this stat quickly so we can send it out.
I mean, put yourself in Geraldine's shoes. It's Tuesday afternoon. She has a mountain of other work.
She has probably seen that 40 percent stat a million times on Twitter and in tech newsletters. It would be so incredibly easy and so human to just quickly add a footnote saying IMF 2024, tell the boss looks good and move on to her next meeting. It would be the easiest thing in the world and it would be a critical career damaging failure.
But Geraldine is a disciplined professional. She knows the rules. So she traces it.
She applies step one. She applies step one. She bypasses the consultant's text and physically locates the primary source.
She opens the actual January 2024 IMF staff discussion note. She sits down, reads the methodology and the executive summary, and immediately three massive red flags jump out at her. Let me guess, based on our framework, the verbs, the population and the missing context.
You nailed it. First, she checks the verb. The actual IMF paper states that AI is likely to affect roughly 40 percent of jobs.
It never anywhere in the text claims it will destroy them. Blatant hedge deletion happened. OK, flag one.
Second, she checks the population. The 40 percent is a global average that ranges wildly from 60 percent in advanced economies to 26 percent in low income countries. The consultant's draft flattened all of that demographic nuance into a single terrifying monolith.
And the third flag, the context we talked about earlier. Right. The productivity gain.
Roughly half of that exposed work is explicitly framed by the IMF economists as being positioned for a productivity gain, not job displacement. The outside consultant dropped that completely, presumably because they thought a purely terrifying statistic would make a better case for the tax credit. But it is directly, fundamentally relevant to a legislative committee trying to weigh the nuances of a reskilling tax credit.
OK, so Geraldine has done the hard work of tracing. She knows the consultant's draft is structural garbage. How does she execute step five? How does she write the honest restatement? She builds it out in her evidence lecture.
She drafts a sentence that protects her company's integrity. Her restatement looks exactly like this. IMF staff economists modeled in a January 2024 analysis that AI is likely to affect roughly 40 percent of jobs worldwide, with the share ranging from about 26 percent in low income countries to about 60 percent in advanced economies.
An estimated that roughly half of that exposed work is positioned to benefit from AI driven productivity gains rather than displacement. OK, let's inject some boardroom reality here. If I am Geraldine's boss, the hard charging head of government relations, and I am looking at that massive multi-clause sentence she just sent back to me, I am furious.
I'm saying, Geraldine, this sentence is way too long. It's way too academic. It's way too soft.
And it completely guts the punchy argument I was trying to make for the tax credit. The consultant's version was scary and scary gets funding. And this is the moment of truth for the policy professional.
But Geraldine has the perfect bulletproof defense ready because she didn't just play grammar police with the words. She traced the underlying reality of the evidence. She lifts her boss in the eye and tells them this restated version is actually a significantly better, more persuasive argument for our tax credit.
How does she justify that? Because it doesn't sound nearly as scary. Because it points out the exact specific mechanism of the policy need. She explains to her boss that the nuanced argument is better.
The tax credit isn't just needed to cushion the blow of inevitable apocalyptic job destruction. The credit is actually needed to help workers actively upskill and move into the productivity gain half of the exposed work. It transforms the company's lobbying position from a defensive panic into a proactive, economically beneficial strategy.
Furthermore, and most importantly, her sentence survives hostile legislative scrutiny. Because if a sharp Senate staffer actually pulls the IMF paper. Exactly.
If a sharp committee staffer pulls the IMF paper to check Thornwick A.I.'s claims, Geraldine's sentence matches the concrete evidence perfectly. Thornwick looks like a deeply informed, trustworthy partner to the government. If they had submitted the consultant's original sentence, their credibility would be instantly destroyed.
The Senate staffer would look at the discrepancy and conclude that Thornwick A.I. either didn't bother to read their own cited source or worse, that they actively chose to lie to a legislative body. That is a phenomenal pivot. She turned a massive reputational liability into a highly specific, defensible policy lever.
She won the argument by being more accurate. Yes, she did. Okay, but let's be honest.
Geraldine's success in that scenario came from tracing an IMF paper, which is a highly transparent grade P model produced by public economists. What happens when the source you are trying to price is inherently biased by its very nature, like a vendor's own marketing claims. This brings us to a deep dive into edge cases.
Let's unpack grade one, Interested Estimates. Interested Estimates are incredibly common in corporate environments, and the temptation for analysts runs in two equally lazy directions. You either accept them completely uncritically because the numbers sound incredibly precise or you arrogantly dismiss them entirely because you assume any vendor trying to sell you something must be lying.
Both extremes represent a failure of analysis. Let's use a hypothetical to ground this. Say I am evaluating a new software vendor and their marketing landing page boldly claims customers using our platform report cutting their compliance review time by 40 percent.
It sounds great. How do I apply the five steps to a marketing claim? You aggressively interrogate the methodology behind the marketing. Step one, is there a primary source that exists beyond the glossy splash page? Is there a named downloadable case study? Is there a published methodology note explaining exactly how they defined and measured compliance review time? What if I dig around the site and I find a tiny little footnote at the bottom that says, based on a survey of customers? Then you immediately apply steps two and three.
You check the date and the population. What is the date of that survey? Is it data from an old beta test in 2023, but it's still living on their 2026 landing page? What is the sample size? If you do the digging and it turns out the data is based on 400 self-selected customers who responded to an email survey in Q2 of 2026, then you absolutely can use the figure. But you must cite it with its fully disclosed population.
So my honest restatement in my internal memo to the buying committee would be a vendor survey of 400 self-selected customers in Q2 2026 reports a 40 percent reduction in compliance review time. Exactly. You wrap the vendor's claim tightly in the context of its creation.
You let the buying committee see exactly how the sausage was made. But, and this is crucial, if there is absolutely no methodology disclosed, no sample size, no date, no explanation of how the data was gathered, then it is grade U. It is an unsupported assertion. You cannot use it to justify a purchase, no matter how specific that 40 percent number sounds.
There is no concrete behind the number for you to check. That makes a lot of sense. What about a different edge case? What about competing studies? Let's say I trace a claim and I actually find two highly credible grade P primary sources, but they completely disagree on the numbers.
One says the market will grow by 10 percent. The other says it will shrink by 5 percent. Do I just average them together, create a blended figure right in the middle so I seem reasonable and moderate? Never, ever average competing studies.
If you average them, you are mathematically inventing a blended figure that neither team of researchers actually published and that neither underlying methodology supports. You are creating a fiction. And on the flip side, you also shouldn't silently pick the more dramatic study just because it helps the narrative of your memo.
So what is the honest move when the experts disagree? You state both. You present the range honestly to your reader and you briefly note the most likely reason for the divergence. Usually, if two rigorous studies disagree, it is because they use different populations, different modeling assumptions or different time frames.
Explain why they differ makes you look incredibly well informed and objective, whereas hiding one of the studies makes you look selective, biased or just plain ignorant of the landscape. That builds immense trust with the reader because you're treating them like an adult who can handle complexity. OK, here is a very modern edge case that everyone is dealing with right now.
Yeah. What about using A.I. assistance for the research itself? Can I just prompt a large language model and ask it to trace the primary source for me to save time? You can absolutely use an LLM to speed up your initial search or to help you brainstorm search terms, but you must be hypersensitive to the mechanics of A.I. hallucinations. An A.I. might generate a beautiful, perfectly formatted citation for you, complete with a highly realistic sounding author name, a plausible academic title and a specific page number.
But that entire document might be completely fake. It might not exist anywhere in reality. But why does it do that? Why does the A.I. lie? It is not lying in the human sense of deception.
It is crucial to understand the mechanics here. Large language models are fundamentally predictive text engines. They are mathematically predicting the next most likely word in a sequence based on their training data.
When you ask it for a citation, it isn't always searching a database of real documents. It is often just predicting what a citation ought to look like in that context. So it stitches together real authors with plausible sounding titles to create a synthetic hallucination.
Which is different from the compression we talked about earlier. Exactly. This is the distinct, different failure mode from citation compression.
Compression is the gradual loss of nuance from a real source. Hallucinations fabricate a source entirely out of thin air. So the rule of thumb here is trust the A.I.'s output, but verify it by clicking the actual link.
Not even trust, just verify. You must personally open the actual document, confirm with your own eyes that the document exists, and read the text to ensure it actually says what the LLM claims it says. An A.I.'s output must be treated as an untraced, unverified claim until you manually put your hands on the concrete.
Got it. Now, what about long range economic projections? Because those seem to haunt policy decks for years. People quote them long after they've gone stale.
They do. And they degrade terribly over time because people forget the date check. Let's look at a classic example.
PWCS famous June 2017 sizing the prize report. It projected that A.I. could contribute up to 15.7 trillion dollars to the global GDP by the year 2030. I have seen that 15.7 trillion number quoted in pitch decks literally this week.
Yes, it is everywhere. But citing a nine year old cumulative projection in the 2026 memo as if it were a near term, current finding is a massive analytical defect. It was always a multi-year cumulative forecast based on the economic realities of 2017.
It was never an annual reality. Another classic example of projection abuse is McKinsey's June 2023 report on generative A.I. They modeled a potential annual value ranging from 2.6 trillion to 4.4 trillion dollars. That is a massive range, almost a two trillion dollar spread between the low end and the high end.
It is a massive spread, which indicates a high degree of uncertainty in the model. But because a wide range is clunky for a headline or a quick slide, secondary coverage routinely collapses that entire range straight to its absolute upper bound. They just say 4.4 trillion, which is misleading, highly misleading.
Doing that intentionally discards the genuine mathematically calculated uncertainty that McKinsey's own economists expressed, collapsing a range to its highest possible figure, falsely overstates the findings precision and its confidence. What if the claim comes from inside my own company? Let's say I'm writing a memo and I'm looking at internal adoption data provided by our own engineering team. Applying all this strict tracing to my own colleagues feels a bit paranoid or counterproductive.
You don't ignore internal data, but internal claims are absolutely not automatically grade M just because they come from inside the building. You have to trace them identically. Really? Even internal data? Yes.
Who exactly calculated that adoption metric? From what underlying database? As of what specific date, did they measure daily active users or just total registered accounts? If you only have a secondhand summary of an internal headline result, you have a professional obligation to track down the specific analysts who ran the query and ask them how they built the query. Internal evidence burns its grade exactly the same way external evidence does. OK, I am entirely sold on the rigorous philosophy here, but let's be practical.
Applying this intensive level of scrutiny, tracing every single citation to the basement, checking the dates, checking the populations and guarding the verbs, doing that for every single claim in a 50 page policy memo seems practically impossible if you're starting from scratch every single time. There are not enough hours in the week. You are exactly right.
If you start from scratch every time, you will burn out in the month, which is exactly why you do not start from scratch. You need to build infrastructure. This brings us to the operational core of this entire discipline, building the evidence ledger and maintaining the currency discipline.
The evidence ledger. This sounds like the holy grail of keeping an organization's facts straight. How do we actually build it? What does it look like? It is a living, working document artifact.
In its simplest form, it's essentially a shared spreadsheet with five explicit columns and one implicit sixth column. The structure directly mirrors the five steps of the framework we just discussed. It forces you to show your work.
OK, let's walk through it column by column. How do I build this? Column one is the claim as it is currently written in your draft or as you found it in the wild. You preserve the original unedited, potentially compressed wording.
So you have an honest before and after record of the transformation. Column two is the primary source traced and cited fully. Author, title, date, publish and the specific page number or hyperlink.
Column three is the check date of measurement and the specific population explicitly written out. Got it. So columns two and three are the unvarnished raw facts of the original source.
Right. Column four is the grade. You assign it an M, P, I or U. And crucially, this column must include a one line written justification tied directly to the specific verb the primary source actually used.
You have to defend your grade. And finally, column five is the restated claim. This is the polished, defensible, highly accurate sentence that will actually appear in your finished document or your presentation.
It is basically a machine for turning bad data into good arguments. You mentioned an implicit sixth column earlier. What is that? The implicit sixth column is the date the row itself was last checked by a human.
This is vital for what we call the currency discipline. Currency discipline, meaning making sure the statistic hasn't gone stale or been debunked since you originally added it to the ledger. Exactly.
A correctly graded claim does not stay correctly graded forever. The world changes. New data emerges.
You need to set review triggers for your ledger. There are specific events in the real world that should trigger you to recheck a row. For instance, a newer annual edition of a study is published or the underlying methodology of the original paper is severely critiqued by peers or the population shifts dramatically.
So how do I manage that? Do I just set a calendar reminder in Outlook to check the whole spreadsheet every December? No, you cannot set a lazy blanket review annually trigger for the whole ledger. You have to tie the trigger to the specific source's actual real world publishing cadence. For example, we talked about the Stanford HAI AI Index report.
That report publishes dependably every single spring. So for any row relying on Stanford data, you set a strict spring trigger. But the WEF Future of Jobs report is completely irregular.
They publish editions in 2018, 2020, 2023, and they're slated for 2025. IMF staff discussion notes have no fixed schedule at all. They just drop when they are ready.
You have to identify and name the specific event that triggers the review for each individual row. OK, this sounds amazing for ensuring absolute accuracy, but I have to put my corporate manager hat on and push back again. If I have to mandate that my team maintains a five column spreadsheet tracking publication cadences for every single number we ever use in any deck, this sounds like a massive soul crushing bureaucratic bottleneck.
My team is going to rebel. They are going to say it slows them down too much. I hear that exact complaint from executives all the time.
But in reality, over a six month timeline, it does the exact opposite. Building a standing ledger speeds up your team's workflow exponentially. How does adding a five step tracing process speed anything up? Because you do not rebuild the trace for every single memo.
Think about your own policy team. Most teams use the same core set of statistics repeatedly across dozens of documents. You have a favorite adoption rate, a standard economic impact estimate, a baseline regulatory count.
You apply the five steps to trace and grade those frequently cited figures exactly once. You do the hard work once. You store them in a central standing ledger that the whole team can access.
From then on, whenever an analyst needs a number for a slide deck, they don't go to Google. They go to the ledger. They just copy and paste the pre-verified, correctly graded, beautifully written restatement directly into their work.
It stops everyone on your team from constantly re-Googling and re-compressing the exact same statistics from memory every single quarter. It is an infrastructure investment that saves hundreds of hours while mathematically guaranteeing your credibility. Ah, I see.
It becomes a verified shared brain for the entire organization. You do the heavy lifting once and it pays dividends forever. But a system this powerful, a shared brain, only works if the people building it apply the rules fairly.
Which leads us to our final core topic, the human element. Even-handed grading and common mistakes. This is the hardest psychological test of the entire method.
The true test of your professionalism isn't how aggressively you trace and grade a claim that you are already suspicious of. The true test is how you handle a claim that you desperately want to be true. Right.
If I am pushing for a specific AI safety regulation and a study drops that perfectly supports my organization's exact lobbying goals, I am highly, highly motivated to just copy paste that headline straight into my deck without looking too closely at the methodology. Exactly. The confirmation bias is overwhelming.
Do you have the discipline to apply the exact same rigorous five-step trace to a statistic that perfectly supports your goals as you do to a statistic that complicates or threatens them? This is the recusal equivalent for a policy analyst. If you give convenient claims a free pass and a generous reading while relentlessly interrogating the inconvenient ones, your ledger becomes completely useless. Worse than useless, actually.
It becomes a massive credibility risk that is wearing the disguise of a credibility tool. You will walk into a boardroom feeling falsely secure because the claim is in the ledger, only to be destroyed by someone who actually read the methodology. That is a terrifying trap.
You feel mathematically rigorous because you have a neat spreadsheet, but you are really just laundering your own pre-existing bias through a fake methodology. What about another trap I see all the time, the single source industry consensus? Yes, this is a very common and very sloppy pattern. A policy memo will cite a single famous executive's public remarks or a tech founder's casual answer on a popular podcast, as if it were a peer-reviewed industry consensus.
They will write a sentence like, studies show AI will replace all entry-level coding work within two years when the only actual source is just one CEO talking on a stage in early 2025. But if the CEO is incredibly famous and runs a trillion dollar company, their opinion carries massive weight in the market, right? Doesn't their status make the claim credible? Their stated expectation is highly relevant market context, yes, but it is fundamentally grade-U evidence until it has an actual empirical data set behind it. You cannot frame a single person's prediction as a finding or a study.
You must frame it honestly to your reader. A senior tech executive stated an expectation that, conflating a famous person's personal credibility with a published rigorous methodology is a classic way that an unsupported assertion borrows unearned authority. OK, I have one final big picture pushback on this entire philosophy.
If I do this perfectly, if I build the standing ledger, if I rigorously hedge my verbs, if I explicitly name my populations and if I refuse to collapse uncertainty, my memos are going to be full of caveats. They're going to sound measured. Meanwhile, my direct competitors are going to be in that same boardroom an hour later, boldly and loudly claiming that their AI solution solves every single problem instantly, perfectly and with 40 percent ROI.
In the real world of corporate theater, won't I lose the argument simply because I sound less confident than the liar? It is a valid fear. You might feel less flashy for about five minutes, but you have to remember your audience. Executive boards, top tier regulators and experienced legislators are heavily, heavily targeted by fluff all day long.
They're drowning in overstated, hyperbolic claims. They have a highly developed immune system against fake confidence. When your competitor makes a wild, unqualified claim and a sharp staffer checks it and finds it completely hollow, that competitor's trust is broken forever.
They are done. When your memo is the only one that survives independent verification by that same sharp staffer, when your claims perfectly match the concrete in the basement, you win the war for long term trust. And in the world of high stakes policy and enterprise sales, long term trust is literally the only asset that actually matters.
Defensibility over flash, winning the long game. That brings us to the end of our deep dive. Let's do a quick summary of this incredible journey we've taken today.
We started with that viral, terrifying IMF headline that AI would destroy 40 percent of jobs. We saw exactly how citation compression driven by the attention economy stripped away the crucial nuance of income brackets and massive productivity gains. We then introduced the solution, the four level evidence ladder, grade M for measured, grade P for model projection, grade Y for interested estimate and grade U for unsupported assertion.
We walked through the rigorous five steps of tracing to the primary source, checking the date and the population, letting the original verb carry the grade and finally writing an honest restatement that respects the boundaries of the evidence. We watched our fictional analyst, Geraldine, save her company's credibility and actually improve their lobbying argument by using a well-placed evidence choice. And we learned how to build a shared standing ledger with currency triggers so our team's statistics never go stale.
It is a brilliant, airtight system. So for the listener who has been with us for this entire hour, what is the single most valuable concrete move they should make when they log into work on Monday morning? On Monday morning, do not try to boil the ocean. Do not try to grade every single claim in your company's entire historical library at once.
That is paralyzing and you will give up. Instead, pick the single most confident, most frequently repeated statistic currently circulating in your organization's own pitch decks or policy materials. The one number that absolutely everyone in the company takes for granted.
Find the golden goose statistic. Yes. Isolate it.
Apply the five step trace to that one single claim. Just one. Go down to the basement and look for the concrete.
If the claim holds up perfectly, that is fantastic. You have verified your company's foundation. But if it doesn't hold up, if you discover it has suffered from severe citation compression or hedge deletion or a correlational shift, you have just found your organization's biggest hidden liability.
And you found it before a regulator, a hostile board member or a ruthless competitor did. You patched the hole in the hull of the ship before it went out to sea and sank. Precisely.
And that makes you incredibly valuable. I love it. But before we sign off, I want to pose one final thought for you to chew on pulling all of this together.
We spent this time talking about how human analysts, journalists, consultants, executives accidentally or intentionally compress nuance because of the incentives of the attention economy. But think about what happens tomorrow. What happens when millions of A.I. agents are explicitly tasked with drafting policy memos and briefing documents at scale? If a language model is mathematically optimized to generate persuasive, highly confident text, it is going to automatically strip away hedges, caveats and population limits faster and more efficiently than any human PR team ever could.
Are we about to enter a recursive loop where A.I. systems ingest compressed data, compress it even further to sound more confident and then seed it back into the policy ecosystem? How do you maintain a ledger of truth when the machine generating the policy is structurally designed to weaponize citation compression at a global scale? That is the exact question that should keep every policy professional awake at night. The discipline we built today isn't just a best practice for today. It is your only survival mechanism for the synthetic future.
An incredibly powerful point to end on. Thank you for joining this deep dive into evidence standards. It is so easy to see a number like 40 percent of jobs destroyed and just accept it because it fits our anxieties or our narratives.
But the next time you see a jagged, terrifying statistic pointed at you in a meeting, don't just stare at it. Trace the foundation, check the population and always, always check the verb. Keep your evidence clean and your credibility will follow.
Catch you on the next deep dive.
Real cases
These examples show the grading method applied to figures a policy professional is likely to encounter directly. Each is cited to its primary or reputable source.
Example 1: The compression chain itself (IMF, January 2024). The gap between IMF Staff Discussion Note SDN/2024/001's carefully hedged, income-group-broken-out finding and the flattened "AI will destroy 40 percent of jobs" version that circulated within days is not a story about any one outlet behaving badly; CNBC, CNN, and Euronews each accurately reported Georgieva's own framing. It is a story about what happens to a Grade P claim after three or four further rounds of informal retelling with no one checking it back against the primary source. This is the anchor case Section 3D traces in full, and it is worth returning to here as the reminder that the compression happens gradually, not in one dramatic misquote.
Example 2: A long-range projection whose range collapses to its upper bound (PwC, 2017). PwC's "Sizing the Prize" report estimated that AI could contribute up to $15.7 trillion to global GDP by 2030, a 14 percent GDP boost, built on an explicit regional and sectoral model (PwC, "Sizing the Prize," June 2017). The figure is Grade P, a transparent long-range economic model, not a measured near-term outcome, yet it is frequently repeated today, nearly a decade after publication and only a few years from its own 2030 horizon, without noting either its age or that it was always a multi-year cumulative projection rather than an annual or already-realized figure.
Example 3: A range that gets quoted as a single number (McKinsey Global Institute, 2023). McKinsey Global Institute's June 2023 analysis of generative AI's economic potential modeled an annual value of $2.6 trillion to $4.4 trillion, built from 63 specific use cases (McKinsey Global Institute, "The economic potential of generative AI," June 2023). Because a range is harder to headline than a single number, secondary coverage routinely quotes only the top of the range, $4.4 trillion, discarding both the lower bound and the explicit use-case methodology that produced it, a Step 3 and Step 4 failure in one move.
Example 4: A projection whose population is a specific employer survey, not "the world" (World Economic Forum, 2023). The Future of Jobs Report 2023's net figure of 14 million jobs lost by 2027 (69 million created, 83 million eliminated) comes from a survey of 803 companies representing 673 million jobs (World Economic Forum, Future of Jobs Report 2023, May 2023). It is frequently restated as what "the global workforce" or "workers worldwide" can expect, when the actual population is a specific, named set of surveyed employers reporting their own hiring expectations, a Step 3 population check every restatement of this figure should carry.
Example 5: A genuinely measured count, for contrast (Stanford HAI, 2025). Not every governance-relevant figure is a projection. Stanford HAI's AI Index Report 2025 counted 59 AI-related regulations introduced by 42 United States federal agencies in 2024, more than double the 25 regulations from 21 agencies the year before (Stanford HAI, AI Index Report 2025, Chapter 6: Policy and Governance, April 2025). This is Grade M: a direct tally of enacted regulatory activity, not a model of future activity, and it is worth holding up beside Examples 2 through 4 specifically because the contrast in verb and confidence is instructive: a measured count earns a flatter, more direct restatement than any projection in this list does.
Example 6: An interested estimate that still cites its own methodology (industry survey practice). A well-run interested estimate discloses its sample size, its selection method, and its sponsor plainly, which is what separates a usable Grade I citation from an unusable one. A vendor or association survey that states "based on responses from 400 self-selected customers surveyed in Q2 2026" gives a policy analyst enough information to grade it honestly and cite it with an accurate population and sponsor disclosed; a vendor claim with no disclosed sample, date, or methodology at all cannot be graded past Grade U no matter how specific the number sounds, because there is nothing behind the number to check.
Example 7: A finding whose income-group breakdown is the whole point (IMF, January 2024, revisited). Return once more to the anchor case, this time for what its own authors treated as the headline finding, not the global average. The IMF's own paper leads with the fact that the 40 percent global figure masks a sharp divide: roughly 60 percent exposure in advanced economies against roughly 26 percent in low-income countries, a gap the paper's authors present as the more policy-relevant number than the flattened global average, because it is the divide, not the average, that drives their call for differentiated policy responses by country income level. A restatement that keeps "40 percent worldwide" but drops the divide has kept the least policy-relevant part of the finding and discarded the part its own authors considered most important.
Example 8: A widely cited figure whose original methodology later shifted (illustrating the currency discipline). Both the World Economic Forum's Future of Jobs Report and Stanford HAI's AI Index Report are actively maintained, regularly updated series, each new edition built on a refined methodology and fresh survey or measurement data rather than a simple rerun of the prior edition's model, but they do not share a cadence: Stanford HAI's AI Index Report publishes every spring, while the Future of Jobs Report has appeared in 2018, 2020, 2023, and 2025, a slower and less regular rhythm. A ledger row built against the 2023 Future of Jobs figures or a prior AI Index edition should carry an explicit note of which edition it cites and a review trigger tied to that specific source's own publishing pattern, not a generic annual assumption, precisely because these are exactly the kind of actively maintained sources Section 3K's currency discipline is built to track, not one-off studies that never get revisited by their own authors.
Example 9: A single-source claim treated as if it were an industry consensus (a pattern, not one incident). A recurring failure across policy submissions in this field is a single individual's public remarks, an executive's conference-stage prediction, a founder's interview answer, cited as if it were a peer-reviewed finding or an industry-wide consensus, when the actual claim traces to one person's stated expectation with no dataset, sample, or published methodology behind it at all. Grading such a claim correctly does not mean the individual's view is worthless; a credible practitioner's stated expectation can be genuinely informative context. It means the restated claim must name exactly what it is, one individual's stated expectation, not a study's finding, which is the distinction Question 11 in Section 10 tests directly.
Where people go wrong
- "If a respected organization is cited, the claim is safe to repeat as stated." The IMF is a credible, careful source; the claim that "AI will destroy 40 percent of jobs" is not what the IMF's own paper says. A respected source's name attached to a claim tells you the primary source is worth tracing; it does not tell you the version of the claim in front of you matches what that source actually found.
- "A range is just a more precise single number, so quoting the top of it is fine." McKinsey's $2.6 trillion to $4.4 trillion range is not the same claim as "$4.4 trillion." A range communicates genuine uncertainty in the underlying model; collapsing it to its highest figure overstates the finding's precision and its confidence.
- "The population doesn't matter if the headline number is accurate." A figure can be numerically correct and still misleading if its actual population (803 surveyed employers, 75 Singapore respondents within a larger global sample, a specific income-group breakdown) gets replaced with a broader implied population ("the world's workers," "businesses everywhere") in the retelling.
- "Softening the verb is just being overly cautious." Changing "could affect" to "will affect" is not caution removed; it is a grade change, from a modeled projection to an unqualified assertion the underlying evidence does not support. The hedge is not decoration; it is the exact word carrying the claim's true evidence grade.
- "An interested party's claim is either fully trustworthy or worthless, no middle ground." Both extremes are wrong. An interested estimate deserves neither blind acceptance nor blanket dismissal; it deserves the same five-step trace as any other claim, graded honestly as Grade I, with extra weight given to whatever independent corroboration exists.
- "If I can't find the primary source quickly, it's safe to assume the secondary source got it right." An untraceable primary source is a finding, not a shortcut. The correct response is to cite the secondary source explicitly as secondary, soften the claim's grade toward Grade U, or keep searching, never to silently borrow the primary source's authority for a claim you were not actually able to verify yourself.
- "A number that has been repeated everywhere for years must be solid by now." Wide repetition is a social fact about how a claim has traveled, not evidence about how it was originally produced. A widely repeated figure still needs its own trace; repetition can just as easily mean a Grade U assertion has achieved consensus-sounding status without ever having been checked.
- "Grading evidence down is always the safer, more rigorous choice." Refusing to trace a claim that happens to support your position, and grading it down out of an excess of caution rather than an actual failed trace, is its own failure of the method. The grade must reflect what the evidence trail actually shows, in either direction, not a thumb on the scale toward whichever conclusion feels more defensible to have reached.
- "Once a claim is graded and restated, it never needs to be revisited." A claim's grade can change: a projection gets superseded by a newer model, a measured count gets revised with better data, a methodology later faces a credible critique. An evidence ledger without review triggers is a snapshot that looks current long after it has stopped being current. (see Topic 6.6) (see Topic 6.10)
- "Restating a claim honestly always makes the argument weaker." As Geraldine's scenario shows, a correctly graded and restated claim can make an argument more precise and, often, more persuasive to a skeptical or well-informed reader, because it survives the exact scrutiny an overstated version invites and fails.
- "I only need to apply this method to claims I am suspicious of." The method applies evenly to every specific, checkable claim in a document, including the ones that already feel obviously true, because a claim's grade depends on how it was produced, not on how confident it feels to the person restating it. A claim that "everyone already knows" is exactly the kind of claim citation compression has had the most opportunity to quietly degrade.
- "A claim that supports my organization's position deserves less scrutiny than one that complicates it." Applying the five-step method unevenly, generous with convenient claims and rigorous with inconvenient ones, defeats the discipline's entire purpose, which is a credibility standard that holds regardless of which direction a correctly graded claim happens to point.
- "One individual's public remarks are as citable as a published study, provided the individual is a recognized expert." A credible individual's stated expectation is worth recording, but it must be restated as exactly that, one person's view, not upgraded into "research shows" or "studies confirm" language the underlying evidence does not support; conflating personal credibility with published methodology is a common way an unsupported assertion picks up borrowed authority it did not earn.
- "A claim my own organization produced internally is automatically Grade M, since we generated it ourselves." An internally produced figure earns its grade from how it was actually produced, a direct count, a model, or an unverified guess passed along as fact, exactly like any external claim; being the source of a number does not exempt it from being traced.
Questions people ask
- What is evidence ledger?
- A working document that records, for every specific, checkable claim in a policy argument, its original wording, its primary source, its checked date and population, its evidence grade, and its restated form, ready to defend line by line. It is the artifact this topic teaches and the third payoff artifact of Module 6.
- What is evidence grade?
- One of four levels (Measured, Modeled Projection, Interested Estimate, Unsupported Assertion) assigned to a claim based on how the underlying figure was actually produced, never on how confidently the claim is stated.
- What is measured (Grade M)?
- A figure counted or directly observed, such as a tally of enacted regulations or a recorded enforcement action, reported by the party that did the counting. Still needs its population and date checked, but is not itself a forecast or a model's output.
- What is modeled projection (Grade P)?
- A figure produced by applying an explicit, published methodology to available data to estimate something that has not yet happened, such as an occupational exposure index or an economic-impact model. Legitimate evidence when the methodology is transparent, but always a model's output, restated with a hedging verb.
- What is interested estimate (Grade I)?
- A figure produced, funded, or prominently amplified by a party with a direct financial or reputational stake in the conclusion it supports, such as a vendor's own ROI case study. Not automatically wrong, but requires disclosed methodology and independent corroboration before it is cited without heavy hedging.
Keep going
This lesson builds Policy analysis and decision-ready writing, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.