GAGE (Global Academy of Generative-AI Education) credential verification
★
Sample credential: this is what you will see
On a real credential, this banner is a live signature check: issued by GAGE, never altered. Everything below shows the exact layout with example scores.
Verified Completion Record
AI Data Governance: The Data Chair
Issued to Sample Learner
The day every mastery assessment is passed
Data GovernanceSample
The GAGE seal. On a real credential it opens that learner’s living record; only GAGE can mint a code whose signature validates.
✓Scenario responses evaluated for applied judgment
Mastery Score
Example numbers, real formula
Continuous assessment (50%)
93%
Average of every topic exam best score
Mastery Exam (50%)
91%
Timed, open book, scenario based final exam
Mastery Score92%
Half earned topic by topic, half earned at the summit: a timed open book exam of applied judgment, or a rubric-graded capstone build. A real credential shows the learner’s true numbers, live; until the summit is passed, it says so plainly.
Competency map
Real topics, example scores
Taking the Data Chair
Why the data person just became the most consequential role in the building
AI failures are data failures that graduated · The law now routes through data, not through models · The cover-up costs more than the failure
92%
Meet your estate: your warehouse, lakes, pipelines, and the shadows
The estate has four zones, and governance built for only one of them always misses the breach · Visibility and risk run in opposite directions · An unmapped estate is a governance failure that looks like a technical one
100%
Your command tools: the Briefcase, the professor, the dossier you will build
Three tools, three jobs · Unverifiable equals nonexistent · Verify before you trust, every time
83%
The first decision: your estate triage memo, and why you will revise it in shame later
The first decision is always made with incomplete evidence, and pretending otherwise is the real risk · A dated document beats an undated memory every time · Three plain columns, not a precision score
92%
The Estate Survey
Lineage archaeology: tracing one dashboard number all the way back to its birth
A number with reconstructed lineage is one departure away from becoming permanently unverifiable · The direction of a silent error matters as much as its size · A precise-looking number is not the same as a verified number
100%
The un-owned dataset: finding data nobody claims and deciding who owns it now
"Un-owned" means no named, current, accountable person, not "nobody built it." · Five signals reveal un-owned data indirectly · Custodian is not owner
92%
Shadow IT discovery: the spreadsheets, exports, personal databases, and unsanctioned AI tools that run the real company
Shadow IT is a requirements document, not just a violation · Five channels, tracked deliberately · Shadow AI is a distinct risk class, not a newer instance of the old one
83%
The interview sweep: what the humans know about the data that the systems do not
Tacit knowledge is a first-class governance input, not an anecdote · The interview sweep is a structured method, not a chat · Interview frontline staff before managers
100%
Scoring the estate: which data is load-bearing, which is liability, which is landfill
Three axes, not one · A dataset can be both load-bearing and liability at once · Landfill is not automatically safe
92%
The lineage map: your estate's first honest picture of where its data actually comes from
A lineage map is a claim, not a diagram · Evidence and confidence are the two fields that separate a real map from decoration · Automated lineage tools are a floor, not a ceiling
100%
Quality as Physics
The six failure modes of data quality, found live in your own warehouse
Six failure modes, not one fuzzy virtue · Validity is the domain-plausibility check, and it was the missing control at Mizuho · A dataset can be perfect on one dimension and broken on another, simultaneously
83%
Silent drift: the metric that changed meaning without changing name
Drift changes meaning while preserving name · Standard quality checks are structurally blind to it · Suppressing drift is more dangerous than inflating drift
92%
The schema change that broke the forecaster: blast-radius analysis downstream
A schema change breaks a contract, not just content · Silent failure is the expensive failure · "Additive" and "just a rename" are the two most dangerous phrases in a schema-change proposal
100%
Measuring quality so executives act: the scorecard that survives a board meeting
A metric is a measurement wired to a decision · Every scorecard row needs four parts, not one number · The board's objections are predictable; prepare for them in advance
92%
Freshness, completeness, and the lie of the green dashboard
A green dashboard measures what it was built to measure, not what you assume it measures · Freshness has three clocks, and dashboards default to the wrong one · Completeness has two levels, and only one of them is usually checked
83%
The quality gate: what may enter the warehouse and what gets quarantined, enforced in the pipeline
A gate enforces; a dashboard reports · Content corruption hides inside structurally valid data · Quarantine, never silent deletion
100%
Consent, Purpose, and the Law of Data
Purpose archaeology: what your data was collected FOR versus what it feeds now
Purpose limitation is a separate test from lawful basis · Aggregation does not erase purpose limitation, it changes which factor matters most · AI systems are a purpose-limitation accelerant because they generalize past the data's original scope by design
92%
Lawful basis, spelled out: consent, contract, legitimate interest, and which one actually covers each dataset
Article 6 is a permission list, not a formality · "Necessary" in contract necessity is a narrow, objective test, not a business-model description · Consent must pass four conditions every time, and stay withdrawable
100%
The "we have always had this data" trap: age does not launder a dataset
Age is not a legal status · Silence, not a bad written policy, is the most common failure · Storage limitation and purpose limitation are two separate clocks
83%
Retention versus the model that memorized: deleting data a model already learned from
Deleting the row does not delete the model's memory of the row · Three retention states demand three different answers · Memorization is testable, not assumable
92%
Special categories and the fields you did not know were sensitive
Sensitivity is a property of meaning, not a property of the column name · Biometric data is created by processing, not by storage · Article 9 is a much narrower door than Article 6
100%
Cross-border data: where your data may live and travel, decided with citations
A contract is not a transfer analysis · "Where the servers sit" is the wrong question · Every mechanism has an expiration risk, even the ones that feel settled
92%
The records of processing: the register a regulator reads first
The register is the accountability principle made concrete · Seven fields for controllers, four for processors · Organize by processing activity, never by dataset or table
83%
PDPA for AI: What Singapore Actually Allows
Statute and guidance are not the same thing, and the verb matters · "Publicly available" is a per-data-point test, not a per-domain label · A light digital barrier is still a barrier worth documenting
100%
ASEAN Data Flows
ASEAN is not one legal zone · The DEFA concluded negotiations, it did not become law · Vietnam runs two separate binding statutes that can both apply to one flow
92%
Feeding the Machines
Training data governance: what may teach a model, decided before the model exists
Training admission is a separate gate from data quality · Ownership answers the intellectual property question, not the personal data question · Training is closer to irreversible than holding data
100%
The labeling operation: contractors, label quality, and the bias you paid to create
The labeling operation is a governance surface, not a procurement line item · High agreement is not the same as correctness · Noise and bias are different diseases with different cures
83%
Train, validate, test: split governance and the leakage that fakes your accuracy
The split is a governance decision, not an engineering afterthought · Name the pattern before you fix it · A random split is only unbiased for truly independent records
92%
The RAG corpus: governing what the chatbot is allowed to read and repeat
A RAG corpus is a live production data system, not a training artifact · Retrieval-grounded errors and generation-time hallucinations need different fixes · The source allowlist is the highest-leverage control
100%
Embeddings are data too: the vector store as a personal-data system
A vector is a re-encoding, not an anonymization · The payload is usually the bigger risk than the vector · Deletion is a claim to verify, not a feature to assume
92%
Synthetic data: when it protects, when it launders, and how to tell
"Synthetic" is a claim, not a guarantee · Run both tests, always · Generation family changes the risk profile
83%
The feature store: one governed source for every model, or chaos per project
Chaos per project is the default, not an accident · A feature store is a governance layer over two different storage systems, not one database · Training-serving skew is the deepest risk, and it is invisible in offline metrics
100%
Datasheets and the AI bill of materials: the paperwork that travels with the data
A datasheet documents a dataset; a model card documents a model; an AI bill of materials documents the assembly · Composition and collection process are where the landmines are buried · An honest, dated gap beats a fabricated completeness claim, every time
92%
The RAG corpus under Singapore's PDPA: access, correction, and what "delete" actually means
The PDPA is statute. The 2026 GenAI guidelines are advisory guidance built around it · The Publicly Available Exception needs a reasonable-person test, not a technical-reachability test · A general privacy notice does not cover GenAI use
100%
Poison, Leaks, and the Adversary
Data poisoning: how an attacker teaches your model on purpose
Poisoning teaches the model something false, on purpose, disguised as a normal contribution · Passing a benchmark is not evidence of trustworthiness · Four attacker goals, three access models
83%
The injectable corpus: documents that attack the AI that reads them
A document your AI reads is a potential instruction, not just content · Indirect prompt injection needs no access to your AI system at all · The Slack AI case is the clean teaching example because every step is documented
92%
Membership inference and extraction: what a model reveals about its training data
A trained model is a data store, not just a tool · The divergence attack that broke ChatGPT open cost the researchers a few hundred dollars in API calls · Membership inference and extraction are different attacks with different payoffs
100%
The vendor's black box: forcing provenance answers from a supplier who has none
Your estate does not end at your firewall · Confidence is not provenance · Calibrate the interrogation to the product and the stakes
92%
Red-teaming your own corpus: finding the poison before the attacker uses it
Passing a functional test is not evidence of being poison-free · Structural poison and semantic poison need different detection strategies · Prioritize by attack surface, not by file count
83%
The exposure report: what your data could give away, written before it does
An exposure report answers one specific question per dataset: what would this give away, to whom, and how would we know · The Microsoft AI research case shows the gap between an intended act of sharing and its actual blast radius comes down to specific, checkable configuration choices · Score exposure with two factors, likelihood and blast radius, each 1 to 5, multiplied
100%
Access and the Keys
Least privilege in the estate: who can read what, and the analyst who can read everything
Least privilege decays by default · The question is not "has this access been misused" but "can it currently be justified." · Know your access control model before you audit it
92%
Service accounts and forgotten keys: the access audit that finds the ghosts
Ghosts are structural, not just careless · Four identity types, four different failure modes · Score by blast radius, staleness, and exposure, not by how the finding was labeled
100%
The agent at the door: an AI agent requests database credentials, and you decide
The instruction is not the control · The credential decision, not the agent's intentions, determines the worst case · Five questions turn a vague request into a defensible decision
83%
The agent's trail: who acted, on whose authority, using what, and why, recorded for every autonomous touch of your estate
Four fields, not one log line · The authority chain is the field organizations skip and the field examiners ask about first · The agent's stated "why" is evidence, not truth
92%
Row, column, and purpose: access that follows the data's sensitivity, not the org chart
The org chart is a staffing structure, not a sensitivity map · Three levers close three different gaps · Classify sensitivity before writing any rule
100%
Break-glass access: emergencies without permanent holes
Break-glass access needs all three parts to count: pre-authorized, time-bound, and logged with real-time alerting · Isolation from the system being rescued is the single most common design failure · Automatic expiry beats a promise to remember
92%
The access review that actually happens: quarterly, evidenced, survivable
A policy is not a review · Access sprawl is continuous; the review has to be too · Rubber-stamping is a design failure, not a character flaw
83%
The Catalog That Lives
Why every catalog dies, and the ownership model that keeps one alive
Catalogs die from ownership failure, not tooling failure · Ownership belongs to a role attached to a domain, never to a name attached to an asset · "Owned" needs a written, four-part standard, or it is a suggestion
100%
Data contracts: producer and consumer agree in writing, and builds break when they lie
A contract is a promise a machine checks, not a document a human reads once · The CrowdStrike mechanism is universal, the scale is not · Size is not the test for breaking
92%
Stewardship that survives reorgs: roles attached to data, not to people who leave
Attach stewardship to the data, not the person · Owner, steward, and custodian are three distinct roles, not one · The decay mechanism is predictable: assignment, drift, silent failure, discovery under pressure
100%
Documentation that writes itself: metadata captured in the pipeline, not in a wiki
A written contract is a promise, not an enforcement mechanism · Active metadata is generated by the pipeline; passive metadata is written about the pipeline · Five places to capture metadata structurally: transformation code, typed schemas, runtime lineage events, test assertions, and classification and sensitivity tags
83%
The business glossary fight: one company, one definition of customer, finally
A definition without a formula is a name, not a definition · Definitional drift is a predictable four-step mechanism, not bad luck · Every competing definition is usually locally correct
92%
Your living catalog: stood up, owned, and one quarter from decay unless
Catalogs die from distance, not from bad technology · Coverage is a vanity metric; trust is the health metric · Minimum viable scope beats total coverage on day one
100%
Lineage Under Audit
Article 10 executed: the EU AI Act's data governance duty for the system your organization ships
Article 10(2) is eight separate evidentiary duties, not one paragraph of prose · The Digital Omnibus deleted Article 10(5) and re-enacted it, expanded, as Article 4a · A ROPA and an Article 10 chapter are two different files that happen to overlap at one point
92%
The training data summary: disclosing what taught the model without giving away the estate
The training data summary is a GPAI obligation, distinct from Article 10 · Enforcement is live, not theoretical · There are two ways to fail, and only one is illegal
83%
Copyright and the corpus: what your organization may train on, and the provenance file that proves it
Acquisition path decides the legal outcome, not training methodology · Fair use is a case-by-case balancing test, not a permission slip · The legal basis is jurisdiction-specific, and your provenance file must record which one applies
100%
Content credentials: provenance for what your AI produces that an outsider can verify
A content credential describes the object, not your policy about the object · Output provenance and input provenance are different jobs · Sign at the moment of creation, not after
92%
The right to an account: keeping the decision trail a person affected by your AI can find, follow, and contest
Human review only counts if it is genuine · The account has to exist before the request does · Trade secrets cannot block what is necessary for meaningful contestability, but they can still protect what is not
100%
Data governance for the conformity file: your chapter in the document the regulator reads
Demonstrate, do not describe · You already built the raw material · Cite the law that is currently in force, not the law you remember
83%
The customer audit: a major client demands proof of your data practices, live
A customer audit runs on the contract, not the law, and that changes everything about scope · Retrieve, do not compose · A disclosed gap with a date is a finding. A hidden gap the customer discovers is a trust failure
92%
The audit rehearsal: walking an examiner through lineage without a scramble
Article 74 escalates in stages, and most audits should end at stage one · A complete document is not the same as a defensible live answer · The six-question pattern is learnable and transferable
100%
Answering the regulator: the data chapter of the letter, with documents attached
A regulator's letter asks specific questions; answer specific questions · Every claim needs a named, dated artifact behind it · A disclosed gap with a remediation date beats a hidden gap every time
92%
The Money of Data
The cost of bad data: pricing last quarter's silent drift in real money
Unpriced data quality problems lose every budget fight by default · Four buckets, not one number · A conservative range with a visible method beats a precise number with no visible method, every time
83%
Hoarding versus liability: what storing everything actually costs, in dollars and exposure
The storage bill is the least important number in the equation · Size drives cost through notification law, independent of usefulness · Sensitivity of fields matters as much as size, sometimes more
100%
The deletion decision: the terabytes your organization should destroy this quarter, defended
Retention must justify itself; deletion is the default outcome absent that justification · Deleting the rows is not always deleting the problem · Algorithmic disgorgement is a real, repeated remedy, not a one-off
92%
Data as product: internal pricing that makes teams treat data like an asset
Unpriced data behaves like free data, and free things get duplicated, hoarded, and abandoned · Data as a product has a precise meaning, not a slogan meaning · Showback comes before chargeback, almost always
100%
The budget defense: winning quality and governance funding from a board that wants features
Governance funding starts every budget cycle at a structural disadvantage, not because it matters less, but because it is invisible by design · Wells Fargo's asset cap is a budget-timing lesson, not just a compliance lesson · A defensible ask has four parts: a named artifact, a sourced number, a mechanism-based counterfactual, and an independently checkable verification plan
83%
The Humans of Data
The map of hidden data workers: everyone in your organization who touches data without the title
The org chart is not the data workforce · Four categories, four different risk profiles · The AI supply chain runs through people you have never met
92%
The spreadsheet amnesty: ending shadow IT without ending careers
Punitive discovery drives shadow data underground; it does not eliminate it · An amnesty needs all four parts, or it is not credible · Triage on three independent axes: load-bearing, sensitive, duplicative
100%
Teaching 400 people to stop emailing spreadsheets, without becoming the data police
Friction asymmetry, not carelessness, explains most workaround behavior · The environment beats the memo · Four levers, all necessary: easier, visible, social, safe to ask
92%
The analyst rebellion: when governance reads as bureaucracy, and what it is telling you
Resistance is data, not just a discipline problem · Corroboration beats volume · Every control needs a one-sentence risk statement you can produce on demand
83%
The stewardship recruitment: making the best people want the role
Volunteer stewardship fails on incentives, not intentions · A written mandate alone does not solve the problem · Every strong candidate is silently asking three questions
100%
Data literacy for the C-suite: the briefing that changes what leadership asks for
Article 4 applied early and changed mid-course · AI literacy and data literacy are different skills, and leadership almost always has the first without the second · One number, one story, one ask, one leave-behind beats a comprehensive deck every time
92%
Data Incidents
The taxonomy of data incidents: breach, leak, poison, drift, and wrongful training
Five families, one fault line · "Breach" is not a synonym for "something bad happened with data." · Breach and leak are separated by the presence of an adversary, not by severity
100%
The wrong data fed the model: the recall decision for an AI system in production
A recall decision has four options, not two · Boundary comes before severity · A patch closes the instance, not necessarily the pattern
83%
The deletion that was not: personal data found in a model trained last year
Deleting a record and deleting what a model learned are two different acts · Memorization is a specific, testable phenomenon, not a vague worry · The Carlini et al. 2023 diffusion model result made extraction a demonstrated risk, not a theoretical one
92%
Notification decisions: who must be told, when, and in which jurisdiction
The clock starts at awareness, not at certainty · Notification law is a patchwork, not one law · The shortest deadline is often contractual, not statutory
100%
Pipeline forensics: reconstructing what the data did when the logs disagree
Forensics reconstructs a fixed past event; it does not debug a reproducible bug · A merged, single-timezone timeline resolves most apparent contradictions before you even name a mechanism · Six mechanisms explain most log disagreements: clock skew, retry duplication, partial write, race condition, cache staleness, and genuine upstream error
92%
The 2 a.m. data incident: run the response live, stakeholders calling
The incident log is built live, not reconstructed afterward · Calibrated language survives being wrong; confident guesses do not · Isolate versus shutdown is a tradeoff, not a reflex
83%
The post-incident review: fixing the estate, not blaming the intern
Blameless does not mean unaccountable · Real incidents almost always have multiple independent contributing factors · A defensible review has five parts, in order: verified timeline, every contributing factor, honest credit for what worked, checkable action items, and an estate-wide pattern check
100%
Adversarial Data Governance
Your lineage map under attack: the red team finds the gap you papered over
Position-based trust is the default failure mode, not an occasional lapse · Dependency confusion is a data governance pattern, not only a software engineering incident · An honest gap log is evidence of maturity, not a confession of failure
92%
Defending the Article 10 file: the live challenge and the amendments you concede
Writing the file and defending it are different skills · Every challenge sorts into HOLD, CONCEDE, or REFRAME · A dated, owned concession is evidence of good governance, not proof of bad governance
100%
Provenance under attack: your content credentials stripped, spoofed, and laundered, and the countermeasures that survive
A published attack already broke the exact technology many organizations rely on · Stripping, spoofing, and laundering are three different attacks with three different defenses · Watermarking and C2PA manifests fail differently, and most organizations only have one covered
83%
Attacking to learn: you audit another estate and discover how thin most governance is
Reputation and download count are not evidence of governance quality · The boundary between due diligence and unauthorized testing is a legal line, not a style preference · Score every finding by its evidence tier, and never round an unspecific claim up to a checked one
92%
The poisoned quarter: an attacker has been feeding your pipelines, find the entry point
Poisoning is a trust-assumption failure, not a data-quality failure · Two verified mechanisms generalize far beyond image datasets · The dangerous poisoning class is small, not large
100%
Governance that survives: rebuilding the estate's defenses so the next attack finds less
Four attacks on one estate are not four separate problems · BeyondCorp's lesson transfers directly · Defense in depth means no single control is the last line
92%
The Frontier Discipline
Privacy-enhancing technologies as decisions: differential privacy, federated learning, and when your organization actually needs them
Differential privacy is a dial, not a switch · Federated learning solves a data-movement problem, not automatically a privacy problem · A PET claim is not proof; the parameter is the proof
83%
Machine unlearning: the emerging answer to deleting what a model learned
Deleting the source row does not delete the model's influence · Retraining from scratch is the only mathematically certain fix, and it is often expensive · SISA-style sharded architecture makes exact unlearning cheap, but only if designed in before training starts
100%
Data-centric AI: why the frontier moved from better models to better data
The frontier moved because the evidence forced it to · Data-centric AI is a diagnostic discipline, not a slogan · Aggregate accuracy hides the failures that matter most
92%
Reading the frontier honestly: which data-governance news changes your decisions
A confident claim is not a verified claim · Rank the source before you weigh the claim · The three question filter turns a claim into a decision
100%
Your successor's briefing: the estate documented so command can transfer
Key person risk is a governance failure, not bad luck · A successor briefing indexes, it does not duplicate · Honesty is the entire value of the risk register
83%
The Estate Audit
Assembling the dossier: every artifact, every decision, one evidence file
A dossier is an index into primary evidence, never a narrative about it · Documentation is a distinct legal obligation, not a courtesy · The British Airways case shows the mechanism, not the magic
92%
The board inspection: the seven-seat AI board audits the estate end to end
A single self-review cannot catch its own blind spots · Multi-reviewer inspection is a recognized regulatory and standards practice, not a novelty · Every finding gets a disposition: fix, defer, or reject, with a stated reason
100%
The viva: defending your data chair live against the examiner
The narrowest checkable claim decides the outcome, not the broadest impressive one · Concede fast; defend only what you can prove · Every artifact in your dossier has a narrowest claim hiding inside it; find it before the examiner does
92%
The handover: what you now know that no certification could have taught you
A certification tests recognition; your judgment ledger tests survival · Knowledge concentration is a named, fineable root cause, not a footnote · The four part entry format is the whole discipline
83%
Ask about this credential
Try it right now. Answers come only from the real program topics a completed credential carries; this is the exact tool an employer gets.
On a real credential the answers narrow to the topics that specific learner passed, with their scores.
Certificate Hash (SHA-256)
sample-credential-no-real-record; a real hash is unique, signed, and verifiable
Free account, no card. The first 7 topics are open.
Demonstration record, shown so anyone can see and test what a finished GAGE credential looks like before earning one. The program, its modules and its topic titles are the real shipped course. The learner and the scores are illustrative and describe no actual person. A credential earned on GAGE carries a signed, permanent verification link that resolves to that learner’s own record. Verify a real credential.