Model provenance
The documented, verifiable chain of custody for a dataset or model, including who created it, how, from what sources, and how its identity can be independently confirmed. Established through mechanisms like datasheets for datasets and platform organization-verification signals; the primary structural defense against poisoning attacks that rely on impersonating a trusted source. (see Topic 4.8)
Defined in 2 GAGE programs, which carry 2 distinct definitions of it. The wording above is taught in AI Data Governance: The Data Chair.
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
The documented, verifiable chain of custody for a dataset or model, including who created it, how, from what sources, and how its identity can be independently confirmed. Established through mechanisms like datasheets for datasets and platform organization-verification signals; the primary structural defense against poisoning attacks that rely on impersonating a trusted source. (see Topic 4.8)
The origin and ownership of the model a product depends on, including whether the core model is the vendor's own or a third party's, which upstream providers and sub-processors are in the chain, and where the training data came from. Undisclosed provenance means depending on a supply chain you cannot see.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- Data poisoning: how an attacker teaches your model on purpose · Poison, Leaks, and the Adversary, AI Data Governance: The Data Chair
- The vendor interrogation: questions that expose what the sales deck hides · Shipping AI and Surviving the Incident, AI Governance: Applied Mastery
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.