Synthetic data
Artificially generated data intended to statistically resemble a real dataset without containing actual real records. Synthetic data carries a formal, testable privacy guarantee only when the generative model producing it was itself trained under differential privacy with a stated epsilon; otherwise its privacy properties are unquantified. (see Topic 4.5)
Defined in 5 GAGE programs, which carry 6 distinct definitions of it. The wording above is taught in AI Data Governance: The Data Chair.
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
Artificially generated data intended to statistically resemble a real dataset without containing actual real records. Synthetic data carries a formal, testable privacy guarantee only when the generative model producing it was itself trained under differential privacy with a stated epsilon; otherwise its privacy properties are unquantified. (see Topic 4.5)
Training or fine-tuning data generated by an algorithm rather than collected from real-world sources. A legitimate acquisition option that reduces most personal-information consent questions but remains subject to the same data-quality obligations as any training data, and can carry its own licensing and memorization risks.
Data that was manufactured by a process rather than measured directly from a real person, transaction, or event. It is designed to resemble real records statistically. It comes in three practical families: rule-based augmentation, generative-model output, and simulation.
Data generated by an AI to fill gaps in real data; useful when real data is scarce, but it can amplify the flaws of the model that produced it (including representation gaps from the original collection), so it is best used cautiously and verified before relying on it.
Artificially generated data resembling real data, used to augment (never replace) real data, valuable for rare edge cases and privacy-preserving sharing, dangerous as a primary training signal in high-stakes domains without real-world validation.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- Data Literacy for AI: Datasets, Quality, and Pipeline Basics · Advanced AI Literacy, AI Literacy & Professional Conduct
- Synthetic data: when it protects, when it launders, and how to tell · Feeding the Machines, AI Data Governance: The Data Chair
- Privacy-enhancing technologies as decisions: differential privacy, federated learning, and when your organization actually needs them · The Frontier Discipline, AI Data Governance: The Data Chair
- Synthetic data: when it saves you and when it launders a bias · Data Reality, AI Governance: Applied Mastery
- Enterprise Data Strategy, From Asset to Competitive Moat · Diagnose the Organization, Business AI Transformation
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.