Calibration
The property that a system's stated confidence matches its real accuracy, so that instances it reports as 90 percent confident are correct about 90 percent of the time on your data. A confidence threshold is only a safe boundary if the confidence is calibrated; an uncalibrated model routes confident errors past the human.
Defined in 5 GAGE programs, which carry 10 distinct definitions of it. The wording above is taught in AI Governance: Applied Mastery.
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
The property of a system's confidence scores accurately reflecting its actual reliability, such that a stated ninety percent confidence corresponds to being correct roughly ninety percent of the time on that class of input. A well-calibrated system is the engineering target this topic argues for, rather than simply a maximally cautious one.
The property that a system's stated confidence matches its real accuracy, so that instances it reports as 90 percent confident are correct about 90 percent of the time on your data. A confidence threshold is only a safe boundary if the confidence is calibrated; an uncalibrated model routes confident errors past the human.
The degree to which a forecaster's stated confidence matches reality, so that of all things called 70 percent likely, about 70 percent occur. The defining trait of Tetlock's superforecasters and the real test of a forecaster, human or machine, above raw accuracy on any single call.
The property of a confidence score being honest: among predictions the model reports at a given confidence, it is correct at roughly that rate. An overconfident model (says 90, right 70) is dangerous because users trust predictions they should question.
The process of characterizing and correcting systematic sensor errors (bias, scale factor, misalignment, distortion) by measuring the sensor against a known ground truth under controlled conditions.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- AI-Suitable vs. Human-Essential Decision Trees · Critical Thinking and Context Engineering, AI Literacy & Professional Conduct
- AI for Innovation: Predictive Analytics Applications · Advanced AI Literacy, AI Literacy & Professional Conduct
- Self-Assessment: Comprehensive Knowledge Check · Assessment and Continuous Learning, AI Literacy & Professional Conduct
- The trust boundary: what this system may decide alone and where a human signs · Evaluation and Trust, AI Governance: Applied Mastery
- The evaluation report: evidence that your trust levels are earned, not hoped · Evaluation and Trust, AI Governance: Applied Mastery
- The viva: defending your command live against the examiner · The Capstone: The Board Audit and Viva, AI Governance: Applied Mastery
- Analytics, Telemetry and Observability · Technology, Platforms and Vendors, Business AI Transformation
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.