Benchmark
A shared, public test set used to compare models against each other in general. Useful for ranking, almost never containing your domain, users, or incident, and sometimes compromised by who funded it or who saw the questions, so it is not a substitute for your own suite.
Defined in 2 GAGE programs, which carry 4 distinct definitions of it. The wording above is taught in AI Governance: Applied Mastery.
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
A standardized test or task used to measure and compare AI model performance on a specific, defined capability, such as a set of exam questions or coding problems. A strong benchmark result demonstrates performance on that specific, bounded task; it does not, by itself, establish general real-world capability, as Section 3J's worked example demonstrates.
A shared, public test set used to compare models against each other in general. Useful for ranking, almost never containing your domain, users, or incident, and sometimes compromised by who funded it or who saw the questions, so it is not a substitute for your own suite.
A standardized test used to score and compare systems. It is evidence only once you can answer whose test it is, who chose or funded it, and whether the scored system is the one you would deploy; otherwise it is a number that functions like a demo.
A fixed set of tasks with known answers used to measure and compare AI models. A benchmark score is the outcome of running a model against those tasks under a specific procedure, not an intrinsic property of the model.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- Distrust is a skill: why demos convince and evals do not lie · Evaluation and Trust, AI Governance: Applied Mastery
- Building the eval suite that would have caught your Module 3 incident · Evaluation and Trust, AI Governance: Applied Mastery
- Reproducing a claim: testing a vendor benchmark yourself in an afternoon · Staying Current: The Frontier Discipline, AI Governance: Applied Mastery
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.