Reinforcement learning from human feedback (RLHF)
A preference-tuning method in which human judgments of which answers are better are used to further train a model toward preferred behavior. It is a major reason two versions of the same base model can behave differently, and it is the kind of tuning that produced the divergence in the Maverick episode.
Defined in 2 GAGE programs, which carry 3 distinct definitions of it. The wording above is taught in Certified AI Governance Professional (CAIGP).
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
A preference-tuning method in which human judgments of which answers are better are used to further train a model toward preferred behavior. It is a major reason two versions of the same base model can behave differently, and it is the kind of tuning that produced the divergence in the Maverick episode.
A specific fine-tuning technique in which human reviewers rank or rate model outputs, and the model is further trained to produce outputs more like the highly rated ones; a major factor in the shift from raw base-model behavior to a consistently helpful, instruction-following deployed assistant.
A weight-layer method for tuning a model in which human ratings of its outputs are used to nudge the model toward the rated-good behaviors. A common place the alignment tax appears.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- Fixing the model, breaking it again: why fixes are never free · Build Before You Govern, Certified AI Governance Professional (CAIGP)
- Prompts, context, and why the same model gives different companies different answers · Build Before You Govern, Certified AI Governance Professional (CAIGP)
- How AI Actually Works Under the Hood: Architectures and Limits for Policy People · Technical Credibility Deep Dive, The AI Lobbyist: Certified AI Policy Strategist
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.