Research
Three Models of RLHF Annotation: Extension, Evidence, and Authority
Identifies three distinct normative frameworks implicit in how RLHF annotators are used — as extensions of a designer's intent (Extension), as evidence of population preferences (Evidence), or as authorities whose judgments define quality (Authority). Each model implies different failure modes, quality metrics, and annotator selection criteria. For alignment practitioners, this matters because conflating these models leads to misaligned training signals: treating consensus-seeking annotators as ground-truth authorities, or treating individual expert judgments as population preferences.
Source
↳ Follow the thread