The Any-Annotator Rule for Moral Content Disagrees With a Calibrated Posterior on 30% of Items
Moral Entropy keeps a full Bayesian posterior over the true label instead of collapsing annotator disagreement into a majority vote or the permissive any-annotator rule, and decomposes entropy into aleatoric (irreducible moral disagreement) and epistemic (noisy or insufficient annotation) components. Auditing standard aggregation rules against that posterior across three corpora and fifteen discourse domains, the any-annotator rule disagrees on roughly 30% of items, almost entirely false positives when pooled, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC). The stricter majority and two-vote rules miss 63-83% of true positives, which is the number anyone building a content-moderation or safety classifier on voted labels should look at.
↳ Follow the thread