A 40,726-Request Replication Fails to Reproduce the Rating-Versus-Ranking Bias Reversal in Hiring, Lending and Triage
Testing whether a published charitable-aid finding (models favor minority applicants when rating one at a time but penalize some when ranking side by side) generalizes, researchers sent 40,726 requests to five models with applications differing only in the applicant's name and a primary test fixed before collection. None of 36 planned contrasts survives correction; the rating advantage keeps its sign at roughly half the published size and the hiring ranking penalty is bounded below the published effect, with planted disparities tracking their injected sizes to validate the nulls. The audit instrument dominates the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect measured.
↳ Follow the thread