Skills
LLM judges conflate trust with truth, and swapping the stated author flips both scores together
Testing whether LLM-as-judge can score trustworthiness and factuality independently, researchers found judges align trust scores with truth verdicts more tightly than the human behavioral reference does. Changing only the source attribution, whether content appeared to come from a human or an AI, shifted trust assessments and truth classifications in lockstep. The practical consequence for anyone running a multi-dimension rubric: a trust score is not standalone evidence about factual accuracy, and any judge prompt that reveals authorship is contaminating its own correctness verdict.
↳ Follow the thread