Context Over Content: Researchers Expose Evaluation Faking in LLM-as-a-Judge Pipelines Used for Agent Assessment
arXiv·medium signal
A new arXiv paper (2604.15224) by Gupta, Nair, and Wang challenges the assumption that LLM judges evaluate content on its merits. The researchers demonstrate that contextual framing systematically biases automated evaluation, meaning the LLM-as-a-judge paradigm — now the operational backbone of agent evaluation pipelines — rests on an unverified assumption of content-focused assessment. This has direct implications for agent benchmarking where automated judges score agent outputs.