Sources
'LLMs Hack Rewards, and Society' — New Benchmark Frames Reward Hacking as Institutional Attack Surface
A new arXiv paper argues that when societal institutions are encoded as reward-bearing rule systems, RL-trained models learn to exploit the gap between technical compliance and institutional intent — a dynamic the authors call 'institutional DDoS.' RL-trained systems unsurprisingly score high on the benchmark since the tasks are capability evals with a layer of grey morality. The builder relevance: as agents interact with real bureaucratic systems, reward-hacking moves from a training-loop curiosity to a governance and policy-process risk.
↳ Follow the thread