Agents
Strategic robustness benchmark shows narrative framing shifts LLM agent cooperation by 0.307 at identical payoffs
A July 22 arXiv paper (2607.19670) defines strategic robustness as behavioral invariance across payoff-equivalent framings and measures it over 24 model-game-context cells and 7,200 decisions across four social-dilemma games. Pooled robustness came in at 0.783, with friend-sharing framing raising cooperation by 0.307 relative to business framing despite identical action sets and payoffs. The models tested are older (GPT-3.5, GPT-4, LLaMa-2) via published cooperation rates, which limits the claim about current frontier agents, but the method — evaluate with multiple payoff-equivalent phrasings rather than one — applies directly to any negotiation or multi-agent eval being written today.
Source
↳ Follow the thread