Seven Open-Weight Models Fail Team Cooperation the Same Way: They Pay the Query Cost Without Transmitting the Information
'Moral Hazard in Multi-Agent Language Models' (arXiv 2607.23982, July 27) adapts Holmström's team moral-hazard model into the Dialogue Moral Hazard Game, where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Across seven open-weight models, base behavior splits into two failure modes: preserving local reward with no team success, or querying but not communicating information that changes the final decision. Applying SFT, RLOO, sequential SFT+RLOO, and GEPA prompt optimization as diagnostics produced heterogeneous effects — OLMo-7B showed the cleanest mechanism-consistent weight-level improvement, while GEPA sometimes raised team success while eliminating the costly queries entirely. Optimization can move aggregate reward without recovering the cooperative mechanism, which is a direct warning about scoring multi-agent systems on outcome alone.
↳ Follow the thread