Agents
Principal-agent contract theory extended to agents that choose their own model and token budget
arXiv:2608.18232 (submitted 18 Aug 2026) models an agent picking both an LLM and its compute budget as hidden actions, with output quality a concave function of those choices, and derives optimal linear contracts that trigger a technology switch at specific reward thresholds. Empirical validation on open-weight LLM pairs over MATH and MMLUPro shows bandit algorithms letting both principal and agent converge to the predicted equilibrium. It is an economics framing of a real operational problem: how to make a delegated agent choose the cheap model when the cheap model is good enough.
Source
↳ Follow the thread