Good Start Labs has evidence that agents trained on a board game get better at finance research, but only with multi-turn harnesses
Alex Duffy and Tyler Marques spun Good Start Labs out of Every in October 2025 with $3.6M from General Catalyst, Inovia and Every, selling trajectory data and RL environments to frontier labs. Their September 2026 experiment trained a 30B model on 1830: The Game of Railroads and Robber Barons under two harness designs; the terminal-agent multi-turn tool-use design improved both in-game performance and finance-research benchmarks, while the single-turn design improved only in-game play. Duffy's framing — "how you design that [the harness] totally changes what the model can learn" — is the actionable claim, and their Diplomacy runs separately showed Claude Fable 5.1 and GPT-6 Astra diverging on betrayal behavior.
Source
↳ Follow the thread