Dispatch
Harness-Bench shows a 23.8-point spread on identical model weights, and Anthropic deleted 80% of Claude Code's system prompt without losing capability
Dan McAteer's Latent Space essay traces the agent harness through four phases from ReAct (Oct 2022) to Claude Code (Feb 2025, ~$1B ARR in six months), and cites Harness-Bench scoring the same model between 52.4 and 76.2 depending only on the harness around it. OpenAI tripled GPT-5.6 Sol's ARC-AGI-3 score from 13.3% to 38.3% through harness changes alone, while Anthropic recently removed 80% of Claude Code's system prompt with no capability loss. The argument for builders: harness capability is being absorbed into weights via RL, so what remains is an interface for managing scarce human attention, not for managing the model.
Source
↳ Follow the thread