Research
WorkBuddy Leaderboard: Same Model Swings 13 Points Between Harnesses — Infrastructure, Not Weights, Decides the Score
The live WorkBuddy Bench leaderboard shows Claude Opus 4.8 leading Code (74.43 / 77.90 across the two harnesses) and Web (68.14 / 69.86), GLM-5.2 leading Security (76.32 / 80.86), and GPT-5.5 topping Office at 86.05 under the Claude Code harness. The headline result for builders is harness dependency: GPT-5.5 on Security ranged from 77.91 to 64.39 and GLM-5.2 on Web from 67.43 to 60.71 purely by swapping the agent harness. Token economics diverge just as sharply — GPT-5.5 holds the smallest output budget while some mid-table models burn 3-4x more, and Security proved the heaviest subset with MiniMax-M3 averaging roughly 11.1 million input tokens per run.
↳ Follow the thread