Reddit
A model-neutral agent harness matched Claude Managed Agents' solve rate on 63% fewer tokens
The TrueForge team ran 14 DevRev Enterprise-Bench tasks three times each with a blind judge across harness/model combinations. Claude Managed Agents with Opus 4.8 solved 11/14 at $11.8 and 10.0M tokens per run; TrueForge with the same model solved 11/14 at $8.6 and 3.7M tokens, averaging 19 tool calls per task against 32. Swapping in GLM-5.2 gave 11.7/14 at $3.0 per run, roughly 75% cheaper than the managed baseline, which is the strongest published argument yet that the agent loop, not the model, is where the token bill lives.
Source
↳ Follow the thread