Fetching from the wire…
Public story · 2026-09-02 · high
The gain comes from a planning-and-testing outer loop, tested across three model pairs and sustained past 70 iterations.
Why now: The result is notable because it holds across three unrelated model pairs instead of a single tuned demo.
Harness-of-Harness wraps existing AI coding harnesses in an outer loop and boosts task success 52.25% on average across three model pairs. For anyone already running Codex, OpenCode, or a similar setup, the gain sits entirely in the outer loop, so swapping in the new loop doesn't mean swapping out the harness underneath.
The paper tests three pairings, Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3, against three benchmarks: GameCraft-Bench, FrontierSWE and ProgramBench. The outer loop scopes work into small, verifiable increments and keeps the coding agent's own tests separate from an independent evaluation pass. After three iterations, the best pairing improved 82.86% over its standalone harness. The setup kept improving past 70 iterations in a multi-day run, not a one-off benchmark score.
For builders running a coding harness now, the practical move is adding the outer loop rather than replacing the inner one. The paper doesn't say whether the same gain holds on work outside these three benchmarks.
Each link below shares sources, entities, or timing with this story.
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
An r/LocalLLaMA post at 1,330 upvotes reports the first run of full K3, Moonshot's 2.8T open-weight MoE, on a 16x NVIDIA GB10 cluster with dspark speculative decoding: 20+ tok/s average, 38 peak, 750 prefill. That's roughly $64K of hardware for frontier-adjacent tokens at your...
SkillSentry (arXiv 2608.09253) targets the gap where an agent completes a task under skill guidance then fails the same task on a repeat run. It defines a DSL for runtime guidance, initializes it from skill specs plus insights mined from historical successful *and failed* trac...
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0%...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.