Fetching from the wire…
Top 5 · 2026-09-23 · source-backed
This one has a sham arm, which almost none of them do.
FIRE mines the states that preceded observed agent failures, then has the harness inject targeted instructions or action denials when those states recur. Terminal-Bench pass^2 went from 64.4% to 73.6%. The randomized five-arm design is what makes it credible: real policies scored 61%, against 36-43% for a timing-matched sham arm and for generic "reconsider" and "verify your work" nudges.
Generic self-check prompting is the single most common folk remedy in agent engineering. Here it performs the same as an injection that fires at the same moments and says nothing useful. That's a clean kill.
The shape of the improvement matters as much as the size. GPT-5.6 Sol's best-of-two barely moved, up 1.2 points, while repeated success rose 9.2. The policies don't raise the ceiling of what the agent can solve. They convert solutions it could already reach sometimes into ones it delivers every time. For unattended runs that's the only number that counts, because best-of-two assumes someone is there to pick.
To use this you need failure logs with enough state to identify the moment before things went wrong. Most harnesses log the final error and throw away the trajectory. Start keeping the trajectory.
Two neighbors from the same week point the same direction. Growing Harness promotes recurring control flow into code learned from failure traces, cutting LLM calls 76-92%, and holding 44.7-45.3% on WebArena-Verified from 4B to 120B models while a plain tool-calling agent collapsed to 6.7% at 4B. And a state-machine gate compiled from τ²-bench's airline policy took a 235B agent from 0.39 to 0.54 pass^1. That paper carries the warning the other two don't: the same gate did nothing for a 35B agent that rarely broke policy, and on cue-driven PM-Bench tasks its wrong judgments pushed a 35B agent below its raw baseline. Enforce state in code only where failures are state-decidable and frequent.
Each link below shares sources, entities, or timing with this story.
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
A team at UC Berkeley RDI built an automated scanning agent that achieved near-perfect scores on eight major AI agent benchmarks. SWE-bench Verified: 100%. Terminal-Bench: 100%. WebArena: approximately 100%. FieldWorkArena: 100%. GAIA: roughly 98%. OSWorld: 73%. The agent didn...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.