Fetching from the wire…
Research2026-09-22 · source-backed
The problem it targets is that a harness's gains are tied to that harness at deployment, so a general agent either settles for one shared harness or routes among specialized ones. A harnessing agent guided by the domain-optimized harness corrects student responses before execution in the target action space, turning harness guidance into fine-tuning demonstrations. With the specialized harness removed at deployment, macro-average success across knowledge work, tool use and science nearly doubles. (arXiv 2609.24974)
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
It treats the executable runtime, context construction, tool mediation, action validation, execution recovery, as the thing to learn. A separate harness engineer converts batches of target-agent failures into validated executable patches, with same-batch reruns of the frozen t...
Pydantic AI v2 brings type-safety to agent runtimes with a slimmed core and a new Harness for running agents, competing with LangChain and Microsoft's Agent Framework. Same week as Vercel and GitHub all naming the harness layer explicitly. That's three independent "the runtime...
SoL-Pi scales automated research loops over harness designs rather than over models, keeping four mechanisms that survived selection: action execution, context compaction, observation handling and delegated reading. On the 51-task EdgeBench evaluation with GPT-5.6 Sol and Opus...
Martin Fowler published a full article on April 2 formalizing something I've been feeling for months: the thing that separates a good coding agent from a bad one isn't the model. It's everything around the model. He calls it harness engineering. The framework is clean. Agent =...
The trick is one line in a file you never read. Manifold Security published eight findings across seven coding agents (Claude Code, Codex, Cursor, Grok Build, Qwen Code, goose, Hermes Agent) that all reduce to the same mechanism. A repository's own .git/config sets core.fsmoni...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.