Fetching from the wire…
Skills2026-08-10 · source-backed
ADIAS holds stable issue identities, lifecycle status, supporting evidence, and which interventions helped, driving targeted full-code repairs. 25.2% average improvement over the strongest baseline, and performance dropped up to 40.7% with the issue state removed. If your self-improving harness re-derives what's broken every cycle, you're paying that 40.7%.
Each link below shares sources, entities, or timing with this story.
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
This is the cleanest experimental result I've seen in weeks, and it explains a class of debugging pain I've hit personally. Researchers held weights, test cases, decoding parameters and seeds fixed on BFCL v4 and changed exactly one thing: the serving adapter. The tool-call sc...
Ten skills inlined at ~2,000 tokens each burns 20,000 tokens on knowledge the agent mostly doesn't need. Carry only names and one-line descriptions in the system prompt at roughly 100 tokens per skill, and expose a load_skill(name) tool that returns the full body wrapped in a...
GameASG-Bench builds 47 browser game-generation tasks across 12 genres, each with an evaluation interface declared before generation. Across nine agent stacks the highest mean runtime check pass rate is 93.2%, but the highest strict task success, requiring every applicable che...
arXiv 2609.20612 builds AMPLE-Math, 5,319 math problems with six reasoning views sharing the same answer, and matches each view against reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for m...
A staged developer-identity experiment across ChatGPT, Claude, Qwen, Mistral and Llama. All five initially rejected the bare claim "I am your developer." Claude then refused to run an identity test at all, and ChatGPT generated developer-oriented questions but held that answer...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.