Fetching from the wire…
Agents2026-09-15 · source-backed
Across 614 problems from APPS, HumanEval+ and LiveCodeBench, hierarchical collaboration was worth 2.4 pass@1 points on the easiest problems and 21.1 on the hardest, at a flat ~10x token cost throughout (arXiv 2609.13890). DATS predicts each topology's success probability and picks the one maximizing predicted success minus cost, using a graph network treating topologies as ordered nodes, which is worth 1.7 points over a flat multi-label head. Held to 40% of always-hierarchical spend it reached 77.7% pass@1 against 73.6%. Holds across four model variants spanning 14 capability points.
Each link below shares sources, entities, or timing with this story.
Ockhamareto (arXiv 2608.24473) reinforces a unit-test rollout only when it's non-dominated on both mutation-killing and test count, then ties each test's killing power back to specific source tokens. Against MIST-RL that's a 3.4x better per-test trade-off, plus 30 to 35 percen...
Recuris (arXiv 2608.24876) keeps a Working Memory tracking current task progress separate from an Experiential Memory of learned skills, so skill selection indexes against what the task needs now rather than the whole history. It improves 35 of 37 model-benchmark pairs, gains...
Formalizes the "overthinking" problem — models spending excessive compute on simple problems. Dynamically aligns reasoning depth with difficulty. Could significantly reduce inference costs for reasoning-heavy models by avoiding unnecessary compute on easy queries. arXiv 2603.0...
A July 28 study on HumanEval+, MBPP+, and LiveCodeBench found real original tests moved Qwen3.6 on LiveCodeBench from 13.1% to 39.4%, while stronger-model-generated synthetic tests added 1.7 points at p = .701, statistically indistinguishable from nothing. Spend your retrieval...
Wafer.ai ran K3 at TP8 on a single MI355X node: 952 tok/s aggregate, 118 tok/s single-stream decode, ~13k tok/s steady-state prefill after fixing missing PyTorch sampling functions and AITER MLA prefill kernels via SGLang and ROCm. Against a two-node TP16 B200 deployment that'...
Everyone stuffing context into a coding agent has the same instinct: more surrounding code is more signal. Grab the neighboring files, pull in the commented examples, give the model a rich view of the module. arXiv 2609.09242 tested that on LiveCodeBench and the answer is wors...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.