Fetching from the wire…
Research2026-09-19 · source-backed
arXiv 2609.19387 hands the same agent the same 15-dimensional accelerator space twice: once as named architectural knobs with simulator counters, once as anonymous variables on [0,1], with evaluator and reachable optima identical. On a nine-kernel FP16 GEMM basket the informed agent beat a modeled H200 by 5.4% and its blind counterpart by 12.3%, using 70.1% fewer simulator calls. A critic loop recovered most of the blind agent's gap and bought the informed one nothing. arXiv Domain knowledge and structured critique act as substitutes. Authors flag five to six runs per condition on one modeled accelerator as preliminary.
Each link below shares sources, entities, or timing with this story.
WebMASLab holds task, tools, and browser fixed and varies only architecture. The Telephone Loop attack exploits cross-agent delegation to create cyclical task loops, averaging 80% success with 0% detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2, and GPT-5.4...
Power availability is now a primary limit on AI infrastructure growth, but making training power-flexible requires knowing how throughput responds to reduction, which nobody had characterized. The index is a normalized metric for the performance cost of a power cut that double...
Activation probes are usually evaluated against agents who don't know they're monitored, which is a generous assumption. This study held models, probes and thresholds fixed and varied only the disclosure: nothing, monitor present, or monitor present plus last round's score. Ac...
The failure they target is specific and under-discussed: a cached error page or a negative price returns in the *expected schema* and gets consumed as fact, unlike a timeout the agent can see. Outcome Monitors check results against contracts mined from task-disjoint traces or...
Skill self-evolution methods revise skill text from execution feedback, but each oracle evaluation needs a full agent rollout, which confines search to patching whatever just failed (arXiv 2609.15396). SkillLift treats ranking as a smoother supervision target than absolute sco...
A controlled study ran five Qwen models over eight cases against a DWSIM simulator, 120 slots per arm, with one instruction as the only difference: request a fresh simulation after a substantive modification. No hard gate. Re-verification happened in 94 of 120 guided slots aga...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.