Fetching from the wire…
Research2026-09-17 · source-backed
AutoTuneBench characterizes four failure modes from a four-day corpus of 619 model calls where agents tuned GPU kernels in a propose-measure-keep loop: strawman baselines manufacture speedups, absolute times don't transfer across machines, saturated tasks nullify comparisons, and infrastructure defects impersonate science. The fix freezes the protocol as code with test-enforced provenance, a database-level validator rejecting out-of-protocol results, anti-cheat checks outside the agent's modification surface, and a 5% cross-run coefficient-of-variation cap. Under it, one config delivers 1.174x on one machine and 1.0049x on another, and KernelBench Level-1 admits 51% of tasks at a median 1.0001x over PyTorch eager. Any self-optimizing agent that reports its own speedups needs a baseline it cannot touch.
Each link below shares sources, entities, or timing with this story.
SOL-ExecBench measures AI-generated GPU kernels against theoretical hardware speed-of-light limits rather than relative rankings. Current agentic systems achieve 40–70% of theoretical hardware efficiency, with clear headroom. As agents increasingly generate and optimize GPU co...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
Gemma 4 12B dropped June 3, and the spec sheet is the kind of thing I read twice to make sure I wasn't misreading it. 11.95 billion params, Apache 2.0, reads text, image, audio, and video. No separate vision encoder. No separate audio encoder. The model handles all of it nativ...
arXiv 2608.03609 formalizes agentic systems over relational data as Stateful Tool-Enabled Agentic Deployments and proves verification against First-Order CTL specs is undecidable. Under a finite-domain restriction it becomes PSPACE-complete, but only if renaming opaque identif...
NVIDIA's GEAR Lab, with CMU and UC Berkeley, released a closed-loop system where coding agents reset physical scenes, run hardware trials, verify outcomes, and rewrite code until a policy works. Jim Fan calls it "AutoResearch in the physical world." Agent teams hit 99% pass@8...
Simulating HBM, DRAM and SSD tiers against a random-forest execution-time predictor across chat, agent and document QA workloads, tiering supported 73.02x more concurrent sessions per GPU at 62.04x lower cost per session. The authors attribute the gains to tier capacities of 1...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.