Fetching from the wire…
Agents2026-07-20 · source-backed
arXiv 2607.15845 organizes agents around explicit knowledge representations for generating multi-step workflows. It's part of a cluster this week (ToolVerse, AnovaX, the ARC ablations) suggesting the field's attention has moved from single-agent capability to the scaffolding that produces and validates plans. That's the right move and it's roughly two years behind where practitioners already are.
Each link below shares sources, entities, or timing with this story.
Sergey Rodionov's paper tests four Codex-based agent variants to isolate what actually drives performance. Verification (simplification plus exact observation reproduction) ranked highest in every setting, but at substantially higher cost. The textual baseline beat the executa...
ZeroDayBench (2603.02297) — GPT-5.2, Claude Sonnet 4.5, and Grok 4.1 all fail at autonomous zero-day vulnerability discovery. Reality check: the CyberStrikeAI threat is automation of *known* exploits, not novel vulnerability discovery. ICLR 2026 Workshop. tau-Knowledge (2603.0...
Chollet confirmed GPT-5.2 at 46-50%, Poetiq at 54%. But he cautions: "Saturating ARC does not mean we have AGI." ARC-AGI-3 launches March 25 with harder abstractions. For builders, the current benchmark race validates that agent scaffolding improvements (not just model improve...
The trajectory-mining pipeline that segments, clusters, and trains a skill-aware policy produced clean skill clusters but only +1.95 points on one benchmark and negligible gains on another (arXiv:2606.20363). A useful negative signal against the hype: generating skills from in...
HiDream-O1-World generates explorable 3D worlds from a text prompt, image or interactive control, scoring 80.9 average on Navi plus 73.3 on Physical and 88.0 on Consistency across 289 multi-turn cases. The basis is 3D priors injected into a memory context plus test-time traini...
Twin has a frontier coding agent write a simulator of an unknown grid game, and the harness refuses to let the agent act until that program reproduces every observed transition, with each mismatch becoming a counterexample used to repair the model (arXiv 2608.14490). It clears...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.