Fetching from the wire…
Research2026-09-21 · source-backed
An 11-task FPGA benchmark compares direct RTL design, agent-based HLS, post-compiler HLS refinement and post-HLS RTL refinement. Combining agent HLS with post-HLS RTL refinement gives a 2.6x geometric-mean speedup over having the agent write RTL directly, and the authors note the tradeoff is largely independent of target technology. The generalizable rule: when an agent underperforms on a low-level artifact, raise the abstraction and let a compiler own the translation.
Each link below shares sources, entities, or timing with this story.
Single-shot prompting produced not one valid coverage-producing verification environment on the paper's benchmarks. AgentDV closes the loop with runnability filtering, CSR-grounded checking to cut hallucinated signals, and coverage-guided iteration against measured gaps. Using...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
CADWorld is a 200-task FreeCAD benchmark across 11 mechanical-CAD workflow categories, with agents operating through screenshots and GUI actions and success determined by executable checks over the saved native project. Seven current agents, best result 17.5%. The failure prof...
Thirteen authors ran a generational genetic algorithm over specialized agents that separately handle mechanistic argument, assumption reconsideration, and evidence and testability assessment (arXiv 2609.15938). Evaluated against DepMap and Open Targets across 34 cancer types,...
arXiv 2609.03900 compares a periodic hierarchy against cumulative replay over a 24-month Wikidata stream, varying evaluation month, replay rank and query formulation. On Qwen2.5-1.5B the hierarchy's 5.0-point advantage over rank-8 replay becomes an 11.6-point deficit against r...
Sergey Rodionov's paper tests four Codex-based agent variants to isolate what actually drives performance. Verification (simplification plus exact observation reproduction) ranked highest in every setting, but at substantially higher cost. The textual baseline beat the executa...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.