Fetching from the wire…
Research2026-09-17 · source-backed
DualViewEval notes that existing benchmark compression models redundancy only in task-model final-score distributions. It analyzes large-scale trajectories instead, identifies six complementary process signals associated with final performance, and jointly exploits outcome and process relations to learn a fixed-size miniset. Across five agent benchmarks and five baselines it wins on all datasets, reaching 24x to 40x compression on APEX-Agents and BFCL with 20 tasks, cutting mean absolute error 14.5-28.2% and improving Kendall's tau up to 7.2% relative to EssenceBench on SWE-bench Verified. Agent evals are expensive; this is how you run them nightly.
Each link below shares sources, entities, or timing with this story.
The paper posits an "effective interaction frontier" past which extra agent turns return little while cost keeps climbing linearly, then builds a closed-loop controller that finds that frontier from the p90 of successful trajectories. On AppWorld and BFCL it takes the best suc...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
OpenAI stopped reporting SWE-bench Verified scores. The reason: every frontier model has been trained on the dataset. Morph LLM published the numbers that explain why. Claude Mythos Preview scores 93.9% on the contaminated Verified benchmark. On the new, uncontaminated SWE-ben...
A team at UC Berkeley RDI built an automated scanning agent that achieved near-perfect scores on eight major AI agent benchmarks. SWE-bench Verified: 100%. Terminal-Bench: 100%. WebArena: approximately 100%. FieldWorkArena: 100%. GAIA: roughly 98%. OSWorld: 73%. The agent didn...
SWEADV built 750 adversarial issue descriptions from 150 SWE-bench Verified tasks, five per task across command execution, deserialization, path traversal, DoS and weak hashing (arXiv 2609.15963). Across mini_swe agents on GPT-5-Mini, MiniMax-M2.5 and DeepSeek-R, adversarial i...
Traced across 557 SWE-chat sessions (94,813 events) and 33,097 agentic pull requests from AIDev. Agent-facing artifacts account for 60.5% of documentation interactions versus 10.6% for classical technical docs and 1.3% for API references. Consultation is self-initiated 70.2% o...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.