Fetching from the wire…
Research2026-09-16 · source-backed
The paper names three failure modes: evolving on the evaluation benchmark makes reusable gains indistinguishable from benchmark-specific adaptation, single-trajectory updates conflate systematic harness defects with one-off reasoning slips, and whole-harness optimization entangles unrelated mechanisms so nothing can be attributed. Their fix is benchmark-disjoint evolution plus contrastive analysis of successful and failed trajectories on the same task, applied to decomposed harness modules. Given how many RSI harness papers have appeared in the past ten days, this is the methodological gate to read them through.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
D-SCAN (SIGIR 2026) found the standard guardrail returns high confidence on compromised output. Their alternative signal is document-level attention dynamics: during a poisoned generation, attention concentrates on the injected document and entropy collapses, versus dispersed...
arXiv 2608.06196 pits lexical+dense ranking against a graph encoding prerequisites, data flow and ordering across 117 realistic non-echoing queries. The ranker hits top-5 in 73.5% ±8.0 of cases; graph neighbours at matched token budget lose 11.2 points at p=0.0007. The mechani...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.