Fetching from the wire…
Agents2026-09-19 · source-backed
arXiv 2609.18366 names the hole: a Proposer repeatedly editing prompts, memory, retrieval, tools and control code against a released benchmark can find a benchmark-wide protocol shortcut, and holdout varies semantic tasks while leaving the protocol fixed. CHASE recasts harness evolution as constraint generation over validity-preserving counterfactuals, where after each Proposer update a Challenger searches for an executable protocol transformation that destroys the gain. arXiv If you're tuning a harness against SWE-style benchmarks, holding out tasks isn't holding out anything.
Each link below shares sources, entities, or timing with this story.
Per-trace debugging experience in agentic code translation doesn't accumulate into reusable knowledge (arXiv 2609.15381). TRAIL runs two agents adversarially, a Translator and a Challenger, distilling trace-specific experience into generalizable translation rules. Against the...
The authors define goal-directed execution as four repeated behaviors: selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, verifying completion against the environment. Post-training Qwen3.5-122B-A10B on 363 long-horizon multi-to...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
SWEADV built 750 adversarial issue descriptions from 150 SWE-bench Verified tasks, five per task across command execution, deserialization, path traversal, DoS and weak hashing (arXiv 2609.15963). Across mini_swe agents on GPT-5-Mini, MiniMax-M2.5 and DeepSeek-R, adversarial i...
RealSWE builds a six-category information taxonomy and four style dimensions, then compares real prompts from SWE-chat against SWE-bench Verified and Pro. Also: 87% of real prompts are casually written, against 94% of benchmark problems written formally. They release 381 multi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.