Fetching from the wire…
Research2026-09-15 · source-backed
Thirteen authors ran a generational genetic algorithm over specialized agents that separately handle mechanistic argument, assumption reconsideration, and evidence and testability assessment (arXiv 2609.15938). Evaluated against DepMap and Open Targets across 34 cancer types, it reached 0.171 DepMap selectivity against 0.115 for the strongest baseline, beating six baselines on both measures. Gains over single-pass generation held on held-out cancer types too, which is the part that matters: the evolutionary loop carried the improvement, not agent specialization alone.
Each link below shares sources, entities, or timing with this story.
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
The failure they target is specific and under-discussed: a cached error page or a negative price returns in the *expected schema* and gets consumed as fact, unlike a timeout the agent can see. Outcome Monitors check results against contracts mined from task-disjoint traces or...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
LangChoiceBench covers 28 projects across seven software areas where Python is a poor default, run against 25 LLMs. Python stays heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models show stronger bias. Analysis of 9,826 reason...
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
First defense modeling multi-turn indirect prompt injection as temporal causal takeover. Uses counterfactual re-executions at tool-return boundaries to detect when tool outputs steer agent behavior. Evaluated on AgentDojo across four task suites. Builder-ready pattern for tool...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.