Fetching from the wire…
Skills2026-08-11 · source-backed
He et al. argue research agents run the same loop as a greybox fuzzer, and fuzzers make progress because coverage gives dense feedback on every execution that directs the next mutation. Scaling the proposer and ranking more samples post-hoc misses this entirely. Use the signal to pick the next intervention, not to sort finished runs, and keep your validation evidence protected from adaptive reuse.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Here's the experiment: a team of cooperating agents rebuilds SQLite in Rust from scratch, using only the 835-page manual. No source code. No test suites. No internet. Then it has to pass a held-out sqllogictest suite. It worked. Cursor published the research (Wilson Lin, July...
Across three models and two environments over a 24-turn horizon, 5x compression produced no statistically significant change in task completion (arXiv 2608.16370): but all six model/regime comparisons showed more retrieval calls, five significant after correction. GPT-5.5 comp...
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation rather than the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success where Meta-SecAlign, the previous state of the art, sit...
arXiv 2609.17394 audited 254 SWE-bench submissions across four splits without running a single model, just by analyzing the published per-instance results. The top two entries both resolve 396 of 500. The top ten agents share 285 successes and 51 failures, leaving 164 instance...
arXiv 2608.03999 holds the Qwen3.5 backbone (0.8B–27B), data, budget, and decoding fixed and swaps only the representation across seven tokenizations. Scaling the backbone 34x barely moves Frechet Music Distance; switching representation halves it. Their PMT stream (10ms timin...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.