Fetching from the wire…
Research2026-09-20 · source-backed
arXiv 2609.19636 makes a point I hadn't considered: an agent in a closed loop writes its own inputs, so SFT and RL checkpoints get scored from different states even on identical tasks, and restricting comparison to shared states selects on an outcome, which flips the sign of the effect in their data (arXiv). Their protocol clones a state one checkpoint reached and hands it to another with no retraining, splitting an endpoint gain into REACH and SOLVE. Across two benchmarks and two independent pipelines the interaction is positive in all five conditions, and on ALFWorld the SFT solver never succeeds where the RL solver fails.
Each link below shares sources, entities, or timing with this story.
SkillPyramid extends Voyager-style libraries with a hierarchical topology plus a self-evolution loop, raising average reward 38% and cutting execution steps 27.7% across ALFWorld, WebShop, and ScienceWorld. The takeaway: don't append every successful trajectory as a new flat s...
arXiv 2609.20784 attacks two failures in self on-policy distillation for multi-turn agents: privileged information doesn't make a teacher reliable, and teacher supervision helps only at certain stages (arXiv). The student drops the teacher on its own once their discrepancy sto...
The method turns tool use from a hardcoded prompt into a learned runtime behavior, then applies cost-aware RL teaching the agent when reading external state is worth the token budget. Qwen3-8B reaches a 96.9% average success rate against SkillOS at 80.2% and SkillRL at 89.9%,...
SkillForge (arXiv 2608.24747) notes that skill-extraction approaches like SkillRL never verify whether a stored skill still works against the current environment, so the bank grows monotonically while quality rots. It makes skill usage explicit during interaction so RL optimiz...
Across three models and two environments over a 24-turn horizon, 5x compression produced no statistically significant change in task completion (arXiv 2608.16370): but all six model/regime comparisons showed more retrieval calls, five significant after correction. GPT-5.5 comp...
The system (arXiv 2609.10712) uses no formal prover, no tools and no internet access. Three Nemotron 3 Ultra checkpoints run a generate-verify-refine loop, and together they scored 30 of 42 at IMO 2026, the gold threshold. NVIDIA posted the math SFT and RL checkpoints on Huggi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.