Fetching from the wire…
Research2026-07-28 · source-backed
arXiv 2607.24720 builds a unified controlled multi-turn environment to isolate where planning ability comes from, across pre-training acquisition, post-training shaping via GRPO and on-policy distillation, and integration through multi-teacher on-policy distillation. Findings: pre-training data containing explicit world models improves generalization, suboptimal trajectories degrade performance specifically over long horizons, and cross-environment learning works when planning patterns are compatible but interferes when they conflict. For anyone curating agent training data, "fewer trajectories beats worse trajectories" is the actionable one. (arXiv 2607.24720)
Each link below shares sources, entities, or timing with this story.
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasonin...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
It turns a course brief into finished slides or a self-contained interactive HTML page in one pass (arXiv 2608.30968). Across 220k production requests the median slide takes 17 seconds and an interactive page 59. A hybrid rule-plus-VLM reward drives GRPO, hardened after the te...
Two specialized RL agents — Memory Manager (ADD/UPDATE/DELETE operations) and Answer Agent — fine-tuned with PPO and GRPO. With only 152 training QA pairs, outperforms baselines across three benchmarks. Directly applicable to persistent agent memory systems. arXiv 2508.19828
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
CodeGrep measures a 30B OpenHands agent averaging 23 rounds and 631K tokens per resolved SWE-Bench Verified issue, much of it grep, glob and view_file. A 14B retrieval agent trained end-to-end with GRPO raises resolve rate to 27.0% from 25.8% while cutting 15% of rounds and 19...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.