Fetching from the wire…
Models2026-09-01 · source-backed
It turns a course brief into finished slides or a self-contained interactive HTML page in one pass (arXiv 2608.30968). Across 220k production requests the median slide takes 17 seconds and an interactive page 59. A hybrid rule-plus-VLM reward drives GRPO, hardened after the team caught a reward-hacking episode producing visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and CogEvol-4B is Apache 2.0. Publishing the reward-hacking incident instead of burying it is the part I want more papers to copy.
Each link below shares sources, entities, or timing with this story.
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasonin...
RobertGolds1/Gradient, created August 23 and at 423 stars in about a day, Apache-2.0, built on OpenPipe ART. It ships a Research Environment containing a reproducible company workspace of emails, contracts, policies, meeting notes and customer records with evidence deliberatel...
Two specialized RL agents — Memory Manager (ADD/UPDATE/DELETE operations) and Answer Agent — fine-tuned with PPO and GRPO. With only 152 training QA pairs, outperforms baselines across three benchmarks. Directly applicable to persistent agent memory systems. arXiv 2508.19828
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 t...
alex000kim/nanoRL (MIT) scales from CartPole on a laptop to async distributed training on GPU clusters with vLLM rollout workers, across 7 files. The author optimizes explicitly for readability and forkability, excluding Megatron-scale parallelism and multi-tenant scheduling....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.