Fetching from the wire…
Research2026-08-21 · source-backed
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 to 2,811 tokens, and routing transferred to easier benchmarks for 76% savings with no retraining. arXiv The per-mode token cap is the load-bearing design choice. Without it the modes collapse into one.
Each link below shares sources, entities, or timing with this story.
Unsloth released GRPO / Shared entity: GRPO / Earlier coverage / Tension
Linked by a graph relationship (Unsloth released GRPO); both cover GRPO; earlier GRPO coverage from 2026-06-20.
Unsloth released GRPO
Linked by a graph relationship (Unsloth released GRPO).
Unsloth released GRPO / Shared topic / Tension
Linked by a graph relationship (Unsloth released GRPO); overlapping topics (model, token); pushes against this story (versus).
Shared entity: Accuracy / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Accuracy; reported by the same outlet (arxiv.org); overlapping topics (accuracy, against, length).
Both cover Accuracy; reported by the same outlet (arxiv.org); overlapping topics (accuracy, against, benchmark).
Unsloth released GRPO / Shared topic
Linked by a graph relationship (Unsloth released GRPO); overlapping topics (benchmark, length, model).
Shared entity: GRPO / Same source domain / Shared topic / Earlier coverage
Both cover GRPO; reported by the same outlet (arxiv.org); overlapping topics (accuracy, benchmark, model, token).
GRPO competes with PPO / Same source domain / Shared topic / Tension
Linked by a graph relationship (GRPO competes with PPO); reported by the same outlet (arxiv.org); overlapping topics (against, baseline, benchmark).