Fetching from the wire…
Public story · 2026-08-19 · high
Agent Lightning's training scripts are public, and the whole run used 3,500 lines of code.
Why now: The paper posted to arXiv and is covered here as of August 19.
A 9B model jumped from 41.8% to 56.4% on SWE-bench Verified after reinforcement learning on 6,000 examples, per the Agent Lightning paper on arXiv. That's a 14.6 point gain, not from a frontier lab's flagship.
The number that matters isn't 56.4%. It's 6,000. That's a dataset size a small team can build without a research budget. The paper ships with the full training and workflow scripts, not just results. Qwen3 is the base model, and the whole setup runs on roughly 3,500 lines of training code.
This is the difference between agentic RL as a lab-only capability and agentic RL as something a small team can run. A 14.6 point gain from 3,500 lines of code and an open-weights base is a large jump for a modest setup.
The evidence here is one paper's reported numbers on one benchmark with one base model. It doesn't say what the training run cost in compute, or whether the gains transfer to a different base model without the same tuning effort. Whether anyone outside the original team reproduces this on a different codebase is the thing worth watching next.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source / Shared topic / Downstream implication
Both cover Qwen3, SWE, Verified; cite the same source (arXiv 2608.17528); overlapping topics (agent, code, exampl, point, training).
Shared entities / Shared topic / Earlier coverage / Tension
Both cover Qwen3, SWE, Verified; overlapping topics (code, cost); earlier Qwen3 coverage from 2026-04-23.
Both cover Qwen3, SWE, Verified; overlapping topics (agentic, beat); earlier Qwen3 coverage from 2026-03-16.
Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Both cover SWE, Task; reported by the same outlet (arxiv.org); overlapping topics (agent, agentic).
Both cover Qwen3, SWE; reported by the same outlet (arxiv.org); overlapping topics (agentic, code).
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (agent, code, point).
Both cover SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (agent, code, training).
Both cover SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (agent, code, cost).