Fetching from the wire…
Public story · 2026-09-12 · high
Six and a half minutes on one H100 gained 30.5 points on MATH-500, but the same recipe barely moved a stronger model.
Why now: The write-up posted with a reproduction notebook and public dataset as of September 12.
A 100-puzzle fine-tune pushed Qwen 3 4B Base from 54% to 84.6% on MATH-500. The run took six and a half minutes on one H100, using zebra logic puzzles as training data, a 30.5-point jump. The method, called PCSS, is described in a Hugging Face post.
Anyone fine-tuning small open models on a limited budget should look at the cost profile here. One GPU and 100 examples produced a 30-point swing on a standard math benchmark, a fraction of the data and compute most reported fine-tuning runs use. The same run also gained 2.57 points on AIME 2025, smaller but real on a harder benchmark.
PCSS is a per-example calibrated sigmoid scaler derived from KTO. It reduces to a standard supervised fine-tuning gradient, multiplied by a scaler that decays as the model masters a given example. In plain terms, it backs off once the model already gets something right, so training effort concentrates on what it still gets wrong.
The gains shrink as the base model gets stronger, and the author says so plainly. Granite 4.1 3B gained 11.1 points under the same recipe. Qwen 3.5 9B, a larger and stronger model, gained only 3.1 points. The more capable the base model, the less the zebra puzzles have to teach it.
The post includes a reproduction notebook and a public dataset, so the 84.6% figure can be checked rather than taken on faith. It doesn't say whether PCSS holds up on tasks further from logic puzzles. It also doesn't say whether the MATH-500 and AIME gains carry over to messier, real-world reasoning tasks.
Each link below shares sources, entities, or timing with this story.
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
The Hugging Face page is marked "Upcoming release" with no model card, license, architecture details, context length or benchmarks, after Alibaba promised both Qwen3.8-Max and the 27B weights for the week of August 10. A ModelScope countdown pointed at August 15. Unsloth signa...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Published August 25, it's IBM's first family of dense decoder-only reasoning models, with the 30B flagship claiming state-of-the-art resolve rates on SWE-Bench Pro and Terminal-Bench. The 8B and 30B went through an agentic training curriculum on real sandboxes for software eng...
The system (arXiv 2609.10712) uses no formal prover, no tools and no internet access. Three Nemotron 3 Ultra checkpoints run a generate-verify-refine loop, and together they scored 30 of 42 at IMO 2026, the gold threshold. NVIDIA posted the math SFT and RL checkpoints on Huggi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.