Fetching from the wire…
Public story · 2026-09-11 · high
Three open Nemotron checkpoints scored 30 of 42 on 2026 olympiad problems using generate-verify-refine, and Nvidia posted the whole recipe.
Why now: Nvidia posted the paper, checkpoints, and benchmark together on September 11.
Three Nemotron 3 Ultra checkpoints scored 30 out of 42 on International Mathematical Olympiad-level problems, clearing the gold medal threshold. No formal prover, no tools, no internet access.
Gold-medal math scores like this usually come from closed labs that publish a number and nothing else. Nvidia published the full paper on arXiv along with math SFT and RL checkpoints on Hugging Face, the inference recipe in NeMo-Skills, and a new 200-problem olympiad benchmark. Anyone with the checkpoints can rerun the 30-of-42 result themselves.
The method is a generate-verify-refine loop. The checkpoints pass proposed solutions back and forth, checking and rewriting until one holds up, instead of one model producing a single answer. That's a different bet than a bigger single model on harder problems. It's also cheaper to test, since you can swap in your own checkpoints and rerun the loop end to end.
What the paper doesn't settle is how this generalizes past olympiad math. Competition problems have a specific shape: clean statements, verifiable answers, no ambiguity about what counts as a solution. Whether generate-verify-refine holds up on messier reasoning tasks, the kind without a clean checkable answer, is a question the benchmark can't answer on its own.
The checkpoints, the recipe, and the benchmark are all sitting there for anyone to run.
Each link below shares sources, entities, or timing with this story.
TechCrunch strings together three deals: Nvidia's reported $13 billion Hugging Face acquisition, its $6 billion Poolside arrangement, and Stripe's acquisition of OpenRouter for over $7 billion about two weeks before August 28. The thesis is acquirers hedging against frontier-l...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
$12,930,300,000. Not "about $13 billion." That exact figure is in Jensen Huang's September 3 post, and the precision is the first sign this is a signed agreement rather than the leak-stage reporting we got in late August. NVIDIA's blog frames the platform in the numbers that m...
The day's highest-scoring r/LocalLLaMA post points out the deal takes the llama.cpp and ggml copyright along with the team Hugging Face hired in February 2026, including Georgi Gerganov (r/LocalLLaMA). The top reply at 957 upvotes is "If it happens, we shall fork and move on....
Exactly matching this year's gold bar, plus a tie for first on MathArena AIME 2026 at 97.1%, and 35/42 on last year's IMO problems (SK Telecom). Xiaohongshu's dots-note-3.0 got a perfect 42 this year, so A.X K2 is at the threshold, not the frontier. It's the only model develop...
Announced July 27 with Microsoft, IBM, Red Hat, Palantir, CrowdStrike, Cloudflare, Databricks, Hugging Face, LangChain, Nous Research, Reflection AI, Thinking Machines Lab, SpaceXAI and the Linux Foundation. Huang's framing is pointed: during the Hugging Face incident "closed...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.