Fetching from the wire…
Public story · 2026-07-31 · high
The 35B model beats GPT-5.5 plus Codex and nears the 2.8-trillion-parameter Kimi K3, with its weights released publicly.
Why now: The result appears in the July 31 briefing, and its public weights let anyone rerun the same 4090 setup to check the score.
Frontis-MA1, a 35-billion-parameter agent, scored 71.21% on MLE-Bench Lite's Medal Average on a single RTX 4090 capped at 12GB of VRAM, per the paper posted to arXiv. That beats a GPT-5.5-plus-Codex combination and closes in on Kimi K3, a 2.8-trillion-parameter model, on the same benchmark.
That matters most for teams without a frontier lab's compute budget. It suggests trained search can close a scaling gap that money alone was assumed to close.
Composing the agent's four trained operators, Draft, Improve, Debug, and Crossover, into search lifted the benchmark's Medal Average from 39.39% to 60.61%. That's under the same 12-hour, single-GPU budget as the final score. Reaching 71.21% took one more layer, an additional search process the paper calls OpenMLE-Evo-Max.
The paper's authors trained each operator separately using execution-grounded supervised fine-tuning and reinforcement learning, on a dataset called OpenMLE. That means the model gets graded on whether its code actually runs and scores, not just whether it looks right. The authors release both the weights and the full stack, making the setup reproducible outside the lab that built it.
The paper doesn't say what Sol is beyond another system Frontis-MA1 still trails. It also doesn't break down how much of the 10.6-point jump from OpenMLE-Evo-Max comes from extra search time versus the operators themselves.
Search beat scale here. A 35B model on one capped consumer GPU outscored a named GPT-5.5-plus-Codex stack and nearly caught a model eighty times its size. The number to watch is how much of Kimi K3's remaining lead disappears if someone runs OpenMLE-Evo-Max against it directly.
Each link below shares sources, entities, or timing with this story.
Kimi K3 competes with OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover Codex, GPT; reported by the same outlet (arxiv.org).
Kimi K3 competes with OpenAI / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover Codex, GPT, Under; reported by the same outlet (arxiv.org).
Claude Code uses Kimi K3 / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Codex, GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Codex, GPT; reported by the same outlet (arxiv.org).
Claude Code uses Kimi K3 / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Codex, Kimi K3; overlapping topics (agent, codex).
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Codex, GPT; overlapping topics (agent, codex).
Kimi K3 competes with OpenAI / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover Codex, GPT; overlapping topics (codex, debug).
Claude Code uses Kimi K3 / Shared entity: Codex / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Kimi K3); both cover Codex; reported by the same outlet (arxiv.org).