Fetching from the wire…
Public story · 2026-07-27 · high
The system beats open baselines of similar size on nearly every benchmark tested using a proxy that captures live model calls as training data.
Why now: OpenForgeRL posted July 23, covered here as part of the July 27 briefing.
Microsoft's OpenForgeRL trains AI agents directly inside the tool harnesses they run in, per a paper posted July 23.
A real gap, closed. Harness agents like Claude Code and Codex handle multi-turn reasoning and tool calls, but training them end-to-end had no matching open infrastructure. The resulting system beats open baselines of similar size on nearly every benchmark tested, scoring 31.7 pass@1 on ClawEval.
The training method is a lightweight proxy sitting between the agent and the model, recording every model call as training data.
Kubernetes handles the rollout orchestration. That decouples training from inference, so the agent learns from what it actually does inside the harness.
OpenForgeRL also posts 55.9 pass@3 on ClawEval, 33.7 on QwenClawBench, and 37.7 on OSWorld-Verified.
On the web-agent tests, it scores 63.0 on Online-Mind2Web and 72.3 on WebVoyager, per the paper.
If independent labs can't reproduce these benchmark gaps outside Microsoft's own numbers, OpenForgeRL is a solid result, not a new training paradigm for agents. Worth watching whether teams building on Claude Code or Codex start training on live tool-call traces from their own harnesses.
Each link below shares sources, entities, or timing with this story.
Microsoft released OpenClaw / Shared entities / Same source / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft released OpenClaw); both cover Claude Code, Codex, OpenForgeRL, Verified; cite the same source (Microsoft's July 23 release).
Microsoft criticizes Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, Codex, July, Microsoft; reported by the same outlet (arxiv.org).
Microsoft criticizes Claude Code / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, Codex, July; overlapping topics (agent, benchmark, claude, code, codex).
Microsoft criticizes Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, July, Microsoft; reported by the same outlet (arxiv.org).
Microsoft criticizes Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, Codex, July; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, Codex, Microsoft; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, Codex; reported by the same outlet (arxiv.org).
Microsoft criticizes Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Claude Code, Codex, July; overlapping topics (agent, call, claude, code, codex).