Sources
RSIAgent lets Kimi-K3 and GLM-5.3 beat GPT-6 on OSWorld-v2 with frozen memory and no parameter updates
arXiv 2609.15364, submitted 2026-09-14, is a training-free multi-agent framework that coordinates curriculum, actor and verifier agents to explore an unfamiliar environment and retain what it learns, including reusable causal relationships between actions, conditions and consequences. It runs broad parallel self-exploration to map environment structure, then deep focused exploration for hard cases, hidden constraints and boundary conditions. The resulting memory is frozen and reused directly downstream, and on OSWorld-v2 and Agent's Last Exam it lifts open-weight Kimi-K3 and GLM-5.3 past closed frontier models including GPT-6, which makes it a cheap retrofit rather than a training project.
↳ Follow the thread