News
Meta's EvoHarness-RL Gets Qwen3-8B to 96.9% on ALFWorld, Within Two Points of Claude Opus 4.5
arXiv 2608.05446 ('EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents') turns tool use from a hardcoded prompt into a learned runtime behavior, then applies cost-aware RL that teaches the agent when reading external state is worth the token budget. On ALFWorld, Qwen3-8B hits a 96.9% average success rate, beating SkillOS (80.2%) and SkillRL (89.9%); the same harness pushes Claude Opus 4.5 to 98.5%. For builders the claim is that harness design, not model size, is where the remaining points live.
Source
↳ Follow the thread