An open harness hits 82.6% on SWE-bench Verified by letting runtime evidence rewrite the plan, not the model policy
openJiuwen (arXiv 2608.27969, Aug 28) separates two harness problems it names Structural Composability and Runtime Adaptivity: developers compose capabilities across single agents, delegated sub-agents and a Swarm Flow over one shared execution substrate, while the framework adapts context, feedback and task control from live evidence (semantic diagnostics, execution outcomes, task progress) under a fixed model policy. It reports 82.6% on SWE-bench Verified and 87.19% on Terminal-Bench 2.1, 3.4 and 3.39 points above the strongest official-leaderboard points it selected. The org's repos are real and active: openJiuwen-ai/jiuwenswarm is at 6,309 stars and agent-core at 420, both pushed Aug 31, so this is a shipped harness rather than a paper artifact.
↳ Follow the thread