ModularRSI argues most self-improving harness results are benchmark overfitting, and splits the harness to prove it
arXiv 2609.14857 (submitted 14 Sep) names three failure modes in recursive harness self-improvement: evolving on the evaluation benchmark makes reusable gains indistinguishable from benchmark-specific adaptation, single-trajectory updates conflate systematic harness defects with one-off reasoning slips, and whole-harness optimization entangles unrelated mechanisms so nothing can be attributed or validated. Their fix is benchmark-disjoint evolution plus contrastive analysis of successful versus failed trajectories on the same task, aggregated across tasks to surface recurring behavioral deficiencies, then applied to decomposed harness modules rather than the monolith. Given how many RSI harness papers have shipped in the past ten days, this is the methodological critique that should gate how you read them.
Source
↳ Follow the thread