An Agent Meant to Run for Weeks Needs Levels, Ticks and Escalation, All in the Harness
arXiv 2609.19519 (submitted 17 Sep 2026) argues a long-horizon agent must run continually without forgetting before it can learn continually, and that this capability lives in the harness rather than the model. The authors derive seven bottlenecks from tasks that outlive any context window, process or human attention interval, and answer them with three parts: levels indexed by time scale where each keeps a bounded file summarizing the level below, a clocked tick as the unit of autonomous action, and cascaded intelligence where work escalates to a more capable model only after failing review. They report on a ten-part deployment rather than a benchmark score, which makes this a design reference more than an evaluation.
↳ Follow the thread