OneDayAgent posts 0.821 on AgentIF-OneDay and holds up across five backends from three model families
arXiv 2608.05013 (Aug 4) argues the open question is not whether individual long-horizon failure modes can be fixed but whether one harness can manage goal drift, state loss and context overflow jointly. OneDayAgent decomposes an open-ended request into bounded subtasks, maintains execution memory under context pressure, and verifies then repairs the final deliverable. On AgentIF-OneDay's 104 cross-environment multimodal tasks it reaches a state-of-the-art 0.821 overall with a GLM-5.2 backend, and the same harness runs unchanged across five backend LLMs from three families — evidence that harness engineering generalizes even as different models produce distinct execution styles under the identical workflow.
Source
↳ Follow the thread