Research
LongWoF-Bench Finds Verifier-Confirmed Experience Beats Distilled Summaries by 8.7-15.5 Points Across Seven Models
LongWoF-Bench comprises 778 machine-verifiable tasks spanning code generation, agent-environment synthesis, mathematical reasoning, and rule following, built to test whether execution experience can be externalized and reused. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Genes beat Skill representations across all seven evaluated models by 8.7-15.5 points, while reference-distilled Genes showed no such advantage, indicating provenance rather than compactness drives the gain. For Claude Opus, Gene reuse completed 39 more tasks than Skill while cutting solve-time token consumption by 9.9%.
↳ Follow the thread