EvolveScaler Generates Long Contexts Where Later Events Revoke Earlier Ones, and Six Models Score Under 10%
EvolveScaler defines information evolution as executable state machines before rendering it as natural language: human-authored specifications set state transitions, record validity, difficulty controls and answer logic, a strong LLM synthesizes a self-contained simulator per specification, and deterministic replay computes reference answers and atomic checklists. Instantiated with 117 task prototypes and 159 final-question operators across five difficulty levels spanning roughly 7 to 1,200 events per instance, it yields about 35,100 training examples and 585 validated evaluation instances. On the very_long tier the strongest model reaches 59.3% avg@5 while six models score below 10%, and training an internal A3B model on 6,000 examples gained 5.25 points average across all eight out-of-distribution benchmarks.
↳ Follow the thread