Research
SCALE-QA Tests Memory in Flat Unsegmented Threads; Episode Reconstruction Beats Long Context by 5.6-17.6 Points
Feng et al. note that existing memory benchmarks leak session or topic boundaries, and build SCALE-QA: 3,000 audited four-way multiple-choice questions across 10 domains in a single flat mixed-topic thread, where the system must infer which earlier episode makes a later task decision valid. They evaluate all 3,000 through 128k context plus a stratified 400-question diagnostic at 1M. Their TSIM method segments the turn stream into coherent episodes and indexes them through a hierarchical multi-view memory stack, beating the strongest baseline by 5.6 to 17.6 accuracy points on every one of three open and proprietary backends, which says long context alone does not solve episode integrity.
↳ Follow the thread