Fetching from the wire…
Research2026-08-27 · source-backed
Existing memory benchmarks leak session or topic boundaries. SCALE-QA gives 3,000 audited four-way questions across 10 domains in a single flat mixed-topic thread, where the system must infer which earlier episode makes a later decision valid (arXiv 2608.25655). All 3,000 run through 128k context plus a stratified 400-question diagnostic at 1M. Their TSIM method segments the turn stream into coherent episodes and indexes them through a hierarchical multi-view memory stack, beating the strongest baseline on all three open and proprietary backends. Long context alone does not solve episode integrity.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (baseline, beat, benchmark, context, memory); pushes against this story (versus).
Shared entity: Scale / Same source domain / Shared topic / Earlier coverage
Both cover Scale; reported by the same outlet (arxiv.org); overlapping topics (context, memory).
Shared entity: Existing / Same source domain / Earlier coverage / Tension
Both cover Existing; reported by the same outlet (arxiv.org); earlier Existing coverage from 2026-07-21.
Both cover Existing; reported by the same outlet (arxiv.org); earlier Existing coverage from 2026-07-17.
Shared entity: Scale / Shared topic / Earlier coverage / Downstream implication
Both cover Scale; overlapping topics (beat, beating); earlier Scale coverage from 2026-03-18.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (alone, beat, benchmark); pushes against this story (versus).
Reported by the same outlet (arxiv.org); overlapping topics (baseline, beat, benchmark); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (alone, benchmark, memory); pushes against this story (but).