Retrieving Reusable Reasoning Memories at Prefill Recovers 21-29 Points Lost to Chain-of-Draft Compression
Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning (arXiv 2608.21265, Aug 21) formalizes a Context-Generation Substitution Law where explicit reasoning context in the prompt substitutes for part of decode-time generation. The training-free method builds reusable memories from historical traces that summarize reasoning patterns, key constraints and critical operations rather than storing raw demonstrations, then retrieves them as prefill-side scaffolds. Layered on prompt-based Chain-of-Draft compression it gained 21.4, 28.0, 29.5 and 6.61 accuracy points on GSM8K, MATH, BBH and MMLU-Sci while delivering a 1.14-1.49x latency speedup over standard CoT, and is compatible with token-level, trace-level and inference-state compression.
↳ Follow the thread