Language Models Need Sleep: Sleep-Like Consolidation Mechanism Converts KV Cache to Persistent Fast Weights
arXiv·high signal
Researchers propose a sleep-like consolidation mechanism where transformers periodically convert recent context into persistent fast weights in state-space model (SSM) blocks via Hebbian learning, then clear the KV cache. This directly addresses the scaling problem of attention with context length for long-horizon agent tasks, offering a biologically-inspired alternative to infinite context windows.