Human-Like Working Memory Constraints Improve Transformer Learning Under Data Scarcity
arXiv·medium signal
Researchers integrate cognitively inspired working memory constraints — fixed-width attention windows and temporal decay mechanisms — into the Transformer architecture, finding these constraints act as beneficial inductive biases under data-scarce conditions. Rather than viewing limited attention as a deficiency, the paper shows that constraining attention mimics human cognitive bottlenecks that force efficient information compression. Relevant to practitioners fine-tuning on small datasets where standard Transformers overfit.