Effective Distillation to Hybrid xLSTM Architectures
arXiv 2603.15590·medium signal
Demonstrates successful distillation from quadratic attention-based LLMs into hybrid xLSTM architectures combining sub-quadratic linear attention with selective state retention, preserving long-context quality where prior linear distillation attempts degraded. Opens a viable path to deploying frontier-quality models on memory-constrained hardware without the quality cliff seen in earlier SSM distillation work.