Reddit
Qwengram-0.8B grafts Qwen3.8 Flash-Next's 51B-parameter n-gram memory onto Qwen3.5-0.8B for 5.05% lower perplexity
A r/LocalLLaMA user froze both the Qwen3.5-0.8B backbone and the roughly 51B-parameter PLE n-gram memory from Qwen3.8-Flash-Next. They trained only a small R=1 reader at decoder layers 3 and 9 with a token-level gate, using 15M tokens on free Kaggle GPUs. Validation perplexity fell from 18.28 to 17.35, and the real memory beat both random and permuted memory controls. A 20M-token reader regressed on math, and strong fixed injection hurt LAMBADA. The GGUFs and a llama.cpp inference path are public. The post reached 353 upvotes and 96 comments.
↳ Follow the thread