Fetching from the wire…
Research2026-07-29 · source-backed
Bits and Memories measures verbatim reproduction of training data rather than membership inference, using Pythia models with known-memorized sequence sets across five precision levels down to four bits. Verbatim memorization drops faster than perplexity at every precision and size under two unrelated quantization algorithms. But at the largest model studied, four-bit quantization still reproduces most memorized sequences while giving up a few percent of capability, and the surviving memorized fraction grows with model size.
Each link below shares sources, entities, or timing with this story.
GSQ uses Gumbel-Softmax sampling to match the accuracy of QTIP and AQLM while keeping the deployment simplicity of GPTQ/AWQ. If you're quantizing models for local inference, this eliminates the accuracy-vs-complexity tradeoff.
The argument is that a structurally compressed model's bfloat16 checkpoint is itself only a distillation-recovered approximation, so training the 4-bit student against it inherits that error. Distilling directly from the original model reaches a comparable peak about 7x faster...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
First rigorous statistical comparison of AIFS against IFS on operational ensemble data (arXiv). The energy number is the headline. AI forecasting is now operationally competitive, not just a research curiosity, and the cost structure is the reason it'll win deployment.
arXiv 2607.29167 describes the mechanism precisely: when an agent consolidates an external observation into long-term memory, the rewrite preserves the action trigger while erasing the low-trust source. The injected instruction resurfaces later looking like user history. Memor...
arXiv 2607.27919 argues long-term memory should be a separately scalable parametric module rather than entangled with reasoning in one weight set, backed by distributed Faiss indexing and sparse batch-wise loading of kNN distributions. The 410M+6.9B pairing lifts the 17-benchm...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.