Fetching from the wire…
Models2026-09-24 · source-backed
Built on Qwen3.5-9B-Base, it renders context as images at 5x, 10x or 15x compression and calls learned tools to decompress only the pages it needs. Apple reports accuracy on par with full text at 4.3x effective compression, and says it beats retrieval and compression baselines up to 10.1x on seven QA benchmarks (Hugging Face). Apple's research model license rules out commercial use, so this is a technique to study rather than a model to deploy.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Compaction is where long sessions go to die. The model summarizes what happened, the summary drops the exact string you needed three hours later, and you don't find out until the agent confidently references a file path that never existed. It's the biggest source of silent con...
Multiverse Computing released HyperNova 60B on Hugging Face for free — a 50% compressed version of OpenAI's gpt-oss-120B using quantum-inspired CompactifAI compression. Memory drops from 61GB to 32GB (fits single consumer GPU) while showing 5x improvement on Tau2-Bench and 2x...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
AutoModelForCausalLM.from_pretrained("unsloth/Qwen3.5-4B-GGUF", gguf_file="Qwen3.5-4B-Q4_K_M.gguf") and everything after is the normal transformers API, with BF16, Q6_K, Q5_K_M and Q4_K_M supported. On a MacBook Pro M2 Max throughput came close to llama.cpp across three checkp...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.