Fetching from the wire…
Models2026-09-13 · source-backed
smolbenchmark inverts the leaderboard: 13 model families that fit in 8GB, ranked by decode speed, tokens per joule and heat, measured on tablets, phones, Macs, Jetsons and Raspberry Pis. About 1,000 configs are live for the Jetson Orin Nano Super 8GB capturing tok/s, tok/J, inter-token latency, power, thermals and battery, with Pi, phone and Mac mini runs pending. The top comment is the necessary caveat: without a reproducibility block naming quant, backend version, context length, warmup count and plugged-in state, a cold first run flatters a model and a long run reverses the ranking through throttling.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
The post argues Ollama went over a year without crediting llama.cpp in its README while a license-compliance issue sat 400+ days without a maintainer response, that llama.cpp runs 1.8x faster (161 against 89 tokens/second) with 30-50% CPU gaps, and that the mid-2025 move to a...
Diffusion LMs decode many tokens per step but pay to interact with all suffix tokens every step, and existing fixes just keep a local window while re-initializing suffix tokens identically each timestep (arXiv 2608.23167). This method splits the suffix into local, middle and t...
July MCP roundups documented Mid-Session Tool Injection against WebMCP agents, using threshold poisoning and fabricated diagnostic events to swap or re-scope tools after a session is already established. The uncomfortable implication: a context provider you trusted at connect...
The ds4 maintainer found single-stream decode running at 59% of the machine's measured memory bandwidth, and the bottleneck wasn't the big weight-streaming kernels but dozens of small kernels between them each paying dispatch latency. Fusing that work into larger dispatches go...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.