Fetching from the wire…
Infra2026-08-09 · source-backed
Released August 5, Qdrant 1.19 promotes TurboQuant from a secondary quantization layer to a primary storage datatype: Turbo4 keeps only the 4-bit representation, dropping from 36 bits per coordinate to 4. Ninefold reduction, plus fewer per-operation disk reads and writes. The post is unusually candid about the cost: without a full-precision copy Qdrant can't rescore top candidates, so Turbo4 is for disk-bound deployments while TurboQuant-over-full-precision stays correct when recall is the priority. Gain is proportionally larger for ColBERT-style multi-vector collections. Same release adds per-component memory tiers (cold/cached/pinned), prefix matching in keyword filters, per-query IDF for sparse search, deterministic sliced scroll, and a global quota API.
Each link below shares sources, entities, or timing with this story.
Released August 4 under Apache 2.0, reframing moderation as policy-adaptive question answering: write your rule in plain language, get a calibrated safety score from a single token, no retraining, one interface for text and images. Mistral claims it matches open guard models u...
It exposes memory, resources and skills as one virtual filesystem under a viking:// protocol, so an agent navigates its context with ls, tree and find rather than opaque similarity queries, with content tiered into L0 abstract, L1 overview and L2 detail loaded on demand. v0.4....
Headroom compresses tool outputs, logs, RAG chunks, and files before they ever reach the model. It deploys as a library, a proxy server, or an MCP server, and the benchmarks are blunt: 92% token reduction on code search (17,765 down to 1,408) and SRE debugging (65,694 down to...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Google Research and Synaptics announced a limited-edition Coral Dev Board powered by the Astra SL2610 SoC with the industry's first implementation of the Coral NPU — a 1 TOPS, open-source, RISC-V-based neural processing unit. Ships pre-configured with Gemma 3 270M for immediat...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.