Fetching from the wire…
Agents2026-09-01 · source-backed
Production skills are directory bundles where only the root loads at activation and references, schemas, scripts and nested subskills load on demand, so compressing the root misses most of the cost while flattening destroys the progressive-loading boundaries (arXiv 2608.30785). Compressing across files, dropping content a reference already gets from the root, removes 38% of bundle tokens and 10.4% of end-to-end per-run tokens with no quality loss on a production content-moderation skill. An unprotected 71% compression setting loses up to 26 accuracy points to one-sided false positives, so the guardrail is doing real work.
Each link below shares sources, entities, or timing with this story.
Paritok-4B (arXiv 2608.24188) is a LoRA on Qwen3-4B distilled from a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories. It's extractive rather than paraphrasing, with 96.0% of emitted identifiers, paths and numbers already present in its input, and intent-conditione...
One engineer, one week, $1,100 in AI costs. Production builds on 33-route app: 1.67s vs 7.38s for Next.js 16 + Turbopack. Client bundles: 72.9 KB vs 168.9 KB. ~94% API coverage, zero source changes required. Cloudflare Blog ---
Speculative decoding usually splits within one machine. SPADE splits across the edge and cloud boundary: a compact draft model on the edge proposes tokens, a large cloud verifier validates them in parallel, and only rejections trigger a cloud correction (arXiv). No retraining,...
arXiv 2606.26083 tested four leading real-time voice systems, including OpenAI's realtime stack, and found they transcribe lexical content well but largely ignore tone, emphasis, and emotion. The same week, SpeechEQ introduced an "emotional intelligence quotient" benchmark sco...
Steve Yegge reframed agentic coding this week and I think he's right, which is annoying because it means the thing I just got good at is already the wrong unit of work. In his latest piece, Yegge lays out a six-wave chart of coding agents and plants a flag: 2026 is the year of...
MiniMax shipped M2.7 on March 18 and made a claim nobody else has made with receipts: the model participated in its own R&D cycle. Not "we used AI to help train it" marketing. MiniMax says M2.7 autonomously handled 30–50% of the development workflow — reading logs, debugging f...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.