Fetching from the wire…
OSS2026-09-01 · source-backed
The Apache-2.0 microVM sandbox for AI agents built on RustVMM and KVM added cross-node pause/resume and control-plane separation on August 28, claiming average cold start under 60ms, sub-150ms delivery under high concurrency, and under 5MB memory overhead per sandbox, with pause and resume at hundred-millisecond granularity (GitHub). 11.6k stars since the April v0.1.0. A credible open alternative under the hosted agent-sandbox services, which is a category that has been quietly expensive.
Each link below shares sources, entities, or timing with this story.
Released August 28 with 78 layers, 77 of them MoE with 256 routed plus one shared expert and top-8 routing, plus a native 10B MTP layer for speculative decoding (GitHub). The attention stack uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse index reus...
Hy4 preview, released August 28: 770B total parameters, 49B active, over 1M token context, open-sourced and simultaneously on Tencent Cloud TokenHub and OpenRouter at $0.834 per million input tokens, $2.501 per million output, $0.042 per million cached. In Tencent's own evalua...
78 layers where layer one is dense FFN and the other 77 are MoE, each with 256 routed experts and 1 shared expert, top-8 routing per token, plus a native 10B MTP layer (0.7B activated) built in for speculative decoding. FP8 and base variants released together on August 28; the...
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
Firecracker + CoW mmap forking spawns KVM-isolated sandboxes in 0.79ms p50 with 265KB memory per instance — versus E2B's ~150ms and ~128MB. At this scale, sandboxed code execution becomes viable inside interactive agent loops where latency directly impacts UX. GitHub ---
Simon Willison ran it and notes the jump from Hy3's 295B/21B/256K, two reasoning effort levels with 'high' as default and 'no_think' available, and a reasoning trace written in slightly truncated English that reads as the model deliberating, considering and rejecting details l...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.