Fetching from the wire…
Security2026-09-03 · source-backed
CVE-2026-84809 and CVE-2026-84811, both published September 2 at CVSS 7.1, hit Tencent's AI-Infra-Guard skill-scan and agentverus-scanner with the same flaw: each hardcodes __pycache__ directories and .pyc/.pyo/.pyd extensions into skip lists. An attacker ships benign Python source next to malicious compiled bytecode that executes on import, and agentverus-scanner returns a CERTIFIED verdict with high trust scores in both static and semantic modes while the payload runs. If your skill-install gate is either of these scanners, the gate does not see the file that runs. Delete compiled artifacts before scanning, or scan the archive rather than the extracted tree.
Each link below shares sources, entities, or timing with this story.
Simon Willison ran it and notes the jump from Hy3's 295B/21B/256K, two reasoning effort levels with 'high' as default and 'no_think' available, and a reasoning trace written in slightly truncated English that reads as the model deliberating, considering and rejecting details l...
The AngelSlim/Hy4-preview-GGUF repo offers Q4_K_M at 435.20 GiB (4.86 bpw) and STQ1_0 at 213.66 GiB (2.38 bpw), benchmarked at 204.56 t/s prefill and 20.47 t/s decode on 8x H20. STQ1_0 comes from llama.cpp PR #22836 and uses ternary weights with 3:4 forced sparsity at 1.3125 b...
Hy4 preview, released August 28: 770B total parameters, 49B active, over 1M token context, open-sourced and simultaneously on Tencent Cloud TokenHub and OpenRouter at $0.834 per million input tokens, $2.501 per million output, $0.042 per million cached. In Tencent's own evalua...
78 layers where layer one is dense FFN and the other 77 are MoE, each with 256 routed experts and 1 shared expert, top-8 routing per token, plus a native 10B MTP layer (0.7B activated) built in for speculative decoding. FP8 and base variants released together on August 28; the...
Released August 28 with 78 layers, 77 of them MoE with 256 routed plus one shared expert and top-8 routing, plus a native 10B MTP layer for speculative decoding (GitHub). The attention stack uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse index reus...
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.