Fetching from the wire…
Public story · 2026-09-08 · high
The engine needs no retraining, runs under an Apache 2.0 license, and any existing MiniMax-H3 LoRA plugs straight in.
Why now: NVIDIA published the Sol-H3 benchmarks under its Sana engine line, giving MiniMax-H3 users the first documented speedup numbers for the 8x B300 configuration.
NVIDIA and MiniMax's new engine turns a five-second AI video clip into 1.653 seconds of render time. Sol-H3 produces 1344x768 video with stereo audio on an 8x B300 Blackwell system, using four diffusion transformer forwards.
That's an 11.04x speedup over the base H3 model. Base H3 needs 18.25 seconds and 50 scheduler points for the same clip. The speedup rises to 15.05x on 15-second outputs.
The speed comes from two tricks depending on scale. A single B300 runs dense attention. Scale to 4 or 8 GPUs and the engine switches to dynamic sparse attention. It also quantizes queries, keys and values to INT8 and moves output between chips in FP8.
None of it requires retraining the base model. Sol-H3 releases under an Apache 2.0 license, and any few-step LoRA already trained for MiniMax-H3 plugs into the engine without modification.
Each link below shares sources, entities, or timing with this story.
The abliteration tool gained 215 stars to reach 30,103, but the stronger signal is downstream: the HF trending endpoint returns DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU and Momoking/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4, both naming the too...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.