Fetching from the wire…
Public story · 2026-08-04 · high
ComfyUI's day-0 build runs on an RTX 3060 after cutting memory needs 66 percent, to a 42.5GB minimum.
Why now: ComfyUI added support the same day MiniMax shipped the weights, skipping the usual lag between an open release and tool integration.
MiniMax released H3, an open-weights video model, with ComfyUI support live the same day, per the ComfyUI Blog.
H3 generates stereo audio synced to the video instead of dubbing it in afterward, the first open-weights model to pair audio and video natively. A 66 percent cut in memory requirements brought the smallest variant to 42.5GB, small enough for an RTX 3060 instead of enterprise-grade GPUs.
The model takes text, image, video, or audio as input and outputs clips up to 15 seconds long at up to 2K resolution. It supports first-and-last-frame control: you set the start and end frames and the model fills in the motion between them. It also does motion transfer from a reference clip.
Weights are posted at Comfy-Org/MiniMax-H3 and require ComfyUI 0.30.0 or newer.
Dubbed audio on generated video tends to drift out of sync with mouth movement and sound effects. Generating it inside the model, on the same timeline as the frames, avoids that failure mode. Test it against dubbed output before trusting the resolution number alone.
Each link below shares sources, entities, or timing with this story.
Open-sourced August 3 under the MiniMax H3 Community License, it now holds four of the top 20 trending slots: base at 47.5k downloads and 3.35k likes, Comfy-Org's mirror at 6.01M, lightx2v's Turbo at 15.1k. It's a 33B dense transformer (about 13B in AdaLN branches that cache a...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Launched July 30: all-modality input (text, image, video, audio), clips up to 15 seconds at 2K with native stereo sound, footage editing and motion transfer from reference material. MiniMax claims 2K generation at under a third of mainstream rivals' cost and says H3 was design...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.