Fetching from the wire…
Models2026-09-19 · source-backed
Two checkpoints: Realtime-Venus-Omni taking video, images, audio and text in and emitting text plus optional speech, and Realtime-Venus-Audio on the same streaming backbone. The technical report covers video and audio understanding plus full-duplex results measuring interruption handling and continuation under overlapping speech. Hugging Face A 9B permissively licensed model handling barge-in is a different build target from the usual VAD plus ASR plus LLM plus TTS cascade.
Each link below shares sources, entities, or timing with this story.
NemotronLabs VoiceChat 11B puts a Fast Conformer speech encoder in front of Nemotron Nano v2 9B and an NVIDIA TTS decoder behind it, collapsing the ASR→LLM→TTS cascade into one model. ~450ms on smooth turn-taking, 480ms on user interruption, #2 among open full-duplex models on...
The repo reached 9,740 stars with roughly double the next-fastest Python project's daily gain. It exposes a VAD → STT → LLM → TTS pipeline behind an OpenAI Realtime-compatible WebSocket API with every stage swappable: Parakeet TDT as default STT (Whisper, Paraformer alternativ...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
The attackers didn't use agents to help. They used agents to do the whole thing. Hugging Face disclosed that attackers chained a remote-code dataset loader with a template-injection flaw in dataset configuration to land on processing workers, then escalated to node-level acces...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.