Fetching from the wire…
Research2026-09-03 · source-backed
This work pairs E2M1 payloads with unsigned E5M3 block scales whose wider range permits periodic tensor scaling, applies selective stochastic rounding only to backward gradients, drops the Hadamard transform entirely, and uses FP4 in every eligible internal linear. Pretraining a Nemotron-H 8B for nearly 190 billion tokens, the block-16 recipe finished with lower final-window training loss and lower held-out validation NLL under each method's own quantized-inference policy. An ablation removing both the Hadamard transform and the BF16 final-block exemption raised measured model-body token throughput 21.2%.
Each link below shares sources, entities, or timing with this story.
Nemotron-Terminal (arXiv:2602.21193, 66 HF upvotes) — First systematic study of data engineering for terminal/CLI agents. Terminal-Task-Gen pipeline with Dockerized environment interaction. Qwen3-initialized 8B model goes from 2.5% to 13.0% on Terminal-Bench 2.0. All checkpoin...
NVIDIA released Star Elastic, a post-training method that nests three submodels (30B, 23B, 12B) inside a single Nemotron Nano v3 checkpoint. The technique uses only 160B tokens (360x reduction vs pretraining) and cuts memory for deploying all three from 126.1GB to 58.9GB in BF...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Re-architected PhysicsNeMo libraries and updated CUDA-X exposed as agent-ready tools, so engineering agents can invoke AI physics models, accelerated solvers and quantum chemistry directly instead of through bespoke wrappers. NVIDIA Research's ACE-RTL agent with Nemotron 3 Ult...
At NVIDIA's July 20 keynote, four major creative tools announced MCP connections letting agents operate inside them rather than through file exports. NVIDIA paired it with a deskside agent supercomputer tying Omniverse, Blender, and Nemotron into a local runtime. Real adoption...
The Series C, led by Aramco Ventures with NVIDIA, Vista, and others, funds a plan to grow capacity roughly 50x over five years, serving DeepSeek, Nemotron, MiniMax, and Kimi. Capital is still flowing hard into open-model serving infrastructure, which is the supply side of the...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.