Fetching from the wire…
Infra2026-07-18 · source-backed
In a July 17 post, NVIDIA pushes "intelligence per dollar" as the agentic-era metric, arguing post-training rather than pretraining is now the central workload because agents need continuous improvement cycles. The claim: a 10-trillion-parameter MoE on 100 trillion tokens in one month using 25% as many GPUs. Supporting open libraries are NeMo Gym (environments as dataset + agent harness + verifier + per-task state) and NeMo RL. Vendor benchmark, treat accordingly, but the reframing of post-training as the main event tracks what I see people actually spending compute on.
Each link below shares sources, entities, or timing with this story.
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
The Hrazdan facility opened August 8, scaling to 300 megawatts and 70,000+ NVIDIA Rubin and Blackwell GPUs by end of 2027, built on NVIDIA DSX (40% more GPUs on the same footprint) with Dell PowerEdge, Schneider Electric power and Vertiv cooling. NVIDIA intends to invest, foll...
NemoClaw lets teams run agents like Hermes and OpenClaw inside NVIDIA OpenShell with managed inference and a security-hardened runtime, sitting at ~21,000 stars (≈248/day, 85 days old). A silicon vendor moving up the stack from "we sell GPUs" to "we run your agents in a sandbo...
NVHBM, announced August 26, relocates NVIDIA's custom memory controller from the compute chip into the HBM base die, claiming up to 30% more bandwidth than standard HBM4E, 15% lower HBM power, and up to 25% more freed area on the XPU compute die (NVIDIA). Next-generation Train...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
OpenAI posted first benchmark results for Jalapeño, its Broadcom co-developed inference ASIC, claiming 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia GB200 and GB300 rack systems, measured on the SemiAnalysis InferenceX suite. T...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.