Fetching from the wire…
Infra2026-09-23 · source-backed
AWS walks through CreateAIBenchmarkJob against a vLLM endpoint on a Blackwell-backed ml.g7e.2xlarge. Nemotron-3 Nano 30B plateaued at 2,823 output tokens/s at 256 concurrent requests. Add a 50s end-to-end and 1.5s TTFT SLA and safe concurrency drops to 80. That gap is the number to size against, not the peak.
Each link below shares sources, entities, or timing with this story.
The six open-source skills take a Hugging Face model reference and return a production real-time SageMaker endpoint, picking the serving container, wiring autoscaling and CloudWatch alarms, and verifying the result (AWS ML Blog). AWS choosing skills as the integration point fo...
NVIDIA launched Nemotron 3 Super — a 120B total / 12B active parameter hybrid Mamba-Transformer MoE, open, designed specifically for multi-agent workloads, and delivering 5x higher throughput than Nemotron 2 at the same active parameter count (NVIDIA Newsroom). It ships with a...
AWS documented serverless fine-tuning for NVIDIA Nemotron 3 on SageMaker, a semantic layer for agentic AI using Stardog on Bedrock AgentCore that grounds agent queries in a governed knowledge graph instead of raw tables, and four deployment patterns for Unsloth-quantized model...
NVIDIA and AWS announced June 23 that NVIDIA's cuVS library now powers GPU-accelerated vector indexing as the default in Amazon OpenSearch Serverless, claiming up to 10x faster index builds at roughly a quarter the cost versus CPU-only, making billion-scale vector DBs buildabl...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.