Fetching from the wire…
Infra2026-09-19 · source-backed
The Kubernetes-native GPU-aware routing add-on for EKS scores pods on KV cache utilization, queue depth, whether the requested LoRA adapter is already loaded, prefix cache hit probability and active request count. AWS reports up to 97% lower time-to-first-token on mixed GPU fleets, up to 98% better P99 under bursty traffic, and 8-50% throughput gains, with no gain over round robin on uniform fleets with steady traffic. AWS The honest null case is stated up front, which I appreciate.
Each link below shares sources, entities, or timing with this story.
HyperPod InstantStart composes EKS orchestration with SageMaker HyperPod behind one API reachable three ways, including an MCP server publishing 38 tools spanning cluster lifecycle, instance groups, storage, model download, inference deployment and node operations. The design...
AWS published the arithmetic builders need to decide whether caching pays. A 10,000-token document reused across 10 questions nets roughly 75% off total input cost. Minimum checkpoint size is 1,024 tokens for Claude Sonnet 4.5 and 4.6, 4,096 for Opus. Default TTL is 5 minutes,...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
First-party pricing, counts toward AWS commitments, Codex via CLI and IDE plugins for VS Code, JetBrains, and Xcode, across commercial and GovCloud (AWS). This removes the procurement and compliance wall for AWS shops that couldn't touch OpenAI under existing contracts. Distri...
You noticed the Mac mini shortage. Here's what caused it. The Information reported, via Cult of Mac, that OpenAI has purchased tens of thousands of M5 Pro and M6 Mac minis plus M5 Max and Ultra Mac Studios over the past several months, running reinforcement learning and comput...
The walkthrough covers implementing MCP tools, wiring authentication, and deploying with AWS CDK against Bedrock AgentCore and Mistral AI Studio. Steal the two-layer JWT pattern: agent identity and end-user identity as separate token layers. Most MCP server tutorials hand-wave...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.