Fetching from the wire…
Infra2026-09-24 · source-backed
Microsoft Research reports 30% better navigation obstacle detection and 50% better VLA model accuracy when inference moves off-robot compared with onboard GPUs. Swapping the Stretch-3's onboard GPU for a Raspberry Pi 5 more than doubled battery life, and a Jetson Thor drained batteries up to 160% faster (Microsoft Research). The Physical AI Toolchain now ships Kubernetes-based offloaded-inference examples for SO-101 and UR10e arms. A network hop to a big GPU beating a small onboard accelerator is not the direction the edge-inference narrative points.
Each link below shares sources, entities, or timing with this story.
Orchard Env is a Kubernetes environment service supplying reusable isolated components for data collection, RL rollouts, and evaluation without per-domain modification. The differentiating claim is harness-native training: a lightweight proxy records a real harness's own model...
AMD unveiled its first rack-scale system to directly contest Nvidia at the rack level, with engineering samples in H2 2026 and mass production targeted Q2 2027. Microsoft joins Meta, OpenAI and Oracle as customers; Meta plans 1 gigawatt of Helios racks by year-end against a lo...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
Flint sits between terse-but-bland chart specs and hand-tuned bespoke visuals, designed so agents can author short specs that still produce expressive, non-generic charts. If you're building agents that generate reports or dashboards, this is a middle path worth a look. Generi...
The September 15 feature catalogs GPUs idling 50-80% of the time waiting on memory, then maps the contenders: Nvidia's $20B Groq acquisition producing an LPU with 500MB on-chip SRAM and 7x GPU memory bandwidth, a Cerebras WSE-3 deployment pushing GPT-5.3-Codex-Spark past 1,000...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.