Fetching from the wire…
Infra2026-09-04 · source-backed
Released in beta at IFA 2026, PAIR distributes inference requests across whichever PCs on a local network have spare capacity, so agentic workflows run in parallel. Supports GeForce RTX 20 Series and newer, RTX PRO workstation GPUs from Turing on, DGX Spark, and Apple M4 or newer, working with Ollama and LM Studio on Windows, macOS and Linux. Nvidia cites 1.9x llama.cpp throughput on RTX 5090 and 1.2x-1.4x vLLM gains, with one-click setup for Hermes Agent and OpenClaw. RTX Spark PCs from Lenovo and Acer in October. NVIDIA
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
Hetzner Experiments runs a token-authenticated API with explicitly no billing, no SLA, no production guarantee. One model live: Qwen3.6-35B MoE with quantized weights, informally measured at ~153ms median TTFT and 224 output tokens/sec. The real question is hardware. Hetzner's...
NemoClaw lets teams run agents like Hermes and OpenClaw inside NVIDIA OpenShell with managed inference and a security-hardened runtime, sitting at ~21,000 stars (≈248/day, 85 days old). A silicon vendor moving up the stack from "we sell GPUs" to "we run your agents in a sandbo...
Portable Computer launched August 26, running the orchestrator LLM, subagent LLM, planner, tool router, scheduler and local search index locally, with local work consuming no billing credits and each cloud escalation requiring separate approval (VentureBeat). Launch platform i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.