Fetching from the wire…
Infra2026-07-24 · source-backed
Hetzner Experiments runs a token-authenticated API with explicitly no billing, no SLA, no production guarantee. One model live: Qwen3.6-35B MoE with quantized weights, informally measured at ~153ms median TTFT and 224 output tokens/sec. The real question is hardware. Hetzner's public GPU lineup is RTX 4000 Ada and RTX PRO 6000 Blackwell, workstation cards that can't serve frontier-size models. Whether they buy datacenter GPUs is the actual signal about a European budget inference tier.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.