Fetching from the wire…
Public story · 2026-07-20 · high
It uses p-bits to run the sampling and optimization work GPUs own, but it's a research machine, not a product yet.
Why now: It surfaces in the July 20 coverage as one of the only non-software answers yet floated for AI's energy cost problem.
A new probabilistic computer treats thermal noise as a computational resource instead of an error to correct, and IEEE Spectrum calls it the largest machine of its kind built so far. Instead of fighting the randomness silicon generates at room temperature, it runs on p-bits, probabilistic bits that use that noise directly to solve sampling and optimization problems.
That's the same class of work GPUs own. Most inference and optimization jobs burn GPU power to do this kind of math, and energy cost is turning into the industry's real constraint on scaling AI. A different physical substrate that handles the same problems is one of the few answers to that cost that isn't just a software optimization.
Research hardware, not a product yet. Spectrum's coverage doesn't say how it compares to a GPU on cost or power per operation, and it doesn't say when, or if, this leaves the lab. I'd want that number before believing the energy story holds up outside a controlled setup.
My bet: the fix for AI's energy problem won't come from a faster GPU or a smarter kernel. It'll come from someone swapping out the substrate entirely, the way this does. Everything else is optimization on top of the same physics.
Worth watching for the first benchmark against a GPU doing the identical sampling job. That number doesn't exist yet.
Each link below shares sources, entities, or timing with this story.
Every tool call, retrieval step and orchestration decision runs on general-purpose silicon, which has made the CPU the emerging bottleneck rather than the GPU (IEEE Spectrum). The piece expects CPU shortages and price increases following the pattern already seen in GPUs and me...
TechCrunch reports the three-Harvard-dropout company raised at that mark on claims its silicon accelerates inference on any model without GPUs. Skeptical read: Etched has been promising silicon for years and a valuation isn't a shipped benchmark. Investor conviction at this si...
Hetzner Experiments runs a token-authenticated API with explicitly no billing, no SLA, no production guarantee. One model live: Qwen3.6-35B MoE with quantized weights, informally measured at ~153ms median TTFT and 224 output tokens/sec. The real question is hardware. Hetzner's...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
NVHBM, announced August 26, relocates NVIDIA's custom memory controller from the compute chip into the HBM base die, claiming up to 30% more bandwidth than standard HBM4E, 15% lower HBM power, and up to 25% more freed area on the XPU compute die (NVIDIA). Next-generation Train...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.