Fetching from the wire…
Public story · 2026-08-07 · high
The deal pairs Taalas chips with Instinct GPUs in Helios racks under ROCm, closing in Q4 weeks after AMD's Cerebras deal.
Why now: The deal was announced August 6 and is set to close in the fourth quarter, weeks after AMD's separate agreement with Cerebras aimed at the same inference-cost problem.
AMD agreed to acquire Taalas on August 6, adding chips etched for a single model's weights to its lineup, per AMD's announcement.
The move targets inference cost, not raw compute. Taalas says its HC1 chip served Llama 3.1 8B at 16,960 tokens per second. That's a claimed 48 times the throughput of Nvidia GPUs and 8.5 times Cerebras's chips on the same task. For AI-native software priced below its own inference bill, that ratio is the gap between a viable product and a subsidized one.
Taalas builds what it calls Hardcore Models. Its chips are physically wired to one model's weights, finished by completing only two of a chip's more than 100 metal layers. Tape-out takes around two months, according to AMD. The tradeoff is fixed: a Taalas chip can't switch models without a new production run.
Terms weren't disclosed. The deal is expected to close in the fourth quarter. After that, Taalas's Toronto-designed silicon runs in AMD's Helios racks alongside Instinct GPUs, under AMD's ROCm stack. It's AMD's second inference-specialized chip deal in weeks, following an earlier agreement with Cerebras.
The bet only pays off for the wider market if AMD sells that speedup by the token, beyond its own racks. Both this deal and the Cerebras one target the same line item: the cost of serving a model that's already trained. That's the COGS number squeezing every AI product priced below what it costs to run.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
AMD unveiled its first rack-scale system to directly contest Nvidia at the rack level, with engineering samples in H2 2026 and mass production targeted Q2 2027. Microsoft joins Meta, OpenAI and Oracle as customers; Meta plans 1 gigawatt of Helios racks by year-end against a lo...
Reuters, via Tech Startups, reports capital released against deployment milestones with Anthropic deploying up to two gigawatts of Instinct MI450 starting 2027. Same structure as Nvidia/OpenAI: compute vendor capital flowing to the lab that commits to buy the silicon. A two-gi...
Canadian startup Taalas unveiled its HC1 chip — a hardwired implementation of Llama 3.1 8B that generates 17,000 tokens per second per user, 73x faster than NVIDIA's H200 at one-tenth the power. Using aggressive quantization (3-bit and 6-bit) on TSMC N6 at 815mm2 die size, dra...
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark (~$4k), AMD's Strix Halo / Ryzen AI Max+ 395 (~$2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly...
Per SSOJet, that cross-company contributor base is the unusual part. OSS coding agents usually orbit one vendor. This one's becoming shared infrastructure, which makes it the self-hostable default to watch against proprietary agents. If you want a coding agent you can run and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.