Fetching from the wire…
Public story · 2026-08-24 · high
Xiaomi says the AI Cube can run 120B and 3B models locally, powered by DRAM stacked directly on the XRING O100 chip.
Why now: As of August 24, the AI Cube specs are the newest number for anyone comparing local-inference hardware against DGX Spark and Strix Halo.
Xiaomi showed off an AI Cube prototype that stacks three of its own chips together, hitting 1.22TB/s of near-memory bandwidth at 150W sustained, according to specs shared on r/LocalLLaMA.
Xiaomi says the AI Cube can run 120B and 3B parameter models on the device itself. That puts it alongside DGX Spark and Strix Halo as a self-contained local-inference box, if it ships as a real product instead of staying a Hot Chips demo.
The chip behind it, the XRING O100, is built on a 6nm process with two DRAM layers stacked wafer-on-wafer directly on top. Xiaomi used a 1.4 micron hybrid-bonding pitch, with 28,672 effective data lines connecting the stack, a bandwidth trick usually reserved for datacenter accelerators.
Xiaomi hasn't said when, or if, the AI Cube reaches customers. Hot Chips is where companies show off architecture, not release dates. If this becomes real hardware, it undercuts the assumption that local inference requires buying into DGX Spark or Strix Halo. If it stays a slide, it's one more sign that phone makers want their own inference silicon, whether the chips ever leave the lab or not.
Each link below shares sources, entities, or timing with this story.
DGX Spark benchmarked against Strix Halo / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (DGX Spark benchmarked against Strix Halo); both cover DGX Spark, Strix Halo; overlapping topics (bandwidth, halo).
NVIDIA released DGX Spark / Shared entity: LocalLLaMA / Same source domain / Shared topic / Earlier coverage / Downstream implication
Linked by a graph relationship (NVIDIA released DGX Spark); both cover LocalLLaMA; reported by the same outlet (reddit.com).
NVIDIA released DGX Spark / Shared entity: LocalLLaMA / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA released DGX Spark); both cover LocalLLaMA; reported by the same outlet (reddit.com).
DGX Spark benchmarked against Strix Halo / Shared entity: DGX Spark / Earlier coverage
Linked by a graph relationship (DGX Spark benchmarked against Strix Halo); both cover DGX Spark; earlier DGX Spark coverage from 2026-04-03.
Linked by a graph relationship (DGX Spark benchmarked against Strix Halo); both cover DGX Spark; earlier DGX Spark coverage from 2026-03-23.
NVIDIA released DGX Spark / Shared entity: LocalLLaMA / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA released DGX Spark); both cover LocalLLaMA; reported by the same outlet (reddit.com).
Linked by a graph relationship (NVIDIA released DGX Spark); both cover LocalLLaMA; reported by the same outlet (reddit.com).
Linked by a graph relationship (NVIDIA released DGX Spark); both cover LocalLLaMA; reported by the same outlet (reddit.com).