NVIDIA RTX Spark Superchip Hits Availability, Running 120B Local Models at 1M-Token Context
devFlokers·medium signal
The NVIDIA RTX Spark Superchip reached availability as a CPU+GPU part with up to 128 GB of unified memory, capable of running local models up to 120 billion parameters with 1-million-token context windows. It materially lowers the bar for self-hosting frontier-scale open-weight models like MiniMax M3 on a single machine. For builders, it shifts large-context local inference from datacenter-only to a desktop-class device.