Reddit
Tinygrad Driver Testing on Blackwell + M3 Ultra RDMA Cluster — Nearly 2TB RAM for MoE Experiments
A r/LocalLLaMA user is testing Tinygrad drivers on a cluster combining NVIDIA Blackwell GPUs and Apple M3 Ultra via RDMA, with nearly 2TB of combined RAM, specifically targeting MoE (Mixture of Experts) model inference speeds. The post hit 95 upvotes and 53 comments (0.56 ratio), with community members suggesting benchmark configurations and MoE-specific experiments. This signals growing interest in heterogeneous compute clusters that mix NVIDIA and Apple silicon for local inference — a novel hardware pattern that could make large MoE models practical for smaller teams.
Source
↳ Follow the thread