Needle: 26M Parameter Model Distills Gemini Tool Calling — 6,000 Tok/s on Consumer Hardware, 489 HN Points
GitHub·medium signal
Cactus Compute released Needle, a 26M-parameter 'Simple Attention Network' that distills Gemini 3.1's function-calling capabilities. The model achieves 6,000 tok/s prefill and 1,200 tok/s decode on consumer CPUs, fits in 14MB at INT4 quantization, and runs inference in under 100ms. Trained on 200B tokens across 16 TPU v6e chips in 27 hours, it replaces MLPs with cross-attention for retrieval-and-assembly. At 489 HN points, it's the highest-engagement Show HN of the day.