Hacker News
Show HN: MoEspresso Squeezes DeepSeek V4 Flash Coder From 84GB to 56.8GB by Deleting ~80B Parameters, Then Writes a C Compiler on a Mac
A Show HN (17 points) packages DeepSeek-V4-Flash-0731-Coder at 56.831GB — 32.63% smaller than the prior 84.355GB / 2.37bpw build — by dropping roughly 80B of the model's parameters and mixing IQ1_S_R4, IQ2_KS, and IQ2_K quantization with q6_K on denser weights, retaining 204B total. It runs only under MoEspresso 2.1+ on Apple MLX, installed via Homebrew, and the demo is the model writing a small C compiler locally. Single-source and unbenchmarked — there is no quality measurement against the full-size weights, which is exactly what a 32% expert-pruning claim needs.
↳ Follow the thread