Tencent's AngelSlim shipped a ternary GGUF that puts the 770B Hy4-preview in 214 GiB at 2.38 bits per weight
Hugging Face (AngelSlim), via r/LocalLLaMA·high signal
The AngelSlim/Hy4-preview-GGUF repo offers Q4_K_M at 435.20 GiB (4.86 bpw) and STQ1_0 at 213.66 GiB (2.38 bpw), roughly half the size, benchmarked at 204.56 t/s prefill and 20.47 t/s decode on 8x H20. STQ1_0 comes from llama.cpp PR #22836 and uses ternary weights with 3:4 forced sparsity at 1.3125 bpw on routed-expert gate/up projections across 29 layers, with IQ2_XXS at 2.0625 bpw on the other 48; an imatrix is mandatory. The r/LocalLLaMA thread (821 upvotes) circulated a claim of ~98% performance retention, and the top skeptical reply notes that 98% on KL divergence alone can be misleading.