Reddit
A Depth-Pruned Qwen3.8-27B Cut to 22.7B With No Fine-Tuning, and the Author Refuses to Claim It Beats Anything
StargazerLabs released Qwen3.8-23B-Mini-Me, a strategic layer-removal prune of Qwen3.8-27B down to roughly 22.7B parameters with no fine-tuning afterward, in bf16, q8 and q4 (MLX only at the moment). The author explicitly declines to publish benchmarks or claim superiority, describing it as a smaller, faster version that is slightly worse at some things. The useful failure mode they do report: it handles standard coding fine but degrades against the original on edge cases and underspecified prompts, where the full model's extra capacity does the inference work.
↳ Follow the thread