Hacker News
nano-llm-posttraining: SFT, DPO and GRPO on One 8GB GPU, Instrumented for Forgetting and Seed Variance
A Show HN repo (created July 21, 15 stars, last pushed August 1) offering minimal readable post-training experiments that fit on a single 8GB GPU. What separates it from the usual minimal-implementation genre is what it measures: catastrophic forgetting, seed-to-seed variance, and the emergence of RL behavior — the failure modes that tutorials usually omit. Zero HN comments on 21 points, so this is an unvalidated find, useful as a cheap local harness for building intuition about post-training rather than as a production recipe.
↳ Follow the thread