AINews Maps Eight Stages of Humans Being Replaced in the ML Pipeline, and Says Verification, Not Generation, Gated Every One
The August 22 issue lays out a progression from reward signals (InstructGPT, 2022) through training data (Phi), teachers (Alpaca's $600 fine-tunes), curricula, researchers (Karpathy's autoresearch stacking 700 overnight experiments to cut GPT-2 training from 2.02 to 1.80 hours), environments (Z.ai), human subjects (Simile), and now cells. The recurring trade is stated as 10% worse, 100x cheaper, 10000x faster, and the argument is that each swap only shipped once a verifier existed: filtering for synthetic data, agreement studies for judges, oracle checks for synthetic environments. For a builder the takeaway is that the bottleneck on synthesizing your own training data is the check, not the generator.
↳ Follow the thread