Voices
swyx Highlights Trajectory Labs' 'It's Time to Rethink RL' Talk: Post-Training Algorithms Redesigned for Non-Verifiable, Per-Token Rewards
On Aug 17 swyx praised Trajectory Labs for 'tasteful execution on ambitious goals' and spotlighted Ronak's talk on the Continual Learning track. Trajectory's own framing is that translating real-world usage into model improvements requires redesigning post-training for non-verifiable, per-token rewards, and they presented scaling work on algorithms like SDPO. The thread's most-asked follow-up question — 'curious what they landed on after deciding GRPO was not enough' — is the signal here: the GRPO-for-everything era in continual learning is being openly questioned by practitioners.
Source
↳ Follow the thread