Sources
The definitive Jev podcast: Diogo Almeida argues all three branches of RLHF were the wrong north star
Latent Space published a 2 hour 21 minute interview on September 21 with Diogo Almeida, CEO of TypeSafe AI and an InstructGPT coauthor, on why he built Jev as a System One decision model. Almeida traces RLHF to three lineages (Christiano et al. 2017, Stiennon et al. 2020, and his own Ouyang et al. 2022) and argues all three are the wrong target, and that the field dropped every mode other than autoregressive chat-tuned LLMs purely because ChatGPT succeeded. The episode also points at his published note on the Tyranny of the KV Cache as the real argument about Jev inside coding agents, and he publicly rejects generic JevBench-style benchmarks in favor of the cookbook patterns.
Source
↳ Follow the thread