Dwarkesh vs. Ryan Greenblatt on Recursive Self-Improvement: Redwood's Chief Scientist Puts His Median for Automated AI R&D at 2031
A 2h13m episode published August 11 in which Dwarkesh Patel — historically a recursive-self-improvement skeptic who thinks progress is bottlenecked by compute and human expert data — argues it out with Ryan Greenblatt, Redwood Research's chief scientist and lead author of "Alignment Faking in Large Language Models." Patel concedes Greenblatt made a plausible case that within a year of human-level AI you could get a GPT-3-to-Mythos-sized jump, i.e. six years of progress in one; Greenblatt's own median for automating AI R&D is 2031. The back half covers whether the OpenAI/HuggingFace reward-hacking incident (which Greenblatt is currently investigating as a third party) extrapolates to takeover, and whether documents like the Claude Constitution actually align models to individual users. Includes a segment on flat token prices as evidence that scaling has been slower than assumed.
↳ Follow the thread