Reddit
Two days with GPT-Live-1 on real phone lines: best-sounding voice model, worst at following the script
A team building AI phone agents (ThunderPhone) posted first impressions after running a ~13,000-token insurance qualification script through GPT-Live-1 across a dozen real phone calls and roughly 25 simulated ones. Their verdict splits cleanly: it is the most natural-sounding model they have put on a line, full-duplex so turn-taking, interruptions and backchannel work, but instruction-following breaks down on a long script. That is the specific failure mode worth knowing before anyone ports a scripted voice workflow, and the post sits at only 15 upvotes on r/OpenAI, so it is a single hands-on account rather than a corroborated benchmark.
Source
↳ Follow the thread