Full-duplex voice agents lose up to 9.7% instruction adherence when behavior comes from a persona instead of a rule
DSB-IFEval (arXiv 2609.03423, 2026-09-03) tests the gap between how voice agents are benchmarked and how they are deployed: benchmarks give explicit turn-management instructions, but production agents are configured by role or persona and must infer when to listen, backchannel, interrupt or yield. Its 1,038 cases span eight assistant roles and five conditioning protocols, scored by a deterministic Instruction Adherence Score and an LLM-judged Persona Adherence Score. Across six real-time systems the trade-off is architecture-dependent: F-Actor and PersonaPlex drop 9.7% and 4.5% under persona-only conditioning, while GPT-Realtime, MiniCPM-o and Fun-Audio-Chat hold persona-consistent content better.
Source
↳ Follow the thread