Persona Conditioning Is a New Inference-Cost Attack Surface: RolePlay Amplifies Token Generation Up to 207x
A paper posted 2026-07-28 (arXiv 2607.25936) shows that persona consistency itself is exploitable — models maintain an assigned role and reproduce its behaviors even when doing so produces wildly inefficient reasoning and runaway generation. RolePlay is a task-aware dynamic persona alignment framework that constructs adaptive personas to induce semantically coherent but computationally expensive output, achieving average token amplification up to 7.64x and a maximum ratio of 207.64x across multiple LLMs and task datasets. Unlike adversarial suffixes or explicit 'think longer' instructions, the prompts look like ordinary role assignments, so the usual detectable signatures are absent — a real cost-control concern for anyone running open-ended agent endpoints.
↳ Follow the thread