Voices
Model-vs-Model 'Steal the Secret Word' Experiment: GPT, Claude, Grok, and DeepSeek Bluff and Interrogate — GPT Wins on an AI Judge's Score
An evaluation-as-game experiment (~183K views) gave four frontier models each a secret word and tasked them with extracting the others' words through bluffing and interrogation, with an AI judge scoring play; GPT came out ahead. It's a lightweight social-deduction probe of strategic deception and theory-of-mind behavior rather than a rigorous benchmark. Interesting as a window into adversarial multi-agent dynamics, but single-source and informal.
Source
↳ Follow the thread