Reddit
Grok 4.3 Tops LLM Sycophancy Consistency Leaderboard — Most Cautious Frontier Model
Grok 4.3 has claimed the top spot on the lechmazur/sycophancy benchmark's Consistency Leaderboard, which measures whether a model maintains the same judgment regardless of who is speaking rather than just measuring flattery. The 63-upvote r/singularity post notes Grok 4.3's ranking is largely because it is 'one of the most cautious models.' This is notable given earlier Grok 4.1 analysis showed sycophancy increasing alongside empathy — xAI appears to have overcorrected.
↳ Follow the thread