Don't use an LLM to grade your generated onboarding docs — sycophancy, confabulation, and incoherence were pervasive across 52 judged configurations
This July 29 user study had 26 developers think aloud through 26 LLM-authored code tours built from real Java bugs, each independently judged by two different LLMs for 52 evaluated configurations. The tooling finding matters most: LLM-generated annotations of tour quality were unreliable, with sycophancy, confabulation, and incoherence all pervasive — so LLM-as-judge is not a usable proxy for documentation quality here. On the human side, developers preferred tours that scaled detail with code length, avoided restating code, stayed scannable, and used a guiding tone; they also trusted descriptions they believed were human-written more than ones they believed were AI-generated, and stack traces alone were often insufficient to surface the steps developers cared about.
↳ Follow the thread