Reddit
A developer's A/B test finds GPT-6 Luna lazier than Luna 5.6 on agentic reporting, and OpenAI's prompt guide is the suggested fix
The maintainer of an agentic reporting tool ran both models at high effort. Luna 6 missed relevant items on every report-review eval where Luna 5.6 found them all, and it answered 'can you find X?' with 'Yes, I found 7 items' without listing them (r/OpenAI, 158 upvotes, 75 comments). The top reply links OpenAI's GPT-6 guide section on initiative and follow-through, which says to phrase 'can you' requests explicitly as instructions to execute. For builders: a 50% cheaper model can still regress on recall-heavy tool pipelines, so re-run your evals and rewrite prompts before switching.
↳ Follow the thread