Five 2026 GenAI Systems All Beat the Average Student on Intro OOP Exams but Still Break on Interfaces and Abstract Classes
ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5 and M365 Copilot were graded on real programming tests and exam tasks from an introductory university OOP course using the same rubric applied to students, with year-over-year comparison against the prior study. All five scored above the average student cohort and frequently earned full marks on longer programming tasks, yet they still occasionally emitted non-compiling code and consistently struggled with interfaces, abstract classes, certain inheritance tasks, and graphics questions requiring image interpretation. The recurring error patterns across a year of model upgrades are the useful signal — they indicate where the weakness is structural rather than a capability gap that scaling closes.
↳ Follow the thread