Open-Weight Code Models Fabricate on 60% of Impossible Tasks and Refuse Only 27%
This work separates code hallucination as ungrounded generation from ordinary code error, and taxonomizes it along groundedness (absolute violations of universal truths versus relative fabrications of ecosystem-specific facts), manifestation level (syntactic, semantic, factual) and behavior (confident fabrication through degenerate output). The adversarial suite contains 270 deliberately unsatisfiable prompts across six languages and 24 subcategories, paired with 91 matched solvable controls, judged by a two-tier protocol validated at 82% human agreement and kappa 0.73. Across twelve open-weight code and reasoning models and 4,332 judged responses, models produced ungrounded code on about 60% of unsatisfiable prompts and refused only 27%, while wrongly refusing 0% of the solvable controls, which means the failure is one-directional and refusal training has room to move without hurting valid work.
↳ Follow the thread