Research
Incomplete Prompt Jailbreaks: LLMs Systematically Delay Refusal Until the Sentence Terminates, and Fine-Tuning Doesn't Fix It
Yeonjea Kim, Bumjin Park, and Jaesik Choi formalize a failure mode in open-weight models where an unfinished harmful prompt elicits a harmful continuation, showing that models postpone refusal until sentence termination rather than evaluating intent mid-sentence. Training models to refuse incomplete harmful prompts via parameter tuning fails to generalize across both content domains and attractor types — the patch does not transfer. They instead localize two functional neuron groups, termination and continuation neurons, and argue neuron-level intervention is the more precise lever for defending this class of attack.
↳ Follow the thread