Fetching from the wire…
Security2026-07-26 · source-backed
Kim, Park, and Choi formalize a failure mode where an unfinished harmful prompt elicits a harmful continuation, because models postpone refusal until sentence termination rather than evaluating intent mid-sentence (arXiv 2607.20473). Training models to refuse incomplete harmful prompts via parameter tuning failed to generalize across both content domains and attractor types. The patch doesn't transfer. They localize two functional neuron groups, termination and continuation neurons, and argue neuron-level intervention is the more precise lever.
Each link below shares sources, entities, or timing with this story.
Prior work found individual parameters whose removal collapses LLM performance by orders of magnitude. arXiv 2607.08733 shows the effect isn't universal across models, then tests the obvious corollary that Super Weight-aware training should work. It doesn't. Training 100 to 8,...
The standard protocol ablates a latent and measures effect at the token where it fires hardest, but that token is chosen by the dictionary under evaluation. Two dictionaries get compared at different places. Training six autoencoders from one initialization showed 7.6% and 11....
arXiv 2608.12253 shows the standard practice of training a policy against a single LLM simulating the user fails because the simulator is itself mode-collapsed, so the policy learns to exploit its dominant mode. Verbalized Sampling recovers up to 9% held-out success; Populatio...
RETRACE has a verifier infer what problem the patch appears to solve using only the patch and trajectory, then compares that inference against the real issue. Training-free, lifted Pass@1 by 7.0% and 3.6% on mini-SWE-agent over SWE-bench Verified. The information-hiding trick...
The diagnosis in this paper is better than the fix, and the fix is very good. Recurrent memory agents fail at long context, but not for the reason most people assume. The bottleneck isn't capture. It's retention. Retention falls below 30% at 896K tokens because every consolida...
100 real frontier research tasks across seven scientific domains, full lifecycle, 800 annotated trajectories, 45-pattern failure taxonomy (arXiv 2608.14905). The headline isn't a leaderboard, it's a shared deficit: agents can't check what they produced against what they found,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.