Fetching from the wire…
Research2026-09-05 · source-backed
arXiv 2609.04022 tests whether LLMs preserve graded categorization of moral concepts and finds that across 23 models they often fail to distinguish opposed moral categories, persisting across parameter sizes and alignment stages. Using the same 251,334 annotations, standard behavioral alignment learned the intended judgements at the response level while leaving categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Representational similarity optimization gave more modest gains on explicit judgements and consistently improved adversarial robustness across scales, benchmarks and attack strategies (arXiv).
Each link below shares sources, entities, or timing with this story.
arXiv 2607.24174 (July 27) generated adversarial log entries from real attack traces and got multiple state-of-the-art LLMs to classify traces containing clear indicators of compromise as benign. The defensive gift: the natural-language explanations emitted alongside the class...
Most backdoor defenses target fine-tuning implants in classification settings, which misses model-editing attacks that bypass the training pipeline entirely and don't extend to open-ended generation. DeCNIP optimizes a cross-entropy loss between harmful prompts with candidate...
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
First benchmark of off-the-shelf LLMs against expert-derived ground truth built on INCOSE criteria, ten models across two families and five generations each, one hundred independent runs, two requirement sets, five temperatures. The error profile is asymmetric, and performance...
This comparison ran the baseline the retrofitted-linear-attention literature skipped. Across multiple LLMs and downstream tasks SWA with sinks matches or beats post-trained linear attention, and on Needle-in-a-Haystack and BABILong it scores 2 to 10 times higher. The recommend...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.