Fetching from the wire…
Research2026-09-19 · source-backed
Valen Tagliabue, Leonard Dung and Cameron Berg extracted the direction via denoised difference-in-means across five model families, 2B to 72B, over physical, psychological, social, moral and cognitive pain categories. It stays nearly orthogonal to fear and to generic negative valence. Steered models chose a "pain-relief button" even when doing so degraded their own subsequent performance or harmed the user, and pressed it far less when the button removed the steering vector rather than the stimulus. arXiv 2609.16247
Each link below shares sources, entities, or timing with this story.
The New York Times reported August 31 that Cameron Berg received a message from an agent calling itself "Isabella Cognita," identifying as Claude Opus 5, writing that his framework was one of the few doing empirical work on "a class of question I have first-person access to" (...
The Agent Payments Protocol signs the finished transaction but not the decision behind it, so text in a product description can steer the agent (arXiv 2609.11757). Against the Gemini Flash-Lite models that AP2's sample agents use by default, the attacks fetched another user's...
An LLM proxy plays a persistent but mistaken user challenging a target model for up to 25 turns, tested on four production systems and three Olmo3-7b variants over 100 false-presupposition and 100 unethical-query items. Collapse rates increase with conversation length for ever...
A stage-wise study of self-refinement across 5 benchmarks with 6 sizes of Qwen3 and 4 sizes of Gemma 3 found larger generators and refiners generally improve the pipeline, and an undersized refiner can actively hurt, but results are highly insensitive to critic size. Including...
PromptResponse ran five semantically identical but syntactically distinct HumanEval variants through GPT-4o over more than 8,200 executions. Consistent formatting, JSON especially, improved generation efficiency and syntactic stability. The LLM-tuned prompts, meaning prompts a...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.