Fetching from the wire…
Research2026-09-16 · source-backed
Four open-weight 7-9B models played repeated Prisoner's Dilemma, Snowdrift, Stag Hunt and Harmony in six framings each. Action policies drifted in all four games, and agent-generated pre-play communication was predominantly stabilizing, with five corrected reversals concentrated in social or team framings. Controlled interventions isolated two channels: reduced action uncertainty and less between-round drift. In Prisoner's Dilemma the authors found a history-balanced policy-content direction in late transformer layers, and projecting it out increased switching during closed-loop play, which makes it causal rather than correlational.
Each link below shares sources, entities, or timing with this story.
Seven models. Five harnesses. Controlled fact-withholding with injected faults. arXiv 2608.16630 is the most operationally direct paper I've read on harness design, and it produces three results that each change what I do this week. One: availability decides outcomes, not dist...
UniTexture backpropagates gradients from a Vision-Language-Action policy's action outputs to the surface texture of a single 3D object through a differentiable renderer, optimizing one shared texture over a distribution of tasks, instructions, states and viewpoints. Tested on...
PRISMA 2020 review, six databases, 743 records screened, 85 retained from 2023–2025 (arXiv 2608.10530). Perception-layer work (prompt injection, jailbreaking, adversarial perturbation) is 66% of papers. Action-layer vulnerabilities (tool misuse, code injection, sandbox escape)...
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
Researchers loaded five systems with a revoked policy and its replacement, then measured retrieval and downstream action across nine policy scenarios, nine models and six defense conditions. Wherever the revocation label was visible to the retrieval layer, the revoked fact cam...
arXiv 2608.02764 targets agents that issue refunds, reserve inventory and move money, where budgets and approval status change between authorization and effect. The authors define policy-state serializability: committed effects must be explainable as authorized against the pol...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.