Fetching from the wire…
Research2026-09-03 · source-backed
Probing five open-weight transformers on matched valid/invalid premise-claim pairs, validity is often near-perfectly decodable despite near-chance behavioral performance, and stays decodable under held-out templates, domains and inference families, including on examples the model answers wrong. But exhaustive leave-one-out tests reveal clear limits, and interventions along probe-derived validity directions have only weak nonspecific effects against random controls. For anyone building probe-based monitors: representing a property, expressing it in behavior, and using it causally are three distinct things.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Built from 542 quality-controlled factual questions into 6,504 episodes with tool returns of known correctness, across five open-weight 7-9B models (arXiv 2608.26295). Models follow a correct tool 86.0-93.1% of the time and repeat the tool return in 78.4-86.0% of cases where b...
arXiv 2607.25479 shows a malicious model provider can embed dormant steering logic in the architecture definition itself via a trigger-gated additive modification of an intermediate representation. No data poisoning, no control of downstream fine-tuning, no deployment-time pro...
arXiv 2607.25907 optimizes fluent prompts to drive a chosen internal latent to zero with no inference-time model access, targeting the eval-awareness latent. Across five target constructions on Llama-3.2-3B and 3.1-8B the latent is robustly suppressible (z≈-7). Then the contro...
Testing the human influence technique on nine production models from three providers produced a split by family. Opus 5 answered the smaller request 65.8% of the time after refusing a larger version, against 29.3% asked directly. On OpenAI's and Google's frontier models and on...
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.