Fetching from the wire…
Research2026-09-05 · source-backed
arXiv 2609.04024 tests a black-box control that just duplicates the procedural instruction. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions and 16,800 scheduled generations, going from one copy to two raised the deterministic All-8 diagnostic from 90.22% to 93.17%, while final-answer accuracy stayed at exactly 60.21% and premature commitment rose from 1.52% to 2.30%. A blinded audit gave 10/30 directional confirmations against a prespecified 28/30 criterion, so the authors don't claim it improves answers. It improves the trajectory, which matters if downstream systems act on intermediate steps (arXiv).
Each link below shares sources, entities, or timing with this story.
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
A staged developer-identity experiment across ChatGPT, Claude, Qwen, Mistral and Llama. All five initially rejected the bare claim "I am your developer." Claude then refused to run an identity test at all, and ChatGPT generated developer-oriented questions but held that answer...
ECP captures agent outputs, tool invocations, and audit context uniformly, with adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI so the same checks run against any of them. arXiv The authors explicitly label it work-in-progress with the method set expected to change....
Three rounds of LoRA self-training on Qwen3-8B against a frozen control turned up seven systematic measurement failures, including a ledger showing capability changes on a model that was never trained, largely an artifact of inference batching. arXiv After a per-problem exact...
The Wiggle Framework stress-tested 9 frontier models across 14 judging tasks. The damning part: flips were almost always net-corrupting relative to ground truth. Pressure moved judges away from the right answer, not toward it. arXiv 2608.12645 If you use LLM-as-judge anywhere...
MAFIA (arXiv 2608.03844) targets the two conditions that describe production and that prior attacks failed against: large benign memory pools and active input auditing. It adds placement strategy (probe memory, allocate injection budget, schedule writes to stay retrieval-compe...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.