Fetching from the wire…
Research2026-09-03 · source-backed
The Memory Trust Gap benchmark uses two suites on a same-family size series (Qwen3 0.6/1.7/4/8B): a Benefit suite unsolvable without the stored fact, and a Safety suite where an authoritative tool always holds the correct value. Models answer with the stale stored value 0.92 to 1.00 of the time at every scale in the Benefit suite, and in the Safety suite harm below the no-memory baseline is capability-gated, with larger models collapsing hardest once a stale note looks current. A 2x2x2x2 factorial shows removing a label amplifies over-trust at every size. This is over-trust, not confusion, so mitigation is itself capability-dependent.
Each link below shares sources, entities, or timing with this story.
Built from 542 quality-controlled factual questions into 6,504 episodes with tool returns of known correctness, across five open-weight 7-9B models (arXiv 2608.26295). Models follow a correct tool 86.0-93.1% of the time and repeat the tool return in 78.4-86.0% of cases where b...
Three rounds of LoRA self-training on Qwen3-8B against a frozen control turned up seven systematic measurement failures, including a ledger showing capability changes on a model that was never trained, largely an artifact of inference batching. arXiv After a per-problem exact...
Microsoft Research showed Qwen3-4B with a "Skeptical-Agent" outperforms 32B models and approaches 235B single-attempt performance. A 50x+ model size compression through inference-time self-refinement. Practical evidence that you can trade model size for inference-time compute...
Accuracy drops 30–50% well before you hit the documented context limit. Not at the limit. Before it. Cross-model testing across GPT-4.1, the Claude 4 family, Gemini 2.5, and Qwen3 quantified what everyone shipping long-context features has felt and couldn't measure (Glasp). Th...
The framing is Gricean: an uncertain cooperative speaker retreats up the specificity hierarchy, trading informativeness for truthfulness. On a T-REx-based benchmark varying entity familiarity and referent specificity, model activations do encode whether a referent falls inside...
arXiv 2607.28457 is an oracle-free multi-turn RL framework where the model emits a solution plus a discrete correctness verdict and a confidence score each turn, keeping its answer only when the verdict is Correct and confidence clears threshold. Ground-truth correctness shape...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.