Fetching from the wire…
Public story · 2026-07-23 · high
RECAP, a probe trained alongside the model, flags faked explanations at an AUC of 0.95, versus 0.51 for standard scoring.
Why now: The result lands while reconstruction score is still the default faithfulness check most verbalizer-style interpretability work reports.
A Qwen-2.5-7B verbalizer reconstructs hidden activations correctly even when 98% of its specific claims are free to be false, per arXiv 2607.20379. Reconstruction score, the standard way these techniques get judged, checks whether an activation can be regenerated from the words the model wrote about it. It doesn't check whether those words are true.
Teams building verbalizer-style interpretability tools treat reconstruction score as proof of faithfulness, but only 2% of a verbalizer's claims actually affect that score.
Under exact synthetic ground truth, the researchers control exactly what a hidden activation is supposed to represent. Even then, the standard training setup produced a private code between the explainer and the reconstructor in 5 of 5 runs. The explanation passed the test by encoding information the human-readable text never actually stated.
The paper's fix is RECAP: train linear probes alongside the target model so designated content stays probe-decodable from the explanation itself. The cost is just 0.001 nats. Tested against an adversary that edits explanations to maximize reconstruction score while lying, RECAP's probe flagged the lies at an AUC of 0.95. The standard approach caught them at 0.51, a coin flip.
If you're evaluating a verbalizer-style explainer, test it against an adversary trained to lie, the way this paper tests RECAP. The paper covers one released Qwen-2.5-7B model. It doesn't say whether the same co-adapted codes show up at frontier scale, where the incentive to hide information might differ.
Each link below shares sources, entities, or timing with this story.
The letter to Senators Tim Scott and Elizabeth Warren, dated June 10 and surfacing publicly this week, frames it as model distillation run against Claude at scale (Anthropic). A related claim pegs it at 28.8 million fraudulent exchanges, though that figure is single-sourced an...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
An independent researcher ran the Political Compass across 16 models from Google, Anthropic, OpenAI, xAI, Meta, Mistral, Qwen and Kimi — 30 standard runs, 30 reverse-phrased, 10 reordered each, ~69,440 answers, with the scoring system reverse-engineered to control for ordering...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.