Fetching from the wire…
Research2026-07-29 · source-backed
arXiv 2607.25907 optimizes fluent prompts to drive a chosen internal latent to zero with no inference-time model access, targeting the eval-awareness latent. Across five target constructions on Llama-3.2-3B and 3.1-8B the latent is robustly suppressible (z≈-7). Then the controls land: a random placebo direction is suppressed just as hard and shifts behavior just as far, and with a real eval passage in context, suppressing the eval direction slightly increases the model's behavioral eval judgment. Readability of a latent is not controllability of the behavior. Worth keeping in mind next time an interpretability result gets framed as a safety guarantee.
Each link below shares sources, entities, or timing with this story.
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Shipped today, giving Meta AI calendar access, guided research, daily briefings and recurring task automation, with the model deciding when to write a script versus drive on-screen controls. The Poly-Guard number is the one builders should read: roughly 1 in 14 multi-turn atta...
arXiv 2607.27180 decouples decision-making from execution: an off-the-shelf VLM issues atomic skill commands, a controller translates them into sub-second chunks of physically simulated full-body motion, so balance and motor failures are factored out. On 1,218 long-horizon ego...
The 2024 idea was that more agents talking to each other equals more intelligence. GroupChat. Everyone wired their agents to message each other. That pattern just lost, and it lost decisively. Anthropic, OpenAI, AutoGen, Cognition, and LangChain independently settled on the sa...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.