Fetching from the wire…
Public story · 2026-07-19 · high
The July 17 paper is framed for robotics, but the same gap, a correct plan followed by a wrong action, shows up in any agent.
Why now: The paper posted to Hugging Face's papers feed on July 17 and had picked up 38 upvotes as of July 19.
BadWAM isolates models that dream right but act wrong, per a paper posted to Hugging Face's papers feed on July 17.
Framed for robotics, but the same gap turns up in any agent that plans correctly and then acts differently. That's the debugging conversation I have with coding agents regularly: the reasoning trace looks right, the diff still doesn't match it.
The model's internal rollout, its prediction of what should happen, comes out correct. The action it actually sends diverges from that rollout anyway, per the paper. The paper doesn't say whether it proposes a fix, only a way to isolate the mismatch.
Whether that check generalizes past robotics is the open question. If it does, it's the first named test for a failure mode a lot of agent debugging boils down to: right answer, wrong output. Worth watching whether anyone runs the same check on a coding agent's trace next.
Each link below shares sources, entities, or timing with this story.
The August 14 report covers January through August 2026: model repos grew from 2.43M to 2.96M, datasets from 711K to 1M, and 85.6% of models have under 200 lifetime downloads (Hugging Face). Chinese labs shipped monthly parameter ceilings of 754B to 2.78T against sub-130B for...
Anaconda acquired Kilo Code on July 15; Kilo supplies planning, coding, and debugging agents inside VS Code, JetBrains, and the CLI. Anaconda's moat is being the default Python environment in enterprises and universities, and that moat is worth very little if the agent layer a...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.