Fetching from the wire…
Research2026-08-05 · source-backed
arXiv 2607.28956, topping HuggingFace Daily Papers today with 74 upvotes, grounds 365 simulated days of e-commerce in 98,843 real product records, forcing agents to coordinate sourcing, pricing, cash management and order handling under mixed-latency feedback. Humans finished around 217,610 RMB; GPT-5.6 Sol reached 40,890 under ReAct and 52,930 under Hermes. The failure taxonomy is the part to steal: "operational coherence" decay, where activity just trails off, and "strategic coherence" breakdown, where goals drift and the agent stops updating on evidence. Humans sustained 100% engagement. LLMs ranged 10.6–66.1%. If you're shipping anything long-running, instrument engagement rate as a first-class metric.
Each link below shares sources, entities, or timing with this story.
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
A10 Networks made its AI Gateway generally available on August 14, pitched as a "centralized control plane for unified routing, cost management, and governance across every AI agent, application and large language model" (Help Net Security). SelectHub launched DataGrout the sa...
A GitHub Issue. No code, no credentials, no access. Just a paragraph of English that tells an AI agent to copy your private repo into a public comment. That's GitLost, and it works whether the agent runs on Copilot, Claude, Gemini, or Codex. (Noma Security) Noma Security discl...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
OpenClaw tagged v2026.8.1 at 03:30 UTC this morning. The release post counts 933 contributors, 569 of them first-time, and more than 16,000 pull requests, roughly half of every PR ever merged into the project, after a seven-week gap against a prior cadence of 106 releases in 2...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.