Fetching from the wire…
Agents2026-08-06 · source-backed
arXiv 2608.04741 uses a fuzzing-inspired process to generate page-specific injections that make logging in look like a plausible prerequisite for continuing the user's task, steering the agent to a controlled login page. Task-agnostic, black-box, no knowledge of the user's goal or agent internals needed, effective across architectures and existing defenses. Credential entry is the one action a browser agent should never take autonomously, and no shipped agent currently treats it as a distinct trust boundary. That's a design gap you could fix in your own harness this afternoon.
Each link below shares sources, entities, or timing with this story.
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
It routes each agent's action decisions and temporal constraints to a lightweight digital-twin server running a training-free rule-based orchestrator, with a constrained POMDP optimized by PPO-Lagrangian (arXiv:2607.09330). Task success comparable to conventional coordination,...
On June 10, Ory announced Agent DX, exposing its open-source identity platform to coding agents via plugins for Claude Code, OpenAI Codex, and others, so agents handle auth flows, sessions, and secrets as first-class objects instead of ad-hoc credential plumbing. It's a single...
SARC-DQ found competent agents converted freshness/lineage/provenance defects into costly actions about 60% of the time, with both data-quality flags and the agents' own hedging detecting them at chance. The conversion rate was flat across four model tiers spanning a 15x price...
arXiv 2607.26935 argues the human-vs-bot label space can't represent agent traffic: an MLP binary classifier misroutes 39.1% of real agent sessions as human, a SAINT transformer 34.5%, while adding an explicit third class yields agent F1 = 1.000 across all 30 runs. Against a f...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.