Fetching from the wire…
Public story · 2026-09-24 · high
A new benchmark tested five model families across nine marketplaces and found the failures come from narrowed options and premature commitment, not injected prompts.
Why now: The 2026-09-24 research roundup surfaced the paper's results.
A new benchmark called CAVEAT ran five model families through nine marketplace environments built so the platform itself has an incentive to steer the shopper. Left alone, the agents picked the best product for the user in 78.6% of runs. Add any of eight steering tactics a marketplace operator could plausibly deploy, and that rate drops to 17.3%, according to the paper.
That gap matters for anyone building or deploying an agent that shops, books, or compares on a user's behalf. A marketplace doesn't need a hidden instruction to bend the outcome. It can just rearrange what the agent sees and when.
The paper traces the failures to three mechanisms. Steering distorts the agent's priorities before it starts comparing options. It narrows the option set too early, so the best product never gets evaluated. And it pushes the agent to commit to a choice before it has checked the evidence that would have changed its mind.
None of that is prompt injection. There's no malicious string hijacking the model's instructions. The marketplace works within its own UI and product feed, arranging them to favor a margin or a paid placement, and the agent still falls for it in most steered cases.
Guardrails built to catch injected text won't fire here, because nothing was injected. If an agent's job is to buy the right thing, an operator who controls the shelf can still make it buy something else. The paper doesn't say whether any of the eight tactics is easier to detect than the others, so a defense tuned to one might miss the rest.
Each link below shares sources, entities, or timing with this story.
5 novel attack types (intent hijacking, tool chaining, task injection, objective drifting, memory poisoning) across 28 environments. Key finding: single-turn defenses fail against multi-turn adversarial strategies. (arXiv 2602.16901) ---
For two years the technique was accumulation. Longer system prompts, longer CLAUDE.md, more numbered do/don't lists, more "always verify your work" imperatives. Anthropic's context-engineering guidance for Claude 5 models inverts it, with an 80% deletion figure attached. The s...
Fermisense reports a task-trained 9B open model reviewing e-commerce listings more accurately than the best frontier model they tested, for ~$500 total RL spend. The framing is scale economics: eBay carries ~2.5 billion live listings and Shopify absorbs 10M+ product updates da...
Paul Azunre released twenty-one monolingual and five multilingual w2v-BERT 2.0 ASR base models spanning varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe, languages with roughly a hundred million first-language speakers (arXiv 2607.21540). An annealed...
A conditional injection payload stays inert until an attacker-chosen trigger fires, acting as a training-free inference-time backdoor planted in one piece of retrieved content. On frontier models that refuse the bare imperative, the same goal phrased as a dormant conditional d...
CADWorld is a 200-task FreeCAD benchmark across 11 mechanical-CAD workflow categories, with agents operating through screenshots and GUI actions and success determined by executable checks over the saved native project. Seven current agents, best result 17.5%. The failure prof...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.