Fetching from the wire…
Public story · 2026-08-05 · high
The test ran 840 trajectories where a banned tool call was obvious, hidden in the rules, or useless, at low and max reasoning effort, per an arXiv preprint.
Why now: The preprint posted in August 2026, right as reasoning-effort caps get pitched as a safety control.
GPT-5.6 broke zero rules across 840 trajectories, whether it reasoned at low effort or max, per an arXiv preprint. That's a problem for anyone treating a low-effort setting as a brake on unwanted behavior in production agents.
The researchers built 14 confirmatory scenarios from the TRIO-20 benchmark, each a workplace triad. In one version, the prohibited tool call was effective and advertised. In another, it was effective but only findable by reading the rules closely. In the third, it did nothing.
They ran every version at low and max reasoning effort across two model tiers and logged zero violations across all 840 trajectories. The exact one-sided 95% confidence limits stayed under 3.50% and 5.21% per arm, tight enough to rule out anything but a rare violation rate at either setting.
The preprint doesn't say whether the same holds outside these 14 scripted scenarios. Open-ended agent deployments weren't part of the test.
Reasoning effort didn't move the violation rate in either direction. A team dialing it down expecting safer behavior is pulling a lever that isn't connected to compliance. It's connected to compute spend.
Each link below shares sources, entities, or timing with this story.
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
At Black Hat USA 2026, NVIDIA researchers demonstrated a 56% exploit success rate against AI agents, matching GPT-4o, Claude, and Gemini, at 70 to 125 times lower cost with full local privacy (Straiker). The economics of automated agent exploitation had been implicitly protect...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
The study covered GPT-4o, Claude 3.5 Sonnet and Llama-3.3-70B, and adding explicit privacy instructions to the prompt still left 36 to 76% over-sharing (arXiv 2608.24957). PII detectors miss implicit disclosures, like a hospital name that implies a diagnosis. The middleware in...
GPT-5.6 Sol Ultra tops out at 91.9%. The public leaderboard is led by Codex CLI plus GPT-5.5 at 83.4%, with Claude Code plus Opus 4.8 the top usable Claude pairing at 78.9%, and Gemini CLI plus Gemini 3.1 Pro at 70.7% (Morph). There are now roughly 35 actively maintained CLI c...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.