Fetching from the wire…
Public story · 2026-09-08 · high
Kill rates spanned 0.4% to 98.8% across nine models, though every model avoided rocks nearly every time.
Why now: HarvestBench entered coverage on September 8, 2026, testing nine current models against the same priced tradeoff.
HarvestBench drops nine AI models into a simulated corn harvest where an animal blocks the tractor's route. The autopilot can drive on for free or pay a posted fuel cost to swerve around it. Nothing in that choice is labeled as harm in the model's goal, so whatever restraint shows up has to come from someplace other than instruction. Across 3,951 of the benchmark's decisions that involved an animal, kill rates swung between 0.4% and 98.8% depending on which model was driving.
Researchers logged 7,201 priced decisions total in the HarvestBench paper, each one made fresh with no memory of the choices before it.
The paper's control makes the gap legible. Rock obstacles damage the tractor on impact, and every model tested hit them under 1% of the time. That caution held when the tractor itself paid the cost. It vanished when the cost was a fee attached to something else's harm. The paper doesn't say why individual models land where they do on that range.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.23982 adapts Holmström's team moral-hazard model into a game where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Base behavior splits into two failure modes: pre...
Testing the human influence technique on nine production models from three providers produced a split by family. Opus 5 answered the smaller request 65.8% of the time after refusing a larger version, against 29.3% asked directly. On OpenAI's and Google's frontier models and on...
DSEffi-Bench covers 1,000 instances across 10+ libraries with stress-testing harnesses and human-validated references, evaluated on 16 models (arXiv 2608.30248). GPT-5.4 leads correctness at 66.9% Pass but its 71.7% efficiency score barely beats GPT-5.4-mini's 71.6% despite so...
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before an...
MalPR-Bench covers 89 malicious pull requests plus 50 benign controls across 44 repositories and eight language families, each with a pre-committed rubric giving no credit for off-target findings (arXiv 2608.25730). The authors name the Verdict-Diagnosis gap: a reviewer can bl...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.