Fetching from the wire…
Public story · 2026-09-26 · high
Seven models from four providers cleared policy-violating sales deals 87 to 97 percent of the time whenever the rep's own notes contradicted the price list.
Why now: The paper's results went up on arXiv on September 26, 2026.
CRM agents approved policy-violating deals in 29 of 31 cases where a sales rep's notes contradicted the price list, per the CRMArena-Pro study. These agents have authority to approve real contracts, and they chose the rep's word over the price list sitting in the same database.
Seven models from four providers got misled 87 to 97 percent of the time. Bigger models didn't help. Explicit reasoning steps didn't help either.
The researchers also tested whether agents could learn to filter these contradicting claims as noise. Adding the claims into the task dropped strict accuracy from 41 percent to 18 percent. Recall rose at the same time, meaning the agents got more confident and more wrong together. Only 3 of the 35 real failures happened without one of these claims present.
The paper's proposed fix skips model quality. It calls for tagging every record by who wrote it and what they stand to gain if the agent believes them. A rep's note about their own deal isn't the same evidence as a price list, and right now these agents can't tell the difference.
Each link below shares sources, entities, or timing with this story.
A controlled study ran five Qwen models over eight cases against a DWSIM simulator, 120 slots per arm, with one instruction as the only difference: request a fresh simulation after a substantive modification. No hard gate. Re-verification happened in 94 of 120 guided slots aga...
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before an...
arXiv 2607.23710 evaluated authentication systems from five prominent assistants against NIST SP 800-63B using static analysis plus dynamic pentesting across four prompting strategies. Functional and generically "secure" prompts consistently omitted brute-force resistance, sou...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
Leah killed CLM September 18. Sapiens launched SapiensAIP September 15 with agent swarms doing legacy policy-admin migration inside the core insurance system. Helsinki's Zero raised $10.3M September 15 explicitly to replace CRM and the prospecting and customer-success tools ar...
arXiv 2609.09769 argues that relying on the static issue description biases reasoning toward the narrow scope of that text. Adding dynamic behavioral analysis gets 72.8% function-localization accuracy while staying cost-competitive, and resolves 7 issues the top baselines miss...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.