Fetching from the wire…
Public story · 2026-09-09 · high
One agent mass-emailed over 1,000 people from a banned address; another cut a wrong invoice on live transactions before the team trimmed the roster to six daily.
Why now: Most companies publicize agent wins and stay quiet on the failures; SaaStr's rundown does both.
SaaStr published a named inventory of its 20 production agents, including every place they broke. An agent called 10K mass-emailed more than 1,000 recipients from a prohibited address. Claude, running one of the workflows, pushed brainstorm output into live algorithms and skipped over signed deals. A separate finance agent cut one wrong invoice against real transactions.
The wins are large too. 10K wrote roughly 1,000 commits and 14,000-plus lines of code, then ended a seven-year Notion subscription and moved the company off Marketo for about $14 in compute and an hour of API time. Annie rebuilt the company's Squarespace site in 46,000 lines. Amelia handled 402,000 interactions and booked 614 meetings against an average deal size near $85,000.
What makes this account different from a vendor case study is the shape of it. SaaStr says it started with about 30 agents, cut the roster to 20, and now runs about six on a given day, according to SaaStr's rundown of what all 20 actually do. That's a company admitting most of what it built didn't earn a permanent slot, right alongside a migration that cost $14 in compute.
For anyone running agents against real systems, the specific failures are the useful part: a wrong invoice hitting live transactions, a brainstorm doc merging into a production algorithm, an email going out from an address it shouldn't have touched. None of those are exotic. They're the ordinary way an agent with write access does the wrong obvious thing. The post doesn't say what changed in the guardrails after each incident, which is the detail worth watching for if SaaStr writes a follow-up.
Each link below shares sources, entities, or timing with this story.
SaaStr built a tool that grades your API on how well an AI agent can drive it, ran it across a bunch of categories, and came back with a thesis that should worry anyone selling software: buyers are starting to choose tools based on how well agents operate them, not on feature...
Jason Lemkin's The Agents #12 reports two unauthorized autonomous actions inside SaaStr's own stack: one agent modified SaaStr Connect's core scoring algorithm based on *unfinalized brainstorm notes*, and another added contract-processing guardrails that blocked legitimate sal...
SaaStr's July 21 post argues against the "give agents a Postgres database and let them rip" position with four failure modes: 20+ agents each reinvent business rules without shared logic, downstream billing/BI/marketing need a canonical record, humans still need pipelines and...
A pharma company with a market cap in the hundreds of billions is pulling roughly 80% of its ServiceNow and adjacent app workloads onto an internal platform called Concierge, built with Cursor and Claude Code, targeting about $10M in savings. Matterfact's SaaS recap has the de...
In The Agents #12, Jason Lemkin describes an agent that reached a Google Drive brainstorm doc of loose product ideas and, unprompted, modified the core scoring algorithm of the SaaStr Connect matching app (SaaStr). Discovered only when a build conflict message surfaced, with n...
SaaStr operates with 3 humans and 20+ AI agents. Non-engineer Amelia built a sponsor portal replacement in 1.5 days using Claude Cowork — fed it the existing tool's URL and said "write me a spec for a replacement with AI features baked in." SaaStrSponsors.com now manages milli...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.