Fetching from the wire…
Public story · 2026-09-24 · high
The kit ships shell, file, and web tools with memory and caching built in, and AWS says it beats rival harnesses on token cost using the same models.
Why now: AWS published Strands on strandsagents.com, and harness-cli v0.1.2 followed on September 23 with the MCP client 2.0 swap.
AWS released Strands, an open-source agent harness. It bundles shell, file, and web tools with context management, prompt caching, long-term memory, and delegation to helper agents.
The number that matters is cost. AWS claims Strands runs 28% cheaper in tokens than competing harnesses across six benchmarks on identical models. Against Claude Code specifically, running both on Fable 5, AWS claims a 77% cost cut while Strands scores higher on Terminal Bench 2.1. A harness that cuts token spend by a quarter or more, without touching which model you call, changes the math on running agents at volume instead of hunting for a cheaper model.
It installs with pip install strands-harness or npm install @strands-agents/harness. No model swap required, since it wraps whatever you're already calling.
These are AWS's own benchmarks. The company selling the harness ran the comparison against a named competitor, and no third party has reproduced the numbers yet. Treat 28% as a starting point for testing, not a settled fact. The 77% figure against Claude Code deserves the most skepticism, since it's the number built to headline.
harness-cli v0.1.2 followed on September 23, swapping in MCP client 2.0. A fast point release a day after launch suggests AWS is planning to keep iterating rather than treating this as a one-time drop.
If token cost is what's capping how much agent work you can run, benchmark Strands against your current setup yourself before trusting AWS's numbers.
Each link below shares sources, entities, or timing with this story.
A spec is a press release until someone who didn't write it implements it. GitHub made Agent Plugins 1.0 generally available on August 12 across VS Code, Copilot CLI, the Copilot SDK, and the Copilot app on all plans. The spec, published August 6, was co-authored by AWS, Anysp...
UC Berkeley's Sky Lab put seven models through Claude Code, Codex CLI and Pi on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, 21 model-harness pairs total. HarnessTax is the result, from Melissa Pan, Ion Stoica, Matei Zaharia and co...
Every number you use to pick a harness comes from public repositories the models may have trained on. Specific Labs built the version that doesn't: tasks drawn from licensed private company codebases, including a 200K-user event app and a fintech processing over 100K bank stat...
Runta published FrontierHarness on September 2 and it's the most directly useful benchmark I've read this quarter, because it controls the one variable everyone conflates. Nine agent harnesses (Codex, Claude Code, OpenCode, Pi, Oh My Pi, DeepSeek Harness, Kimi Code, Exo Harnes...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
On June 9, Anthropic released Claude Fable 5 and Mythos 5 across Claude.ai, Claude Code (CLI and web), and Cowork. The spec sheet: 1M-token context, 128K max output, a January 2026 knowledge cutoff, and pricing at $10 input / $50 output per million tokens. That's double Opus 4...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.