Fetching from the wire…
Infra2026-09-16 · source-backed
AWS published the arithmetic builders need to decide whether caching pays. A 10,000-token document reused across 10 questions nets roughly 75% off total input cost. Minimum checkpoint size is 1,024 tokens for Claude Sonnet 4.5 and 4.6, 4,096 for Opus. Default TTL is 5 minutes, simplified cache management covers about 20 preceding content blocks per checkpoint, and multiple checkpoints must be ordered longest TTL to shortest. Time-to-first-token gains become pronounced past 10,000 tokens.
Each link below shares sources, entities, or timing with this story.
Two concrete recipes for regulated customers who need inference in a single region, not merely in-geography, since cross-Region inference is the throughput-friendly default. Path one: CLAUDE_CODE_USE_MANTLE=1 plus AWS_REGION, pinning models by plain ID, supported in Ireland, S...
Comparing gpt-5.6-luna, terra and sol on Bedrock against gpt-5.4-mini and nano: on AIME, Luna cost $0.0021 per passing answer against Mini's $0.0139 despite Mini's lower nominal token price. On DeepSearchQA multi-turn agent trajectories the gap widened to $0.05 against $0.40 p...
Bedrock invocation logs to S3 carrying model ID, token counts and IAM Identity Center user identity; an Athena view computing per-user daily spend; a Lambda on a 15-minute EventBridge schedule rewriting Customer Managed Policies via iam:CreatePolicyVersion (AWS). Denials take...
Anthropic launched Claude Sonnet 4.6 claiming performance comparable to Opus 4.5 at $3/$15 per million tokens (vs. Opus's $5/$25). SWE-bench Verified: 79.6% (near Opus 4.6's 80.8%). OSWorld-Verified: 72.5% (tied with Opus 4.6's 72.7%). 1M token context window in beta. Now the...
The walkthrough uses utility bills queried at scale as its worked example, and argues for preprocessing before ingestion rather than a better embedding model. That matches what I hit building Rayni: the retrieval was usually fine and the extraction was wrong, and no amount of...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.