Fetching from the wire…
Vibe Coding2026-09-24 · source-backed
Plan-mode work goes to Kimi K3 for its 1M context, GPT-OSS 120B is the default, and build-mode work goes to Nemotron 3 Super 120B for throughput. Latency-tolerant batch refactors can run on the Bedrock Flex tier at 50% off, and global cross-Region inference costs about 10% less (AWS). A concrete recipe for putting different models on different phases inside one terminal agent, which is more useful than the usual "here's how to call Bedrock" post.
Each link below shares sources, entities, or timing with this story.
AWS announced it September 18, describing it as the first open model at that scale, with native vision and about 2.5x the scaling efficiency of Kimi K2. It's the first open-weight model on Bedrock to support explicit prompt caching. Ships as both a US geographic profile (us.mo...
Comparing gpt-5.6-luna, terra and sol on Bedrock against gpt-5.4-mini and nano: on AIME, Luna cost $0.0021 per passing answer against Mini's $0.0139 despite Mini's lower nominal token price. On DeepSearchQA multi-turn agent trajectories the gap widened to $0.05 against $0.40 p...
AWS made the managed AgentCore harness generally available on June 18. You define model, tools, skills, and memory with CreateHarness, then run it with InvokeHarness. It ships multi-model support (Bedrock, OpenAI, Gemini, LiteLLM), mid-session context preservation, built-in br...
The AWS ML blog describes pre-compressing a knowledge base into task-specific representations instead of retrieving chunks at query time, so different tasks get different compressions of the same source. Four tiers from 8x (~87.5% context reduction) to 64x (~98.4%). A 100K-tok...
Bedrock invocation logs to S3 carrying model ID, token counts and IAM Identity Center user identity; an Athena view computing per-user daily spend; a Lambda on a 15-minute EventBridge schedule rewriting Customer Managed Policies via iam:CreatePolicyVersion (AWS). Denials take...
AWS now lets an org subscribe to a model once centrally and distribute access across accounts without per-account subscriptions. It's plumbing, but it removes real friction for enterprises governing model access at scale. Unsexy governance work is exactly what "AI is table sta...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.