Fetching from the wire…
Models2026-09-01 · source-backed
Announced August 31, exposing SFT, DPO, reward finetuning and RL across three surfaces (Fireworks). Harvey post-trained Kimi K3 into "Harvey Tenet" scoring 19.7% all-pass on LAB against 10.8% for the base model at comparable cost. Vercel reports a 93% error-free generation rate for v0 with a 40x end-to-end latency improvement using RFT plus speculative decoding. Heidi Health reached production in four weeks at 3.5x lower latency, and Factory's adapters caught about 70% of real secrets against about 59% for GPT-5.5 at a 5% false-alarm budget. Unusually specific for a launch post, which is why I am repeating the numbers.
Each link below shares sources, entities, or timing with this story.
OpenRouter and Vercel's AI Gateway both list $2.50/M input and $15.00/M output, down from $5.00/$30.00, while OpenAI's own API docs show standard pricing. Vercel's changelog dates it precisely: Aug 17 through Sep 18, all requests through AI Gateway. Gateway-only and time-boxed...
K3 and K3 Fast are now available from Baseten and Fireworks among others, which matters for teams wanting the leading open-weights model without routing prompts to China-based inference. ZDR toggles globally in the dashboard or per-request via zeroDataRetention. Vercel publish...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
At Black Hat USA 2026, NVIDIA researchers demonstrated a 56% exploit success rate against AI agents, matching GPT-4o, Claude, and Gemini, at 70 to 125 times lower cost with full local privacy (Straiker). The economics of automated agent exploitation had been implicitly protect...
Wafer.ai ran K3 at TP8 on a single MI355X node: 952 tok/s aggregate, 118 tok/s single-stream decode, ~13k tok/s steady-state prefill after fixing missing PyTorch sampling functions and AITER MLA prefill kernels via SGLang and ROCm. Against a two-node TP16 B200 deployment that'...
OpenAI's system card classifies GPT-5.3-Codex as "High" for cybersecurity — meaning it can automate end-to-end cyber operations against hardened targets. This is explicitly dual-use: the same capability that makes it an excellent security auditor also makes it an unprecedented...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.