Fetching from the wire…
Public story · 2026-07-21 · high
The launch also confirms DeepMind has started pre-training Gemini 4, its most ambitious training run yet.
Why now: Google shipped both models on July 21 and used the same announcement to confirm Gemini 4 pre-training is underway.
Google shipped Gemini 3.6 Flash on July 21, priced at $1.50 per million input tokens and $7.50 per million output tokens. DeepSWE code precision rose from 37% to 49%, and OSWorld-Verified computer-use accuracy climbed to 83% from 78.4%, the report says. The knowledge cutoff also moved off January 2025, to March 2026, closing a real gap for anyone building on top of it.
The number worth tracking is the 17% drop in output tokens versus 3.5 Flash, per 9to5Google. Benchmarks measure how smart a model is. Token count measures what it costs to run, and that cut applies to every call already running through Flash, no prompt changes required.
Pricing tells a sharper story on the cheap end. Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens. It beats the full 3 Flash on SWE-Bench Pro, 54.2% versus 49.6%, per 9to5Google. The cheap tier now beats last generation's mid tier on a real coding benchmark. If you route work by cost, the bar for good-enough coding performance dropped in price and rose in capability at the same time.
A 3.5 Flash Cyber variant, restricted to government and trusted partners, handles vulnerability detection and patching, 9to5Google reports. The report doesn't say what qualifies a partner or when that access might widen. The bigger signal sits in that footnote: DeepMind has begun what Google calls its most ambitious pre-training run yet, for Gemini 4. No date, no scale details surfaced. Flash 3 through 3.6 shipped while the actual next generation was already underway behind it.
Each link below shares sources, entities, or timing with this story.
Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, the last purpose-built to find and patch vulnerabilities and pitched as a cheap alternative to large security-specialized models like Mythos (DeepMind). Splitting a cheap tier into a security SKU is new packa...
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tok...
Willison's August 13 release adds Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite plus gemini-embedding-2 and -001, rebuilding on LLM 0.32's structured message and streaming APIs so reasoning, tool calls and results emit as typed stream events while preserving Gemini thought si...
"Gemini is Cooked but GCP is Cooking" argues Google quietly shelved 3.5 Pro, which industry chatter placed at roughly Opus 4.5 level, shipping Gemini 3.6 Flash as a bridge the authors call worse than Muse Spark 1.2, Grok 4.5, and tier-1 Chinese open-source models. The hard num...
The flagship slipped past July 17, the third postponement since an original June 2026 date. Bloomberg-sourced reporting attributes it to coding benchmarks failing to match GPT-5.6, plus hallucination and output-consistency problems, with a retraining data refresh aimed at codi...
Arena Mode is the standout. Run two agents on the same prompt with hidden model identities, vote on which performs better. 40K+ votes so far. Key finding from the community leaderboard: speed beats intelligence in most real-world tasks (Gemini 3 Flash beats Gemini 3 Pro). Plan...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.