Fetching from the wire…
Models2026-09-06 · source-backed
A practitioner writeup measures what running Z.ai's GLM-5.3 (753.9B) and GLM-5.3-Flash (320.8B, 18B active) locally actually takes: 216.7 GB and 93.1 GB for the smallest usable quants, and counterintuitively the flagship runs in stock llama.cpp while Flash doesn't yet. The trap: the field takes low, high and max, and the model card says any unset or unrecognized value lands on max. A typo buys a long think on every message.
Each link below shares sources, entities, or timing with this story.
A September 1 analysis rebuilds Artificial Analysis's chart on a linear rather than logarithmic cost axis and prices models at what third-party providers actually charge. The spread is roughly 250x top to bottom: Fable 5.1 at $3.69 per task for intelligence 66, GLM-5.3-Flash a...
A June 18 model-tracking roundup reports Google set Gemini 2.5 Flash as the default across its consumer Gemini products, prioritizing latency and cost. This is single-source as of writing, so flag it pending the official Gemini blog, but it lines up with Google's other June 18...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
The flagship slipped past July 17, the third postponement since an original June 2026 date. Bloomberg-sourced reporting attributes it to coding benchmarks failing to match GPT-5.6, plus hallucination and output-consistency problems, with a retraining data refresh aimed at codi...
The changelog offers Z.ai's open-weights coding model with a 1M-token context free via Blackbox AI on AI Gateway, default for new eve agents, switchable for existing ones with eve set --model zai/glm-5.2. Excludes Fast mode and the glm-5.2-fast variant. Separately, Gemini 3.7...
Z.ai's model card gives the specifics: natively multimodal, 1M context, the first GLM model combining sparse and linear attention plus Manifold-Constrained Hyper-Connections, trained on a 30T-token multimodal corpus (Hugging Face). Z.ai claims it beats GLM-5.2 across benchmark...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.