GLM-5.3-Flash ships as 320B total / 18B active under MIT with a hybrid sparse-plus-linear attention stack and a 30T-token multimodal corpus
Z.ai's model card, live on Hugging Face since 2026-08-25 with 1,179 likes, gives the specifics behind the Ox Alpha reveal: 320B total parameters, 18B active, natively multimodal, 1M context, MIT license, and the first GLM model to combine sparse and linear attention plus Manifold-Constrained Hyper-Connections for scaling efficiency. Z.ai claims it beats GLM-5.2 across benchmarks at one tenth the price and approaches Claude Opus 4.8 on coding and agentic work. The card also documents a `reasoning_effort` parameter with low/high/max levels that defaults to max, and warns that `clear_thinking` defaults to false and should be set true for chat, which are the two settings most likely to blow up a builder's token bill on first use.
↳ Follow the thread