Z.ai (formerly Zhipu AI) launched GLM-5.3-Flash on August 26, the first natively multimodal model in the GLM-5 series. The model has 320 billion total parameters but activates only 18 billion per token, with a context window of 1 million tokens. On the DeepSWE v1.1 coding benchmark, GLM-5.3-Flash scored 63.4 compared to 46.2 for its predecessor GLM-5.2. The model runs entirely on domestically produced Chinese AI accelerators, achieving a threefold improvement in end-to-end performance through custom quantization techniques. Weights are open source under the MIT license on Hugging Face and ModelScope.
Source: Read the original article

