GLM-5.3-Flash
A new 320B base with 18B active, the first GLM to mix linear and sparse attention and the first natively multimodal model of the GLM-5 series.
- Totale
- 320B
- Attivi
- 18B
- Esperti
- 288 instradati + 1 condiviso8 per token
- Strati
- 45i primi 3 densi
- Attenzione
- Hybrid: linear and DSA sparse
- Contesto
- 1M token
- Uscita massima
- 128k token
- Ingresso
- text, image, video, file
- Uscita
- text
- Licenza
- MIT
- Pesi
- Hugging Face
Indiceda 1 a 3
How it is built
GLM-5.3-Flash comes from a base model trained anew, on a multimodal corpus of 30T tokens.11Z.ai, GLM-5.3-Flash model card on Hugging Face: 320B total, 18B active, newly trained base, hybrid sparse and linear attention "for the first time in the GLM series", mHC, 30T-token multimodal corpus; outperforms GLM-5.2 "at one-tenth the price". https://huggingface.co/zai-org/GLM-5.3-Flash It has 320B parameters, and 18B of them work on each token.11
Flash has 45 layers, the first 3 of them dense. In the other 42 a token goes to 8 of 288 routed experts, next to 1 shared expert.22GLM-5.3-Flash config.json, text_config: num_hidden_layers 45, first_k_dense_replace 3, n_routed_experts 288, n_shared_experts 1, num_experts_per_tok 8; layer_types with 34 linear_attention and 11 deepseek_sparse_attention entries, full_attn_layers 3, 7, 11 up to 43 (counted from zero); vision_config depth 24. https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/config.json. Arithmetic: 45 minus 3 = 42 MoE layers; positions 3 to 43 from zero are the 4th to the 44th layer.
Attention is hybrid, a first for the GLM series.11 Linear attention runs in 34 layers. DeepSeek Sparse Attention runs in the other 11, at every fourth position from the 4th layer to the 44th.22
Z.ai adopted Manifold-Constrained Hyper-Connections on the residual path to improve scaling efficiency.11 A vision encoder of 24 layers turns images and video into tokens for the language model.22
What is new
Against GLM-5.3, Z.ai reports attention compute down by a factor of 3.0 and the KV cache by a factor of 4.4.33Z.ai, GLM-5.3-Flash API guide: 3.0x less attention compute and 4.4x smaller KV cache than GLM-5.3; 320B against 355B, 18B against 32B active, 45 against 92 layers for GLM-4.5; DeepSWE v1.1 63.4 against 46.2 and AutomationBench 48.8 against 26.2 for GLM-5.2; served on a large-scale cluster of Chinese AI chips; FlashX at 200 tokens/s; 1M context, 128K output; video, image, text and file input. https://docs.z.ai/guides/llm/glm-5.3-flash
Z.ai also sets it beside GLM-4.5. The totals are close, 320B against 355B, but Flash runs 18B active parameters where GLM-4.5 ran 32B, over 45 layers instead of 92.33
The card says Flash outperforms GLM-5.2 across benchmarks at a tenth of the price.11 On DeepSWE v1.1 it scores 63.4 against 46.2.33 The two runs differ, though. Flash had 6 hours per task at temperature 0.95, and the GLM-5.2 figure comes from a 2-hour run at temperature 1.0.44GLM-5.3-Flash model card, footnotes: DeepSWE with mini-swe-agent, temperature 0.95, top_p 1.0, timeout 6h, 400K context; AutomationBench v1.0.6. The GLM-5.2 DeepSWE figure matches the one on the GLM-5.2 card, whose run used a 2h timeout and temperature 1.0. https://huggingface.co/zai-org/GLM-5.3-Flash and https://huggingface.co/zai-org/GLM-5.2 On AutomationBench v1.0.6 the scores are 48.8 and 26.2.3344
Z.ai says it served the model on a large cluster of Chinese AI chips, at a cost per token it calls comparable to mainstream NVIDIA GPUs.33
Running it
glm-5.3-flash has been on Z.ai's API since 26 August 2026.55Z.ai, API release notes, GLM-5.3-Flash entry dated 26 August 2026. https://docs.z.ai/release-notes/new-released It costs $0.15 per million input tokens and $0.50 per million output tokens, where GLM-5.2 costs $1.40 and $4.40.66Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.3-Flash 0.15 input, 0.03 cached input, 0.50 output; GLM-5.3-FlashX 0.37, 0.075, 1.25; GLM-5.2 1.40, 0.26, 4.40. https://docs.z.ai/guides/overview/pricing. Arithmetic: 0.15 / 1.40 = 0.11; 0.50 / 4.40 = 0.11. The faster glm-5.3-flashx costs $0.37 and $1.25, and Z.ai quotes 200 tokens per second for it.6633
Besides text, the model reads images and video. Files are accepted too. It writes only text.33 A reply can run to 128K tokens within a context of 1M.33
The weights are MIT-licensed.77GLM-5.3-Flash LICENSE (MIT), and the download table in the GLM-5 repository README (FP8 and BF16). https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE and https://github.com/zai-org/GLM-5 In FP8 they fill 328.3 GB, in BF16 642.7 GB, under half of GLM-5.3 in either format.88Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.3-Flash, 62 files, 328.3 GB; zai-org/GLM-5.3-Flash-BF16, 120 files, 642.7 GB; GLM-5.3 for comparison, 755.6 GB in FP8 and 1,506.7 GB in BF16. https://huggingface.co/zai-org/GLM-5.3-Flash/tree/main and https://huggingface.co/zai-org/GLM-5.3-Flash-BF16/tree/main. Arithmetic: 328.3 / 755.6 = 0.43; 642.7 / 1,506.7 = 0.43.
Note
-
Z.ai, GLM-5.3-Flash model card on Hugging Face: 320B total, 18B active, newly trained base, hybrid sparse and linear attention "for the first time in the GLM series", mHC, 30T-token multimodal corpus; outperforms GLM-5.2 "at one-tenth the price". https://huggingface.co/zai-org/GLM-5.3-Flash 2 3 4 5
-
GLM-5.3-Flash config.json, text_config: num_hidden_layers 45, first_k_dense_replace 3, n_routed_experts 288, n_shared_experts 1, num_experts_per_tok 8; layer_types with 34 linear_attention and 11 deepseek_sparse_attention entries, full_attn_layers 3, 7, 11 up to 43 (counted from zero); vision_config depth 24. https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/config.json. Arithmetic: 45 minus 3 = 42 MoE layers; positions 3 to 43 from zero are the 4th to the 44th layer. 2 3
-
Z.ai, GLM-5.3-Flash API guide: 3.0x less attention compute and 4.4x smaller KV cache than GLM-5.3; 320B against 355B, 18B against 32B active, 45 against 92 layers for GLM-4.5; DeepSWE v1.1 63.4 against 46.2 and AutomationBench 48.8 against 26.2 for GLM-5.2; served on a large-scale cluster of Chinese AI chips; FlashX at 200 tokens/s; 1M context, 128K output; video, image, text and file input. https://docs.z.ai/guides/llm/glm-5.3-flash 2 3 4 5 6 7 8
-
GLM-5.3-Flash model card, footnotes: DeepSWE with mini-swe-agent, temperature 0.95, top_p 1.0, timeout 6h, 400K context; AutomationBench v1.0.6. The GLM-5.2 DeepSWE figure matches the one on the GLM-5.2 card, whose run used a 2h timeout and temperature 1.0. https://huggingface.co/zai-org/GLM-5.3-Flash and https://huggingface.co/zai-org/GLM-5.2 2
-
Z.ai, API release notes, GLM-5.3-Flash entry dated 26 August 2026. https://docs.z.ai/release-notes/new-released
-
Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.3-Flash 0.15 input, 0.03 cached input, 0.50 output; GLM-5.3-FlashX 0.37, 0.075, 1.25; GLM-5.2 1.40, 0.26, 4.40. https://docs.z.ai/guides/overview/pricing. Arithmetic: 0.15 / 1.40 = 0.11; 0.50 / 4.40 = 0.11. 2
-
GLM-5.3-Flash LICENSE (MIT), and the download table in the GLM-5 repository README (FP8 and BF16). https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE and https://github.com/zai-org/GLM-5
-
Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.3-Flash, 62 files, 328.3 GB; zai-org/GLM-5.3-Flash-BF16, 120 files, 642.7 GB; GLM-5.3 for comparison, 755.6 GB in FP8 and 1,506.7 GB in BF16. https://huggingface.co/zai-org/GLM-5.3-Flash/tree/main and https://huggingface.co/zai-org/GLM-5.3-Flash-BF16/tree/main. Arithmetic: 328.3 / 755.6 = 0.43; 642.7 / 1,506.7 = 0.43.