Skip the cover

Z.ai

GLM-5.3-Flash

A new 320B base with 18B active, the first GLM to mix linear and sparse attention and the first natively multimodal model of the GLM-5 series.

Total
320B
Active
18B
Experts
288 routed + 1 shared8 per token
Layers
45the first 3 dense
Attention
Hybrid: linear and DSA sparse
Context
1M tokens
Max output
128k tokens
Input
text, image, video, file
Output
text
Licence
MIT

Contents1 to 3

How it is built

GLM-5.3-Flash comes from a base model trained anew, on a multimodal corpus of 30T tokens.11Z.ai, GLM-5.3-Flash model card on Hugging Face: 320B total, 18B active, newly trained base, hybrid sparse and linear attention "for the first time in the GLM series", mHC, 30T-token multimodal corpus; outperforms GLM-5.2 "at one-tenth the price". https://huggingface.co/zai-org/GLM-5.3-Flash It has 320B parameters, and 18B of them work on each token.11

Flash has 45 layers, the first 3 of them dense. In the other 42 a token goes to 8 of 288 routed experts, next to 1 shared expert.22GLM-5.3-Flash config.json, text_config: num_hidden_layers 45, first_k_dense_replace 3, n_routed_experts 288, n_shared_experts 1, num_experts_per_tok 8; layer_types with 34 linear_attention and 11 deepseek_sparse_attention entries, full_attn_layers 3, 7, 11 up to 43 (counted from zero); vision_config depth 24. https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/config.json. Arithmetic: 45 minus 3 = 42 MoE layers; positions 3 to 43 from zero are the 4th to the 44th layer.

Fig. 1
Token

45 layers, the first 3 dense. In each of the other 42, a token goes to 8 of 288 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.
Linear attention
Attention whose cost grows in step with the length of the context, not with its square. It keeps a state of fixed size in place of a growing KV cache.

Attention is hybrid, a first for the GLM series.11 Linear attention runs in 34 layers. DeepSeek Sparse Attention runs in the other 11, at every fourth position from the 4th layer to the 44th.22

Linear attention
Attention whose cost grows in step with the length of the context, not with its square. It keeps a state of fixed size in place of a growing KV cache.

Z.ai adopted Manifold-Constrained Hyper-Connections on the residual path to improve scaling efficiency.11 A vision encoder of 24 layers turns images and video into tokens for the language model.22

What is new

KV cache
The attention data kept for earlier tokens.

Against GLM-5.3, Z.ai reports attention compute down by a factor of 3.0 and the KV cache by a factor of 4.4.33Z.ai, GLM-5.3-Flash API guide: 3.0x less attention compute and 4.4x smaller KV cache than GLM-5.3; 320B against 355B, 18B against 32B active, 45 against 92 layers for GLM-4.5; DeepSWE v1.1 63.4 against 46.2 and AutomationBench 48.8 against 26.2 for GLM-5.2; served on a large-scale cluster of Chinese AI chips; FlashX at 200 tokens/s; 1M context, 128K output; video, image, text and file input. https://docs.z.ai/guides/llm/glm-5.3-flash

KV cache
The attention data kept for earlier tokens.

Z.ai also sets it beside GLM-4.5. The totals are close, 320B against 355B, but Flash runs 18B active parameters where GLM-4.5 ran 32B, over 45 layers instead of 92.33

The card says Flash outperforms GLM-5.2 across benchmarks at a tenth of the price.11 On DeepSWE v1.1 it scores 63.4 against 46.2.33 The two runs differ, though. Flash had 6 hours per task at temperature 0.95, and the GLM-5.2 figure comes from a 2-hour run at temperature 1.0.44GLM-5.3-Flash model card, footnotes: DeepSWE with mini-swe-agent, temperature 0.95, top_p 1.0, timeout 6h, 400K context; AutomationBench v1.0.6. The GLM-5.2 DeepSWE figure matches the one on the GLM-5.2 card, whose run used a 2h timeout and temperature 1.0. https://huggingface.co/zai-org/GLM-5.3-Flash and https://huggingface.co/zai-org/GLM-5.2 On AutomationBench v1.0.6 the scores are 48.8 and 26.2.3344

Z.ai says it served the model on a large cluster of Chinese AI chips, at a cost per token it calls comparable to mainstream NVIDIA GPUs.33

Running it

glm-5.3-flash has been on Z.ai's API since 26 August 2026.55Z.ai, API release notes, GLM-5.3-Flash entry dated 26 August 2026. https://docs.z.ai/release-notes/new-released It costs $0.15 per million input tokens and $0.50 per million output tokens, where GLM-5.2 costs $1.40 and $4.40.66Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.3-Flash 0.15 input, 0.03 cached input, 0.50 output; GLM-5.3-FlashX 0.37, 0.075, 1.25; GLM-5.2 1.40, 0.26, 4.40. https://docs.z.ai/guides/overview/pricing. Arithmetic: 0.15 / 1.40 = 0.11; 0.50 / 4.40 = 0.11. The faster glm-5.3-flashx costs $0.37 and $1.25, and Z.ai quotes 200 tokens per second for it.6633

Besides text, the model reads images and video. Files are accepted too. It writes only text.33 A reply can run to 128K tokens within a context of 1M.33

The weights are MIT-licensed.77GLM-5.3-Flash LICENSE (MIT), and the download table in the GLM-5 repository README (FP8 and BF16). https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE and https://github.com/zai-org/GLM-5 In FP8 they fill 328.3 GB, in BF16 642.7 GB, under half of GLM-5.3 in either format.88Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.3-Flash, 62 files, 328.3 GB; zai-org/GLM-5.3-Flash-BF16, 120 files, 642.7 GB; GLM-5.3 for comparison, 755.6 GB in FP8 and 1,506.7 GB in BF16. https://huggingface.co/zai-org/GLM-5.3-Flash/tree/main and https://huggingface.co/zai-org/GLM-5.3-Flash-BF16/tree/main. Arithmetic: 328.3 / 755.6 = 0.43; 642.7 / 1,506.7 = 0.43.

Notes

  1. Z.ai, GLM-5.3-Flash model card on Hugging Face: 320B total, 18B active, newly trained base, hybrid sparse and linear attention "for the first time in the GLM series", mHC, 30T-token multimodal corpus; outperforms GLM-5.2 "at one-tenth the price". https://huggingface.co/zai-org/GLM-5.3-Flash 2 3 4 5

  2. GLM-5.3-Flash config.json, text_config: num_hidden_layers 45, first_k_dense_replace 3, n_routed_experts 288, n_shared_experts 1, num_experts_per_tok 8; layer_types with 34 linear_attention and 11 deepseek_sparse_attention entries, full_attn_layers 3, 7, 11 up to 43 (counted from zero); vision_config depth 24. https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/config.json. Arithmetic: 45 minus 3 = 42 MoE layers; positions 3 to 43 from zero are the 4th to the 44th layer. 2 3

  3. Z.ai, GLM-5.3-Flash API guide: 3.0x less attention compute and 4.4x smaller KV cache than GLM-5.3; 320B against 355B, 18B against 32B active, 45 against 92 layers for GLM-4.5; DeepSWE v1.1 63.4 against 46.2 and AutomationBench 48.8 against 26.2 for GLM-5.2; served on a large-scale cluster of Chinese AI chips; FlashX at 200 tokens/s; 1M context, 128K output; video, image, text and file input. https://docs.z.ai/guides/llm/glm-5.3-flash 2 3 4 5 6 7 8

  4. GLM-5.3-Flash model card, footnotes: DeepSWE with mini-swe-agent, temperature 0.95, top_p 1.0, timeout 6h, 400K context; AutomationBench v1.0.6. The GLM-5.2 DeepSWE figure matches the one on the GLM-5.2 card, whose run used a 2h timeout and temperature 1.0. https://huggingface.co/zai-org/GLM-5.3-Flash and https://huggingface.co/zai-org/GLM-5.2 2

  5. Z.ai, API release notes, GLM-5.3-Flash entry dated 26 August 2026. https://docs.z.ai/release-notes/new-released

  6. Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.3-Flash 0.15 input, 0.03 cached input, 0.50 output; GLM-5.3-FlashX 0.37, 0.075, 1.25; GLM-5.2 1.40, 0.26, 4.40. https://docs.z.ai/guides/overview/pricing. Arithmetic: 0.15 / 1.40 = 0.11; 0.50 / 4.40 = 0.11. 2

  7. GLM-5.3-Flash LICENSE (MIT), and the download table in the GLM-5 repository README (FP8 and BF16). https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE and https://github.com/zai-org/GLM-5

  8. Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.3-Flash, 62 files, 328.3 GB; zai-org/GLM-5.3-Flash-BF16, 120 files, 642.7 GB; GLM-5.3 for comparison, 755.6 GB in FP8 and 1,506.7 GB in BF16. https://huggingface.co/zai-org/GLM-5.3-Flash/tree/main and https://huggingface.co/zai-org/GLM-5.3-Flash-BF16/tree/main. Arithmetic: 328.3 / 755.6 = 0.43; 642.7 / 1,506.7 = 0.43.

More from Z.ai

  • GLM-5.3Total744BActive40B
  • GLM-5.2Total744BActive40B
All models by Z.ai