Skip the cover

Z.ai

GLM-5.2

A 744B mixture of experts with 40B active that took the GLM-5 line to a 1M-token context and shared its attention indexer across layers.

Total
744B
Active
40B
Experts
256 routed + 1 shared8 per token
Layers
78the first 3 dense
Attention
DeepSeek Sparse Attention, IndexShare
Context
1M tokens
Max output
128k tokens
Input
text
Output
text
Licence
MIT

Contents1 to 3

How it is built

GLM-5.2 keeps the size set by GLM-5: 744B parameters, with 40B active per token.11Z.ai, GLM-5.2 model card on Hugging Face: 1M-token context; IndexShare, per-token FLOPs cut by 2.9x at a 1M context (arXiv 2603.12201); MTP acceptance length up to 20% longer; multiple thinking effort levels; SWE-bench Pro 62.1 against 58.4 for GLM-5.1, run with OpenHands, temperature 1, top_p 1, max_new_tokens 32k, 400K context; MIT licence with "no regional limits". https://huggingface.co/zai-org/GLM-5.2 22Z.ai, GLM-5 repository on GitHub, README: GLM-5 scales from 355B parameters (32B active) to 744B (40B active), pre-training data from 23T to 28.5T tokens, integrates DeepSeek Sparse Attention, "largely reducing deployment cost"; download table with GLM-5.2 in BF16 and GLM-5.2-FP8 at 744B-A40B. https://github.com/zai-org/GLM-5 GLM-5 had grown from the 355B of GLM-4.5, which ran 32B per token. It also raised pre-training data from 23T to 28.5T tokens.22

Three dense layers open the stack of 78. The 75 below them are MoE layers with 256 routed experts and 1 shared expert, and a sigmoid router keeps 8 experts per token.33GLM-5.2 config.json: num_hidden_layers 78, first_k_dense_replace 3, n_routed_experts 256, n_shared_experts 1, num_experts_per_tok 8, scoring_func sigmoid, index_topk 2048, indexer_types with 21 "full" and 57 "shared" entries. https://huggingface.co/zai-org/GLM-5.2/blob/main/config.json. Arithmetic: 78 minus 3 = 75 MoE layers.

Fig. 1
Token

78 layers, the first 3 dense. In each of the other 75, a token goes to 8 of 256 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.
Indexer
A small scoring network that decides which earlier tokens an attention step reads.

GLM-5 also brought in DeepSeek Sparse Attention, which Z.ai credits with a large cut in deployment cost.22 For each query an indexer chooses 2,048 earlier tokens, and attention reads those alone.33

Indexer
A small scoring network that decides which earlier tokens an attention step reads.

What is new

The context reached 1M tokens.44Z.ai, GLM-5.2 API guide: 1M context, 128K maximum output, text input and output. https://docs.z.ai/guides/llm/glm-5.2 GLM-5 and GLM-5.1 stopped at 202,752 positions.55config.json of GLM-5 and GLM-5.1, max_position_embeddings 202752. https://huggingface.co/zai-org/GLM-5/blob/main/config.json and https://huggingface.co/zai-org/GLM-5.1/blob/main/config.json

IndexShare lets one indexer serve four sparse attention layers.11 In the config, 21 layers compute an indexer and 57 reuse one.33 Z.ai puts the saving at 2.9 times fewer operations per token at a 1M context.11

Speculative decoding
A cheap draft guesses several tokens ahead and the full model checks them in one pass. The acceptance length counts how many guesses survive.

Z.ai also reworked the multi-token prediction layer, which drafts tokens for speculative decoding. The acceptance length grew by up to 20%.11

Speculative decoding
A cheap draft guesses several tokens ahead and the full model checks them in one pass. The acceptance length counts how many guesses survive.

On SWE-bench Pro Z.ai reports 62.1, against 58.4 for GLM-5.1. The run used OpenHands with a 400K context and up to 32K new tokens per reply.11 Several thinking effort levels trade performance against latency.11

GLM-5.3 was post-trained from this same base model, and the GLM-5.2 card now points to it as the newer version.66Z.ai, GLM-5.3 model card ("the same base model as GLM-5.2"), and the GLM-5.2 card metadata, new_version: zai-org/GLM-5.3-BF16. https://huggingface.co/zai-org/GLM-5.3 and https://huggingface.co/zai-org/GLM-5.2

Running it

glm-5.2 reached Z.ai's API on 16 June 2026.77Z.ai, API release notes, GLM-5.2 entry dated 16 June 2026. https://docs.z.ai/release-notes/new-released It reads text only. The context holds 1M tokens, and a reply up to 128K.44 Prices match GLM-5.3: $1.40 per million input tokens and $4.40 per million output tokens.88Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.2 and GLM-5.3 both 1.40 input, 0.26 cached input, 4.40 output. https://docs.z.ai/guides/overview/pricing

Z.ai publishes 1,506.7 GB of weights in BF16, and a second repository holds an FP8 copy of 755.6 GB.2299Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.2, 282 files, 1,506.7 GB; zai-org/GLM-5.2-FP8, 141 files, 755.6 GB. https://huggingface.co/zai-org/GLM-5.2/tree/main and https://huggingface.co/zai-org/GLM-5.2-FP8/tree/main

The licence is MIT, and the card adds that it has "no regional limits".11 GLM-5.3 later moved to a licence of its own. It asks for a security review from Model as a Service groups with revenue above US$10 billion over 12 consecutive months.1010Z.ai, GLM-5.3 License, clause 2. https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE

Notes

  1. Z.ai, GLM-5.2 model card on Hugging Face: 1M-token context; IndexShare, per-token FLOPs cut by 2.9x at a 1M context (arXiv 2603.12201); MTP acceptance length up to 20% longer; multiple thinking effort levels; SWE-bench Pro 62.1 against 58.4 for GLM-5.1, run with OpenHands, temperature 1, top_p 1, max_new_tokens 32k, 400K context; MIT licence with "no regional limits". https://huggingface.co/zai-org/GLM-5.2 2 3 4 5 6 7

  2. Z.ai, GLM-5 repository on GitHub, README: GLM-5 scales from 355B parameters (32B active) to 744B (40B active), pre-training data from 23T to 28.5T tokens, integrates DeepSeek Sparse Attention, "largely reducing deployment cost"; download table with GLM-5.2 in BF16 and GLM-5.2-FP8 at 744B-A40B. https://github.com/zai-org/GLM-5 2 3 4

  3. GLM-5.2 config.json: num_hidden_layers 78, first_k_dense_replace 3, n_routed_experts 256, n_shared_experts 1, num_experts_per_tok 8, scoring_func sigmoid, index_topk 2048, indexer_types with 21 "full" and 57 "shared" entries. https://huggingface.co/zai-org/GLM-5.2/blob/main/config.json. Arithmetic: 78 minus 3 = 75 MoE layers. 2 3

  4. Z.ai, GLM-5.2 API guide: 1M context, 128K maximum output, text input and output. https://docs.z.ai/guides/llm/glm-5.2 2

  5. config.json of GLM-5 and GLM-5.1, max_position_embeddings 202752. https://huggingface.co/zai-org/GLM-5/blob/main/config.json and https://huggingface.co/zai-org/GLM-5.1/blob/main/config.json

  6. Z.ai, GLM-5.3 model card ("the same base model as GLM-5.2"), and the GLM-5.2 card metadata, new_version: zai-org/GLM-5.3-BF16. https://huggingface.co/zai-org/GLM-5.3 and https://huggingface.co/zai-org/GLM-5.2

  7. Z.ai, API release notes, GLM-5.2 entry dated 16 June 2026. https://docs.z.ai/release-notes/new-released

  8. Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.2 and GLM-5.3 both 1.40 input, 0.26 cached input, 4.40 output. https://docs.z.ai/guides/overview/pricing

  9. Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.2, 282 files, 1,506.7 GB; zai-org/GLM-5.2-FP8, 141 files, 755.6 GB. https://huggingface.co/zai-org/GLM-5.2/tree/main and https://huggingface.co/zai-org/GLM-5.2-FP8/tree/main

  10. Z.ai, GLM-5.3 License, clause 2. https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE

More from Z.ai

All models by Z.ai