GLM-5.3
Z.ai's flagship, a 744B mixture of experts with 40B active, post-trained from the GLM-5.2 base for coding and long agent tasks.
- Total
- 744B
- Active
- 40B
- Experts
- 256 routed + 1 shared8 per token
- Layers
- 78the first 3 dense
- Attention
- DeepSeek Sparse Attention, IndexShare
- Context
- 1M tokens
- Max output
- 128k tokens
- Input
- text
- Output
- text
- Licence
- GLM-5.3 License
- Weights
- Hugging Face
Contents1 to 3
How it is built
GLM-5.3 starts from the same base model as GLM-5.2. Z.ai says every gain comes from post-training.11Z.ai, GLM-5.3 model card on Hugging Face: "the same base model as GLM-5.2", every gain from post-training; reasoning_effort defaults to max; benchmark table and footnotes (Terminal-Bench 3.0: Claude Code 2.1.207, avg@3; CyberGym: Claude Code 2.1.207, max reasoning effort, no web tools, single-run Pass@1 over 1,507 tasks). https://huggingface.co/zai-org/GLM-5.3 That base is a mixture of experts of 744B parameters, with 40B active for each token.22Z.ai, GLM-5 repository on GitHub, README, download table: GLM-5.3 at 744B-A40B in FP8, GLM-5.3-BF16 at 744B-A40B in BF16. https://github.com/zai-org/GLM-5. Hugging Face counts 753.33B parameters in the weight files; Z.ai does not explain the gap.
Of the 78 layers, the first 3 are dense. The other 75 carry 256 routed experts and 1 shared expert each, and the router keeps 8 per token.33GLM-5.3 config.json: num_hidden_layers 78, first_k_dense_replace 3, n_routed_experts 256, n_shared_experts 1, num_experts_per_tok 8, scoring_func sigmoid, topk_method noaux_tc, index_topk 2048, indexer_types with 21 "full" and 57 "shared" entries. https://huggingface.co/zai-org/GLM-5.3/blob/main/config.json. Arithmetic: 78 minus 3 = 75 MoE layers. It scores experts with a sigmoid. A bias on each score balances the load, with no auxiliary loss.33
Attention is DeepSeek Sparse Attention: an indexer picks 2,048 earlier tokens for each query, and attention reads only those.33 GLM-5.3 keeps IndexShare from GLM-5.2, which reuses one indexer across every four sparse attention layers.44Z.ai, GLM-5.2 model card on Hugging Face (IndexShare reuses the same indexer across every four sparse attention layers; paper arXiv 2603.12201). https://huggingface.co/zai-org/GLM-5.2
What is new
Thinking can no longer be switched off through the API. Effort runs from low to max, and max is the default.55Z.ai, GLM-5.3 API guide: 1M context, 128K maximum output, text-only input, reasoning always on with low, high and max effort. https://docs.z.ai/guides/llm/glm-5.3 11
On Terminal-Bench 3.0, averaged over three runs per task, GLM-5.2 scored 4.6. GLM-5.3 scores 28.3.11 On CyberGym Z.ai reports 84.5, against 77.2 for GLM-5.2, in the Claude Code 2.1.207 harness at maximum effort with no web tools. Each of the 1,507 tasks got one attempt.11
On 17 September 2026 the Center for AI Standards and Innovation (CAISI) at NIST called GLM-5.3 "the most cyber-capable open-weight model released to date".66NIST, "CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities", 17 September 2026. https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities Its evaluators also put the model about four months behind US frontier models on aggregate cyber tasks.66
Running it
Z.ai announced GLM-5.3 on 14 August 2026 and opened it that day to its GLM Coding Plan and ZCode. API access and open weights were to follow in stages, after safety evaluations.77Z.ai on X, 14 August 2026 at 05:17 UTC: "GLM-5.3 is available now through GLM Coding Plan and ZCode. API access and open weights will be released in stages following rigorous safety evaluations." https://x.com/Zai_org/status/2088132969630212372. CAISI also gives 14 August as the release date. https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities The API opened on 18 August, as glm-5.3.88Z.ai, API release notes, GLM-5.3 entry dated 18 August 2026. https://docs.z.ai/release-notes/new-released It takes text only, with 1M tokens of context and up to 128K of output.55 Input costs $1.40 per million tokens and output $4.40.99Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.3 1.40 input, 0.26 cached input, 4.40 output. https://docs.z.ai/guides/overview/pricing
The FP8 weights on Hugging Face come to 755.6 GB, and the BF16 copy to 1,506.7 GB.221010Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.3, 141 files, 755.6 GB; zai-org/GLM-5.3-BF16, 282 files, 1,506.7 GB. https://huggingface.co/zai-org/GLM-5.3/tree/main and https://huggingface.co/zai-org/GLM-5.3-BF16/tree/main That is memory for the weights alone, before any KV cache.
The GLM-5.3 License grants MIT-style rights plus one clause for the largest users.1111Z.ai, GLM-5.3 License, clause 2. https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE It covers a licensee whose group runs a Model as a Service business, such as inference sold through an API, and earns more than US$10 billion over any 12 consecutive months.1111 Such a licensee must pass Z.ai's security review before any commercial use.1111 GLM-5.3-Flash and GLM-5.2 stay under plain MIT.1212MIT License files of GLM-5.3-Flash and GLM-5.2. https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE and https://huggingface.co/zai-org/GLM-5.2
Notes
-
Z.ai, GLM-5.3 model card on Hugging Face: "the same base model as GLM-5.2", every gain from post-training; reasoning_effort defaults to max; benchmark table and footnotes (Terminal-Bench 3.0: Claude Code 2.1.207, avg@3; CyberGym: Claude Code 2.1.207, max reasoning effort, no web tools, single-run Pass@1 over 1,507 tasks). https://huggingface.co/zai-org/GLM-5.3 2 3 4
-
Z.ai, GLM-5 repository on GitHub, README, download table: GLM-5.3 at 744B-A40B in FP8, GLM-5.3-BF16 at 744B-A40B in BF16. https://github.com/zai-org/GLM-5. Hugging Face counts 753.33B parameters in the weight files; Z.ai does not explain the gap. 2
-
GLM-5.3 config.json: num_hidden_layers 78, first_k_dense_replace 3, n_routed_experts 256, n_shared_experts 1, num_experts_per_tok 8, scoring_func sigmoid, topk_method noaux_tc, index_topk 2048, indexer_types with 21 "full" and 57 "shared" entries. https://huggingface.co/zai-org/GLM-5.3/blob/main/config.json. Arithmetic: 78 minus 3 = 75 MoE layers. 2 3
-
Z.ai, GLM-5.2 model card on Hugging Face (IndexShare reuses the same indexer across every four sparse attention layers; paper arXiv 2603.12201). https://huggingface.co/zai-org/GLM-5.2
-
Z.ai, GLM-5.3 API guide: 1M context, 128K maximum output, text-only input, reasoning always on with low, high and max effort. https://docs.z.ai/guides/llm/glm-5.3 2
-
NIST, "CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities", 17 September 2026. https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities 2
-
Z.ai on X, 14 August 2026 at 05:17 UTC: "GLM-5.3 is available now through GLM Coding Plan and ZCode. API access and open weights will be released in stages following rigorous safety evaluations." https://x.com/Zai_org/status/2088132969630212372. CAISI also gives 14 August as the release date. https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities
-
Z.ai, API release notes, GLM-5.3 entry dated 18 August 2026. https://docs.z.ai/release-notes/new-released
-
Z.ai, API pricing, USD per million tokens, read 29 September 2026: GLM-5.3 1.40 input, 0.26 cached input, 4.40 output. https://docs.z.ai/guides/overview/pricing
-
Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: zai-org/GLM-5.3, 141 files, 755.6 GB; zai-org/GLM-5.3-BF16, 282 files, 1,506.7 GB. https://huggingface.co/zai-org/GLM-5.3/tree/main and https://huggingface.co/zai-org/GLM-5.3-BF16/tree/main
-
Z.ai, GLM-5.3 License, clause 2. https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE 2 3
-
MIT License files of GLM-5.3-Flash and GLM-5.2. https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE and https://huggingface.co/zai-org/GLM-5.2
More from Z.ai
- GLM-5.3-FlashTotal320BActive18B
- GLM-5.2Total744BActive40B