Skip the cover

Moonshot AI

Kimi K3

Moonshot AI's open-weight mixture of experts, with 2.8T parameters and 104B of them active for each token.

Total
2.8T
Active
104B
Experts
896 routed + 2 shared16 per token
Layers
93the first one dense
Attention
Hybrid: KDA and Gated MLA
Context
1M tokens
Input
text, image
Output
text
Licence
Kimi K3 License

Contents1 to 3

How it is built

Kimi K3 uses 104B of its 2.8T parameters for each token.11Moonshot AI, Kimi K3 model card, July 2026. Terminal-Bench 2.1 in the Kimi Code harness, reasoning effort max, temperature 1.0, top-p 1.0 for agentic tasks. https://huggingface.co/moonshotai/Kimi-K3 Only the first of its 93 layers is dense.22Moonshot AI, Kimi K3 config.json: first_k_dense_replace 1; full attention layers 4, 8 and every fourth layer to 92, plus 93; MXFP4 with group size 32 and uint8 scales; attention, shared experts, dense MLP, lm_head and vision tower left out of quantization, in bfloat16. https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json The other 92 route every token to 16 of 896 experts, with 2 shared experts always on.11

Linear attention
Attention that keeps a fixed-size state for earlier tokens, instead of a cache that grows with the context.

Most layers, 69 of the 93, use Kimi Delta Attention (KDA), a linear form of attention. The remaining 24 use Gated MLA and sit at every fourth layer, with one more at the top.1122

Linear attention
Attention that keeps a fixed-size state for earlier tokens, instead of a cache that grows with the context.

Attention Residuals (AttnRes) change how depth is used. A layer draws on the representations of earlier layers selectively instead of adding them all up.33Moonshot AI, Kimi K3 blog post: "The full model weights will be released by July 27, 2026." https://www.kimi.ai/blog/kimi-k3 Attention runs at a hidden size of 7,168 with 96 heads. The vocabulary holds 160K tokens.11 Images pass through MoonViT-V2, a vision encoder of 401M parameters.11

Fig. 1
Token

93 layers, the first one dense. In each of the other 92, a token goes to 16 of 896 routed experts and to 2 shared experts. The experts it meets here are simulated, because the real router depends on the trained weights.

What is new

Moonshot launched K3 on 16 July 2026, and it went live the same day in Kimi Code and on the Kimi API.44Moonshot AI (Kimi.ai) on X, 16 July 2026 at 18:58 UTC: "Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API." https://x.com/Kimi_Moonshot/status/2077830229968683203. The Kimi blog index and the Kimi Code changelog also date the release 16 July. The blog post page shows 17 July, the date in Beijing, where it was already 02:58. https://www.kimi.ai/blog and https://www.kimi.com/code/docs/en/kimi-code/whats-new.html

Kimi K2 routed each token to 8 of 384 experts.55Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. https://arxiv.org/abs/2507.20534 K3 routes it to 16 of 896, within what Moonshot calls a Stable LatentMoE framework. Moonshot claims about 2.5 times the overall scaling efficiency of K2.11

The context grew to 1,048,576 tokens, from 256K in Kimi K2.6.1166Moonshot AI, Kimi K2.6 model card: context length 256K, native INT4 quantization. https://huggingface.co/moonshotai/Kimi-K2.6 Quantization-aware training starts at the SFT stage. It uses MXFP4 weights with MXFP8 activations.11 K2.6 used native INT4.66

On Terminal-Bench 2.1 Moonshot reports 88.3, run in its own Kimi Code harness at max effort and temperature 1.0.11

Running it

The weights reached Hugging Face on 27 July, the deadline the blog post had set.3377Hugging Face, moonshotai/Kimi-K3 commit history: "Initial commit" on 27 July 2026 at 13:31 UTC. https://huggingface.co/moonshotai/Kimi-K3/commits/main Under the Kimi K3 License, a company selling the model as a service needs a separate agreement for commercial use once its revenue passes US$20 million over twelve months.88Moonshot AI, Kimi K3 License, sections 2 and 4. https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE Internal use is exempt.88

MXFP4
A floating-point format of four bits per weight, in which each small group of weights shares one scale factor.22

In the checkpoint only the routed experts are MXFP4, in groups of 32 weights that share one 8-bit scale. Attention and the shared experts stay in BF16.22

MXFP4
A floating-point format of four bits per weight, in which each small group of weights shares one scale factor.22

Moonshot recommends supernodes of 64 or more accelerators.33 Thinking cannot be turned off. Effort has three levels, and max is the default.11

On the Kimi API, kimi-k3 has a context of 1,048,576 tokens.99Kimi API Platform, chat model pricing, per million tokens, as listed on 29 September 2026. https://platform.kimi.ai/docs/pricing/chat Input costs $3.00 per million tokens, or $0.30 on a cache hit. Output costs $15.00.99

Notes

  1. Moonshot AI, Kimi K3 model card, July 2026. Terminal-Bench 2.1 in the Kimi Code harness, reasoning effort max, temperature 1.0, top-p 1.0 for agentic tasks. https://huggingface.co/moonshotai/Kimi-K3 2 3 4 5 6 7 8 9 10

  2. Moonshot AI, Kimi K3 config.json: first_k_dense_replace 1; full attention layers 4, 8 and every fourth layer to 92, plus 93; MXFP4 with group size 32 and uint8 scales; attention, shared experts, dense MLP, lm_head and vision tower left out of quantization, in bfloat16. https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json 2 3 4

  3. Moonshot AI, Kimi K3 blog post: "The full model weights will be released by July 27, 2026." https://www.kimi.ai/blog/kimi-k3 2 3

  4. Moonshot AI (Kimi.ai) on X, 16 July 2026 at 18:58 UTC: "Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API." https://x.com/Kimi_Moonshot/status/2077830229968683203. The Kimi blog index and the Kimi Code changelog also date the release 16 July. The blog post page shows 17 July, the date in Beijing, where it was already 02:58. https://www.kimi.ai/blog and https://www.kimi.com/code/docs/en/kimi-code/whats-new.html

  5. Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. https://arxiv.org/abs/2507.20534

  6. Moonshot AI, Kimi K2.6 model card: context length 256K, native INT4 quantization. https://huggingface.co/moonshotai/Kimi-K2.6 2

  7. Hugging Face, moonshotai/Kimi-K3 commit history: "Initial commit" on 27 July 2026 at 13:31 UTC. https://huggingface.co/moonshotai/Kimi-K3/commits/main

  8. Moonshot AI, Kimi K3 License, sections 2 and 4. https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE 2

  9. Kimi API Platform, chat model pricing, per million tokens, as listed on 29 September 2026. https://platform.kimi.ai/docs/pricing/chat 2

More from Moonshot AI

All models by Moonshot AI