Kimi K3
Moonshot AI's open-weight mixture of experts, with 2.8T parameters and 104B of them active for each token.
- Totale
- 2,8T
- Attivi
- 104B
- Esperti
- 896 instradati + 2 condivisi16 per token
- Strati
- 93il primo denso
- Attenzione
- Hybrid: KDA and Gated MLA
- Contesto
- 1M token
- Ingresso
- text, image
- Uscita
- text
- Licenza
- Kimi K3 License
- Pesi
- Hugging Face
Indiceda 1 a 3
How it is built
Kimi K3 uses 104B of its 2.8T parameters for each token.11Moonshot AI, Kimi K3 model card, July 2026. Terminal-Bench 2.1 in the Kimi Code harness, reasoning effort max, temperature 1.0, top-p 1.0 for agentic tasks. https://huggingface.co/moonshotai/Kimi-K3 Only the first of its 93 layers is dense.22Moonshot AI, Kimi K3 config.json: first_k_dense_replace 1; full attention layers 4, 8 and every fourth layer to 92, plus 93; MXFP4 with group size 32 and uint8 scales; attention, shared experts, dense MLP, lm_head and vision tower left out of quantization, in bfloat16. https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json The other 92 route every token to 16 of 896 experts, with 2 shared experts always on.11
Most layers, 69 of the 93, use Kimi Delta Attention (KDA), a linear form of attention. The remaining 24 use Gated MLA and sit at every fourth layer, with one more at the top.1122
Attention Residuals (AttnRes) change how depth is used. A layer draws on the representations of earlier layers selectively instead of adding them all up.33Moonshot AI, Kimi K3 blog post: "The full model weights will be released by July 27, 2026." https://www.kimi.ai/blog/kimi-k3 Attention runs at a hidden size of 7,168 with 96 heads. The vocabulary holds 160K tokens.11 Images pass through MoonViT-V2, a vision encoder of 401M parameters.11
What is new
Moonshot launched K3 on 16 July 2026, and it went live the same day in Kimi Code and on the Kimi API.44Moonshot AI (Kimi.ai) on X, 16 July 2026 at 18:58 UTC: "Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API." https://x.com/Kimi_Moonshot/status/2077830229968683203. The Kimi blog index and the Kimi Code changelog also date the release 16 July. The blog post page shows 17 July, the date in Beijing, where it was already 02:58. https://www.kimi.ai/blog and https://www.kimi.com/code/docs/en/kimi-code/whats-new.html
Kimi K2 routed each token to 8 of 384 experts.55Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. https://arxiv.org/abs/2507.20534 K3 routes it to 16 of 896, within what Moonshot calls a Stable LatentMoE framework. Moonshot claims about 2.5 times the overall scaling efficiency of K2.11
The context grew to 1,048,576 tokens, from 256K in Kimi K2.6.1166Moonshot AI, Kimi K2.6 model card: context length 256K, native INT4 quantization. https://huggingface.co/moonshotai/Kimi-K2.6 Quantization-aware training starts at the SFT stage. It uses MXFP4 weights with MXFP8 activations.11 K2.6 used native INT4.66
On Terminal-Bench 2.1 Moonshot reports 88.3, run in its own Kimi Code harness at max effort and temperature 1.0.11
Running it
The weights reached Hugging Face on 27 July, the deadline the blog post had set.3377Hugging Face, moonshotai/Kimi-K3 commit history: "Initial commit" on 27 July 2026 at 13:31 UTC. https://huggingface.co/moonshotai/Kimi-K3/commits/main Under the Kimi K3 License, a company selling the model as a service needs a separate agreement for commercial use once its revenue passes US$20 million over twelve months.88Moonshot AI, Kimi K3 License, sections 2 and 4. https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE Internal use is exempt.88
In the checkpoint only the routed experts are MXFP4, in groups of 32 weights that share one 8-bit scale. Attention and the shared experts stay in BF16.22
Moonshot recommends supernodes of 64 or more accelerators.33 Thinking cannot be turned off. Effort has three levels, and max is the default.11
On the Kimi API, kimi-k3 has a context of 1,048,576 tokens.99Kimi API Platform, chat model pricing, per million tokens, as listed on 29 September 2026. https://platform.kimi.ai/docs/pricing/chat Input costs $3.00 per million tokens, or $0.30 on a cache hit. Output costs $15.00.99
Note
-
Moonshot AI, Kimi K3 model card, July 2026. Terminal-Bench 2.1 in the Kimi Code harness, reasoning effort max, temperature 1.0, top-p 1.0 for agentic tasks. https://huggingface.co/moonshotai/Kimi-K3 2 3 4 5 6 7 8 9 10
-
Moonshot AI, Kimi K3 config.json: first_k_dense_replace 1; full attention layers 4, 8 and every fourth layer to 92, plus 93; MXFP4 with group size 32 and uint8 scales; attention, shared experts, dense MLP, lm_head and vision tower left out of quantization, in bfloat16. https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json 2 3 4
-
Moonshot AI, Kimi K3 blog post: "The full model weights will be released by July 27, 2026." https://www.kimi.ai/blog/kimi-k3 2 3
-
Moonshot AI (Kimi.ai) on X, 16 July 2026 at 18:58 UTC: "Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API." https://x.com/Kimi_Moonshot/status/2077830229968683203. The Kimi blog index and the Kimi Code changelog also date the release 16 July. The blog post page shows 17 July, the date in Beijing, where it was already 02:58. https://www.kimi.ai/blog and https://www.kimi.com/code/docs/en/kimi-code/whats-new.html
-
Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. https://arxiv.org/abs/2507.20534
-
Moonshot AI, Kimi K2.6 model card: context length 256K, native INT4 quantization. https://huggingface.co/moonshotai/Kimi-K2.6 2
-
Hugging Face, moonshotai/Kimi-K3 commit history: "Initial commit" on 27 July 2026 at 13:31 UTC. https://huggingface.co/moonshotai/Kimi-K3/commits/main
-
Moonshot AI, Kimi K3 License, sections 2 and 4. https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE 2
-
Kimi API Platform, chat model pricing, per million tokens, as listed on 29 September 2026. https://platform.kimi.ai/docs/pricing/chat 2
Altri modelli di Moonshot AI
- Kimi K2.7 CodeTotale1TAttivi32B
- Kimi K2.6Totale1TAttivi32B