Skip the cover

Moonshot AI

Kimi K2.6

A 1T mixture of experts with 32B active per token, built for long coding runs and for agent swarms of up to 300 sub-agents.

Total
1T
Active
32B
Experts
384 routed + 1 shared8 per token
Layers
61the first one dense
Attention
MLA
Context
262k tokens
Input
text, image
Output
text
Licence
Modified MIT

Contents1 to 3

How it is built

Kimi K2.6 has the shape of Kimi K2.5: 1T parameters, 32B active per token.11Moonshot AI, Kimi K2.6 model card, April 2026. Runs with thinking on, temperature 1.0, top-p 1.0 and a 262,144-token context. SWE-Bench Pro and Verified with an in-house framework adapted from SWE-agent, coding scores averaged over 10 runs; Terminal-Bench 2.0 with Terminus-2; BrowseComp with the agent swarm; HLE-Full with search, code-interpreter and web-browsing tools. https://huggingface.co/moonshotai/Kimi-K2.6 Moonshot says a K2.5 deployment can be reused as it is.11

Sixty of the 61 layers carry experts: 384 routed and 1 shared in each, with 8 routed experts chosen per token.11 An expert is 2,048 wide.11

MLA
Multi-head Latent Attention stores keys and values as a compressed latent vector, which shrinks the KV cache.

Attention is Multi-head Latent Attention (MLA), with 64 heads at a hidden size of 7,168.11 The vocabulary has 160K tokens. Images enter through MoonViT, an encoder of 400M parameters.11

MLA
Multi-head Latent Attention stores keys and values as a compressed latent vector, which shrinks the KV cache.
Fig. 1
Token

61 layers, the first one dense. In each of the other 60, a token goes to 8 of 384 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.

What is new

The Kimi blog dates K2.6 to 20 April 2026.22Moonshot AI, Kimi blog index, entry dated 2026-04-20. https://www.kimi.ai/blog

Moonshot measured it against K2.5. SWE-Bench Pro rose from 50.7 to 58.6, averaged over 10 runs in an in-house framework adapted from SWE-agent. That framework gives the agent six tools, bash among them.11 In the same framework SWE-Bench Verified went from 76.8 to 80.2.11 The larger jump is on Terminal-Bench 2.0 with the Terminus-2 agent, from 50.8 to 66.7. That run kept thinking preserved across turns and used the default JSON parser.11

Agent swarm
One task split among many copies of the model, each working on a part in parallel.

The agent swarm scales to 300 sub-agents working through 4,000 coordinated steps.11 With the swarm, BrowseComp reaches 86.3. With search and code tools, HLE-Full rose from 50.2 to 54.0.11

Agent swarm
One task split among many copies of the model, each working on a part in parallel.

Thinking can be switched off. The card calls that instant mode and recommends a temperature of 0.6 for it.11

Running it

Modified MIT is the MIT licence with one condition. Above 100 million monthly active users, or US$20 million of monthly revenue, a product has to show "Kimi K2.6".33Moonshot AI, Modified MIT License of Kimi K2.6. https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE

The weights are native INT4, quantized the way Kimi K2 Thinking was.11 Moonshot suggests vLLM or SGLang, and offers a Kimi Vendor Verifier to check that a third-party deployment behaves correctly.11 Video input is experimental and works only on Moonshot's own API.11

kimi-k2.6 costs $0.95 per million input tokens, or $0.16 on a cache hit, and $4.00 per million output tokens. Its context is 262,144 tokens.44Kimi API Platform, chat model pricing, per million tokens, as listed on 29 September 2026. https://platform.kimi.ai/docs/pricing/chat

Kimi K2.7 Code is a coding model built on K2.6.55Moonshot AI, Kimi K2.7 Code model card: "a coding-focused agentic model built upon Kimi K2.6". https://huggingface.co/moonshotai/Kimi-K2.7-Code

Notes

  1. Moonshot AI, Kimi K2.6 model card, April 2026. Runs with thinking on, temperature 1.0, top-p 1.0 and a 262,144-token context. SWE-Bench Pro and Verified with an in-house framework adapted from SWE-agent, coding scores averaged over 10 runs; Terminal-Bench 2.0 with Terminus-2; BrowseComp with the agent swarm; HLE-Full with search, code-interpreter and web-browsing tools. https://huggingface.co/moonshotai/Kimi-K2.6 2 3 4 5 6 7 8 9 10 11 12 13 14 15

  2. Moonshot AI, Kimi blog index, entry dated 2026-04-20. https://www.kimi.ai/blog

  3. Moonshot AI, Modified MIT License of Kimi K2.6. https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE

  4. Kimi API Platform, chat model pricing, per million tokens, as listed on 29 September 2026. https://platform.kimi.ai/docs/pricing/chat

  5. Moonshot AI, Kimi K2.7 Code model card: "a coding-focused agentic model built upon Kimi K2.6". https://huggingface.co/moonshotai/Kimi-K2.7-Code

More from Moonshot AI

All models by Moonshot AI