Skip the cover

MiniMax

MiniMax-M3

MiniMax's natively multimodal flagship, about 428B parameters with 23B active, built on MiniMax Sparse Attention for a 1M-token context.

Total
428B
Active
23B
Experts
128 routed + 1 shared4 per token
Layers
60the first 3 dense
Attention
MiniMax Sparse Attention on GQA
Context
1M tokens
Input
text, image, video
Output
text
Licence
MiniMax Community License

Contents1 to 3

How it is built

MiniMax-M3 is a mixture of experts, and its card gives the size only roughly: about 428B parameters, about 23B of them active for each token.11MiniMax, MiniMax-M3 model card on Hugging Face: "~428B parameters and ~23B activated parameters"; mixed-modality training "from the very first step"; 9x prefill and 15x decode against M2 at 1M context, per-token compute cut to 1/20; thinking modes enabled, adaptive, disabled. https://huggingface.co/MiniMaxAI/MiniMax-M3. The MXFP8 repository is https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8

Three dense layers come first. In each of the other 57, the router keeps just 4 of 128 routed experts, beside 1 shared expert.22MiniMax-M3 config.json, text_config: num_hidden_layers 60, moe_layer_freq 0 for the first 3 layers and 1 for the other 57, num_local_experts 128, n_shared_experts 1, num_experts_per_tok 4, scoring_func sigmoid, use_routing_bias true, num_attention_heads 64, num_key_value_heads 4, sparse attention on the same 57 layers with sparse_block_size 128 and sparse_topk_blocks 16; vision_config num_hidden_layers 32. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/config.json. Arithmetic: 60 minus 3 = 57. MiniMax-M2.7 sends a token to 8 of 256 and has no shared expert.33MiniMax-M2.7 config.json: num_local_experts 256, shared_intermediate_size 0, num_experts_per_tok 8. https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/config.json

Fig. 1
Token

60 layers, the first 3 dense. In each of the other 57, a token goes to 4 of 128 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.
Grouped query attention
Several query heads share one set of keys and values, which shrinks the KV cache.

Every layer uses grouped query attention, with 64 query heads over 4 key-value heads.22 The first 3 layers attend in full. The other 57 use MiniMax Sparse Attention.22

Grouped query attention
Several query heads share one set of keys and values, which shrinks the KV cache.

An index branch scores blocks of earlier tokens and picks a few for each head group.44MiniMax, "MiniMax Sparse Attention", arXiv 2606.13392, 11 June 2026: blockwise sparse attention built on GQA, index branch with per-group top-k block selection; on a 109B test model, 28.4x less attention compute per token at 1M context, on par with GQA, 14.2x prefill and 7.6x decode wall-clock speedups on H800. https://arxiv.org/abs/2606.13392 The main branch then runs exact attention over those blocks only. In M3 a block is 128 tokens, and 16 blocks are kept.22

Images and video entered training alongside text "from the very first step".11

What is new

Against M2 at a 1M context, MiniMax reports prefill 9 times faster and decoding 15 times faster.11 Compute per token falls to a twentieth.11

The attention paper tested the method on a separate 109B model. At 1M tokens it cut attention compute per token 28.4 times.44 On H800 GPUs, prefill ran 14.2 times faster and decoding 7.6 times faster.44 Those numbers describe the test model, not M3.

The M2 series accepts only text and tool calls. M3 adds images and video.55MiniMax, Anthropic-compatible API reference: M3 supports text, image and video blocks; "The M2.7, M2.5, M2.1, and M2 series support text and tool-call content blocks only". https://platform.minimax.io/docs/api-reference/text-anthropic-api

MiniMax puts M3 at 83.5 on BrowseComp, above Claude Opus 4.7 at 79.3, and publishes no settings for the run.66MiniMax, MiniMax M3 model page: context up to 1M with a guaranteed minimum of 512K; BrowseComp 83.5 against 79.3 for Opus 4.7, no run settings given. https://www.minimax.io/models/text/m3

Running it

MiniMax-M3 opened on the API on 1 June 2026.77MiniMax, API release notes, MiniMax M3 entry dated 1 June 2026. https://platform.minimax.io/docs/release-notes/models The context runs to 1M tokens, with at least 512K guaranteed.6688MiniMax, text generation guide, model table: MiniMax-M3 context window 1,000,000. https://platform.minimax.io/docs/guides/text-generation Reasoning can be turned off or left to the model to decide.11

MiniMax marks its M3 prices "Permanent 50% off". With the discount, prompts up to 512K tokens cost $0.30 per million input tokens and $1.20 per million output tokens.99MiniMax, pay-as-you-go pricing, USD per million tokens, read 29 September 2026, marked "Permanent 50% off". Up to 512K input: 0.30 input, 1.20 output, 0.06 cache read, against list prices of 0.60, 2.40 and 0.12. Above 512K: 0.60, 2.40 and 0.12, against 1.20, 4.80 and 0.24. https://platform.minimax.io/docs/guides/pricing-paygo. Arithmetic: 0.60 / 0.30 = 2; 2.40 / 1.20 = 2. Longer prompts cost twice that.99

MXFP8
A number format of one byte per value, in which small blocks of values share a scale factor.

The BF16 weights on Hugging Face fill 854.2 GB. An MXFP8 copy in a second repository takes 443.7 GB.111010Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: MiniMaxAI/MiniMax-M3, 59 files, 854.2 GB; MiniMaxAI/MiniMax-M3-MXFP8, 31 files, 443.7 GB. https://huggingface.co/MiniMaxAI/MiniMax-M3/tree/main and https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8/tree/main

MXFP8
A number format of one byte per value, in which small blocks of values share a scale factor.

Non-commercial use is free under the MiniMax Community License.1111MiniMax Community License, clauses 2 and 3 and the appendix of prohibited uses. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE A commercial product must carry a visible "Built with MiniMax M3" credit, and its maker must send MiniMax a one-time notice by email. Above US$20 million of yearly revenue, written authorization comes first. Military use is prohibited.1111

Notes

  1. MiniMax, MiniMax-M3 model card on Hugging Face: "~428B parameters and ~23B activated parameters"; mixed-modality training "from the very first step"; 9x prefill and 15x decode against M2 at 1M context, per-token compute cut to 1/20; thinking modes enabled, adaptive, disabled. https://huggingface.co/MiniMaxAI/MiniMax-M3. The MXFP8 repository is https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8 2 3 4 5 6

  2. MiniMax-M3 config.json, text_config: num_hidden_layers 60, moe_layer_freq 0 for the first 3 layers and 1 for the other 57, num_local_experts 128, n_shared_experts 1, num_experts_per_tok 4, scoring_func sigmoid, use_routing_bias true, num_attention_heads 64, num_key_value_heads 4, sparse attention on the same 57 layers with sparse_block_size 128 and sparse_topk_blocks 16; vision_config num_hidden_layers 32. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/config.json. Arithmetic: 60 minus 3 = 57. 2 3 4

  3. MiniMax-M2.7 config.json: num_local_experts 256, shared_intermediate_size 0, num_experts_per_tok 8. https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/config.json

  4. MiniMax, "MiniMax Sparse Attention", arXiv 2606.13392, 11 June 2026: blockwise sparse attention built on GQA, index branch with per-group top-k block selection; on a 109B test model, 28.4x less attention compute per token at 1M context, on par with GQA, 14.2x prefill and 7.6x decode wall-clock speedups on H800. https://arxiv.org/abs/2606.13392 2 3

  5. MiniMax, Anthropic-compatible API reference: M3 supports text, image and video blocks; "The M2.7, M2.5, M2.1, and M2 series support text and tool-call content blocks only". https://platform.minimax.io/docs/api-reference/text-anthropic-api

  6. MiniMax, MiniMax M3 model page: context up to 1M with a guaranteed minimum of 512K; BrowseComp 83.5 against 79.3 for Opus 4.7, no run settings given. https://www.minimax.io/models/text/m3 2

  7. MiniMax, API release notes, MiniMax M3 entry dated 1 June 2026. https://platform.minimax.io/docs/release-notes/models

  8. MiniMax, text generation guide, model table: MiniMax-M3 context window 1,000,000. https://platform.minimax.io/docs/guides/text-generation

  9. MiniMax, pay-as-you-go pricing, USD per million tokens, read 29 September 2026, marked "Permanent 50% off". Up to 512K input: 0.30 input, 1.20 output, 0.06 cache read, against list prices of 0.60, 2.40 and 0.12. Above 512K: 0.60, 2.40 and 0.12, against 1.20, 4.80 and 0.24. https://platform.minimax.io/docs/guides/pricing-paygo. Arithmetic: 0.60 / 0.30 = 2; 2.40 / 1.20 = 2. 2

  10. Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: MiniMaxAI/MiniMax-M3, 59 files, 854.2 GB; MiniMaxAI/MiniMax-M3-MXFP8, 31 files, 443.7 GB. https://huggingface.co/MiniMaxAI/MiniMax-M3/tree/main and https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8/tree/main

  11. MiniMax Community License, clauses 2 and 3 and the appendix of prohibited uses. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE 2

More from MiniMax

All models by MiniMax