MiniMax-M3
MiniMax's natively multimodal flagship, about 428B parameters with 23B active, built on MiniMax Sparse Attention for a 1M-token context.
- Totale
- 428B
- Attivi
- 23B
- Esperti
- 128 instradati + 1 condiviso4 per token
- Strati
- 60i primi 3 densi
- Attenzione
- MiniMax Sparse Attention on GQA
- Contesto
- 1M token
- Ingresso
- text, image, video
- Uscita
- text
- Licenza
- MiniMax Community License
- Pesi
- Hugging Face
Indiceda 1 a 3
How it is built
MiniMax-M3 is a mixture of experts, and its card gives the size only roughly: about 428B parameters, about 23B of them active for each token.11MiniMax, MiniMax-M3 model card on Hugging Face: "~428B parameters and ~23B activated parameters"; mixed-modality training "from the very first step"; 9x prefill and 15x decode against M2 at 1M context, per-token compute cut to 1/20; thinking modes enabled, adaptive, disabled. https://huggingface.co/MiniMaxAI/MiniMax-M3. The MXFP8 repository is https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8
Three dense layers come first. In each of the other 57, the router keeps just 4 of 128 routed experts, beside 1 shared expert.22MiniMax-M3 config.json, text_config: num_hidden_layers 60, moe_layer_freq 0 for the first 3 layers and 1 for the other 57, num_local_experts 128, n_shared_experts 1, num_experts_per_tok 4, scoring_func sigmoid, use_routing_bias true, num_attention_heads 64, num_key_value_heads 4, sparse attention on the same 57 layers with sparse_block_size 128 and sparse_topk_blocks 16; vision_config num_hidden_layers 32. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/config.json. Arithmetic: 60 minus 3 = 57. MiniMax-M2.7 sends a token to 8 of 256 and has no shared expert.33MiniMax-M2.7 config.json: num_local_experts 256, shared_intermediate_size 0, num_experts_per_tok 8. https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/config.json
Every layer uses grouped query attention, with 64 query heads over 4 key-value heads.22 The first 3 layers attend in full. The other 57 use MiniMax Sparse Attention.22
An index branch scores blocks of earlier tokens and picks a few for each head group.44MiniMax, "MiniMax Sparse Attention", arXiv 2606.13392, 11 June 2026: blockwise sparse attention built on GQA, index branch with per-group top-k block selection; on a 109B test model, 28.4x less attention compute per token at 1M context, on par with GQA, 14.2x prefill and 7.6x decode wall-clock speedups on H800. https://arxiv.org/abs/2606.13392 The main branch then runs exact attention over those blocks only. In M3 a block is 128 tokens, and 16 blocks are kept.22
Images and video entered training alongside text "from the very first step".11
What is new
Against M2 at a 1M context, MiniMax reports prefill 9 times faster and decoding 15 times faster.11 Compute per token falls to a twentieth.11
The attention paper tested the method on a separate 109B model. At 1M tokens it cut attention compute per token 28.4 times.44 On H800 GPUs, prefill ran 14.2 times faster and decoding 7.6 times faster.44 Those numbers describe the test model, not M3.
The M2 series accepts only text and tool calls. M3 adds images and video.55MiniMax, Anthropic-compatible API reference: M3 supports text, image and video blocks; "The M2.7, M2.5, M2.1, and M2 series support text and tool-call content blocks only". https://platform.minimax.io/docs/api-reference/text-anthropic-api
MiniMax puts M3 at 83.5 on BrowseComp, above Claude Opus 4.7 at 79.3, and publishes no settings for the run.66MiniMax, MiniMax M3 model page: context up to 1M with a guaranteed minimum of 512K; BrowseComp 83.5 against 79.3 for Opus 4.7, no run settings given. https://www.minimax.io/models/text/m3
Running it
MiniMax-M3 opened on the API on 1 June 2026.77MiniMax, API release notes, MiniMax M3 entry dated 1 June 2026. https://platform.minimax.io/docs/release-notes/models The context runs to 1M tokens, with at least 512K guaranteed.6688MiniMax, text generation guide, model table: MiniMax-M3 context window 1,000,000. https://platform.minimax.io/docs/guides/text-generation Reasoning can be turned off or left to the model to decide.11
MiniMax marks its M3 prices "Permanent 50% off". With the discount, prompts up to 512K tokens cost $0.30 per million input tokens and $1.20 per million output tokens.99MiniMax, pay-as-you-go pricing, USD per million tokens, read 29 September 2026, marked "Permanent 50% off". Up to 512K input: 0.30 input, 1.20 output, 0.06 cache read, against list prices of 0.60, 2.40 and 0.12. Above 512K: 0.60, 2.40 and 0.12, against 1.20, 4.80 and 0.24. https://platform.minimax.io/docs/guides/pricing-paygo. Arithmetic: 0.60 / 0.30 = 2; 2.40 / 1.20 = 2. Longer prompts cost twice that.99
The BF16 weights on Hugging Face fill 854.2 GB. An MXFP8 copy in a second repository takes 443.7 GB.111010Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: MiniMaxAI/MiniMax-M3, 59 files, 854.2 GB; MiniMaxAI/MiniMax-M3-MXFP8, 31 files, 443.7 GB. https://huggingface.co/MiniMaxAI/MiniMax-M3/tree/main and https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8/tree/main
Non-commercial use is free under the MiniMax Community License.1111MiniMax Community License, clauses 2 and 3 and the appendix of prohibited uses. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE A commercial product must carry a visible "Built with MiniMax M3" credit, and its maker must send MiniMax a one-time notice by email. Above US$20 million of yearly revenue, written authorization comes first. Military use is prohibited.1111
Note
-
MiniMax, MiniMax-M3 model card on Hugging Face: "~428B parameters and ~23B activated parameters"; mixed-modality training "from the very first step"; 9x prefill and 15x decode against M2 at 1M context, per-token compute cut to 1/20; thinking modes enabled, adaptive, disabled. https://huggingface.co/MiniMaxAI/MiniMax-M3. The MXFP8 repository is https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8 2 3 4 5 6
-
MiniMax-M3 config.json, text_config: num_hidden_layers 60, moe_layer_freq 0 for the first 3 layers and 1 for the other 57, num_local_experts 128, n_shared_experts 1, num_experts_per_tok 4, scoring_func sigmoid, use_routing_bias true, num_attention_heads 64, num_key_value_heads 4, sparse attention on the same 57 layers with sparse_block_size 128 and sparse_topk_blocks 16; vision_config num_hidden_layers 32. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/config.json. Arithmetic: 60 minus 3 = 57. 2 3 4
-
MiniMax-M2.7 config.json: num_local_experts 256, shared_intermediate_size 0, num_experts_per_tok 8. https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/config.json
-
MiniMax, "MiniMax Sparse Attention", arXiv 2606.13392, 11 June 2026: blockwise sparse attention built on GQA, index branch with per-group top-k block selection; on a 109B test model, 28.4x less attention compute per token at 1M context, on par with GQA, 14.2x prefill and 7.6x decode wall-clock speedups on H800. https://arxiv.org/abs/2606.13392 2 3
-
MiniMax, Anthropic-compatible API reference: M3 supports text, image and video blocks; "The M2.7, M2.5, M2.1, and M2 series support text and tool-call content blocks only". https://platform.minimax.io/docs/api-reference/text-anthropic-api
-
MiniMax, MiniMax M3 model page: context up to 1M with a guaranteed minimum of 512K; BrowseComp 83.5 against 79.3 for Opus 4.7, no run settings given. https://www.minimax.io/models/text/m3 2
-
MiniMax, API release notes, MiniMax M3 entry dated 1 June 2026. https://platform.minimax.io/docs/release-notes/models
-
MiniMax, text generation guide, model table: MiniMax-M3 context window 1,000,000. https://platform.minimax.io/docs/guides/text-generation
-
MiniMax, pay-as-you-go pricing, USD per million tokens, read 29 September 2026, marked "Permanent 50% off". Up to 512K input: 0.30 input, 1.20 output, 0.06 cache read, against list prices of 0.60, 2.40 and 0.12. Above 512K: 0.60, 2.40 and 0.12, against 1.20, 4.80 and 0.24. https://platform.minimax.io/docs/guides/pricing-paygo. Arithmetic: 0.60 / 0.30 = 2; 2.40 / 1.20 = 2. 2
-
Sum of the .safetensors file sizes listed by Hugging Face, read 29 September 2026: MiniMaxAI/MiniMax-M3, 59 files, 854.2 GB; MiniMaxAI/MiniMax-M3-MXFP8, 31 files, 443.7 GB. https://huggingface.co/MiniMaxAI/MiniMax-M3/tree/main and https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8/tree/main
-
MiniMax Community License, clauses 2 and 3 and the appendix of prohibited uses. https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE 2
Altri modelli di MiniMax
- MiniMax-M2.7Totale229BAttiviNon dichiarato