Mistral Small 4
Mistral's 119B mixture of experts with 6.5B active per token, holding Magistral's reasoning in the same weights as Instruct.
- Total
- 119B
- Active
- 6.5B
- Experts
- 128 routed + 1 shared4 per token
- Layers
- 36
- Attention
- Multi-head latent attention
- Context
- 256k tokens
- Input
- text, image
- Output
- text
- Licence
- Apache 2.0
- Weights
- Hugging Face
Contents1 to 3
How it is built
Mistral Small 4 has 119B parameters, and Mistral gives two counts for the active part. The documentation and the model card say 6.5B per token, about 5.5%.11Mistral AI, Mistral Small 4 119B A6B model card, Hugging Face. https://huggingface.co/mistralai/Mistral-Small-4-119B-2603. Arithmetic: 6.5 / 119 = 5.5%. 22Mistral AI documentation, Mistral Small 4 model card, version 26.03, released 16 March 2026, prices per million tokens, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03 The launch post says 6B, or 8B with the embedding and output layers.33Mistral AI, launch post for Mistral Small 4, 16 March 2026. https://mistral.ai/news/mistral-small-4/
All 36 layers are mixture of experts layers, with no dense layer in front.44Mistral AI, params.json in the official repository: n_layers 36, first_k_dense_replace 0, num_experts 128, num_shared_experts 1, num_experts_per_tok 4, kv_lora_rank 256, q_lora_rank 1024, max_position_embeddings 1048576, qformat_weight fp8_e4m3, vision encoder with 24 layers. https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/blob/main/params.json The expert counts match Mistral Large 3: 128 routed experts and 1 shared expert per layer, and 4 routed experts for each token.44
Keys and values are stored in a compressed latent of rank 256.44 This is multi-head latent attention, and the vLLM command on the card selects an MLA backend for it.11
Images go through a vision encoder of 24 layers.44
What is new
The reasoning models Mistral used to call Magistral now live in the same weights as Instruct. So does Devstral.11
Reasoning is chosen per request with reasoning_effort. At none the replies read like Mistral Small 3.2. At high they run as long as the old Magistral models.11
Against Mistral Small 3, Mistral reports a 40% cut in end-to-end completion time in a latency-optimized setup.11 Optimized for throughput instead, it serves 3 times as many requests per second.11
Mistral also counts output length. On AA LCR, Small 4 with reasoning scores 0.72 with 1.6K characters of output, and Mistral says Qwen models need 3.5 to 4 times more for comparable scores.11 On LiveCodeBench it beats GPT-OSS 120B while writing 20% less.11
Running it
The weights are FP8, under Apache 2.0.1144 That comes to about 119 GB (119B x 1 byte).55Arithmetic: 119B x 1 byte = 119 GB, weights only, before the KV cache. https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/blob/main/params.json The card's vLLM command spreads them over 2 GPUs.11 An NVFP4 checkpoint and an EAGLE head for speculative decoding are also published.11
The config allows 1,048,576 positions.44 Both the documentation and the card stop at 256k tokens.1122
mistral-small-2603 costs $0.15 per million input tokens and $0.6 per million output tokens on the API.22
Six weeks later came Mistral Medium 3.5, a dense model of 128B.66Mistral AI documentation, Mistral Medium 3.5 model card, released 28 April 2026. https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04. Arithmetic: 16 March to 28 April 2026 is 43 days, about six weeks. It is only a little larger, and it uses every parameter on every token.
Notes
-
Mistral AI, Mistral Small 4 119B A6B model card, Hugging Face. https://huggingface.co/mistralai/Mistral-Small-4-119B-2603. Arithmetic: 6.5 / 119 = 5.5%. 2 3 4 5 6 7 8 9 10 11 12
-
Mistral AI documentation, Mistral Small 4 model card, version 26.03, released 16 March 2026, prices per million tokens, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03 2 3
-
Mistral AI, launch post for Mistral Small 4, 16 March 2026. https://mistral.ai/news/mistral-small-4/
-
Mistral AI, params.json in the official repository: n_layers 36, first_k_dense_replace 0, num_experts 128, num_shared_experts 1, num_experts_per_tok 4, kv_lora_rank 256, q_lora_rank 1024, max_position_embeddings 1048576, qformat_weight fp8_e4m3, vision encoder with 24 layers. https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/blob/main/params.json 2 3 4 5 6
-
Arithmetic: 119B x 1 byte = 119 GB, weights only, before the KV cache. https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/blob/main/params.json
-
Mistral AI documentation, Mistral Medium 3.5 model card, released 28 April 2026. https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04. Arithmetic: 16 March to 28 April 2026 is 43 days, about six weeks.
More from Mistral AI
- Mistral Medium 3.5Total128BActive128B
- Ministral 3 14BTotal13.9BActive13.9B
- Mistral Large 3Total675BActive41B