Ministral 3 14B
The largest Ministral 3 model, a dense 14B that Mistral pruned from Mistral Small 3.1 and then trained by distillation.
- Totale
- 13,9B
- Attivi
- 13,9B
- Esperti
- Nessunomodello denso
- Strati
- 40
- Attenzione
- GQA, 32 query and 8 KV heads
- Contesto
- 256k token
- Ingresso
- text, image
- Uscita
- text
- Licenza
- Apache 2.0
- Pesi
- Hugging Face
Indiceda 1 a 3
How it is built
The name says 14B. The card splits the model into a 13.5B language model and a 0.4B vision encoder, 13.9B in all.11Mistral AI, Ministral 3 14B Instruct 2512 model card, Hugging Face. https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512. Arithmetic: 13.5B + 0.4B = 13.9B.
It is a dense model of 40 layers with a hidden size of 5,120.22Mistral AI, params.json in the official repository: n_layers 40, dim 5120, n_heads 32, n_kv_heads 8, hidden_dim 16384, max_position_embeddings 262144. https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512/blob/main/params.json Each layer has 32 query heads and 8 key-value heads, and the feed-forward block is 16,384 wide.22
Mistral did not train it from scratch. The Ministral 3 paper calls the method Cascade Distillation, "an iterative pruning and continued training with distillation technique".33Mistral AI, "Ministral 3", arXiv 2601.08584, January 2026. https://arxiv.org/abs/2601.08584 The 14B started as a pruned copy of Mistral Small 3.1, a 24B model that was also its teacher in pre-training.33 In pre-training, distilling from Small 3.1 worked better than distilling from the much stronger Mistral Medium 3. Medium 3 became the teacher only in post-training.33
The family was trained on between 1 and 3 trillion tokens.33 Its context was stretched from 16,384 to 262,144 positions with YaRN and position-based softmax temperature scaling.33
What is new
Ministral 3 came out on 2 December 2025. The 14B is the largest of its three sizes, and the others are 8B and 3B.44Mistral AI, "Introducing Mistral 3", 2 December 2025. https://mistral.ai/news/mistral-3
Every size comes as a base model and an instruction-tuned model, with a reasoning version on top. All of them read images.44
Mistral's documentation calls its performance comparable to the larger Mistral Small 3.2 24B.55Mistral AI documentation, Ministral 3 14B model card, version 25.12, released 2 December 2025, prices per million tokens, read 29 September 2026. https://docs.mistral.ai/models/model-cards/ministral-3-14b-25-12 The reasoning version scores 0.850 on AIME 2025, where Qwen3-14B with thinking scores 0.737.11 Its GPQA Diamond score is 0.712, and LiveCodeBench gives 0.646.11
The instruct version is set against Qwen3 14B without thinking. It leads on Arena Hard, 0.551 to 0.427, and on WildBench, 68.5 to 65.1.11
Running it
The instruct weights ship in FP8. The card says the model fits "in 24GB of VRAM in FP8, and less if further quantized".11 The weights alone come to about 14 GB (13.9B x 1 byte), before the KV cache.66Arithmetic: 13.9B parameters x 1 byte in FP8 = 13.9 GB, weights only. https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512 Base and reasoning versions are in BF16.11
Mistral recommends a temperature below 0.1 in production. Images should have an aspect ratio close to 1:1.11
Input and output cost the same on the API, $0.2 per million tokens, under the id ministral-14b-2512.55 The context window is 256k tokens.2255
Every version is under Apache 2.0.44 So is Mistral Small 4, a mixture of experts of 119B.77Mistral AI documentation, Mistral Small 4 model card. https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03
Note
-
Mistral AI, Ministral 3 14B Instruct 2512 model card, Hugging Face. https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512. Arithmetic: 13.5B + 0.4B = 13.9B. 2 3 4 5 6 7
-
Mistral AI, params.json in the official repository: n_layers 40, dim 5120, n_heads 32, n_kv_heads 8, hidden_dim 16384, max_position_embeddings 262144. https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512/blob/main/params.json 2 3
-
Mistral AI, "Ministral 3", arXiv 2601.08584, January 2026. https://arxiv.org/abs/2601.08584 2 3 4 5
-
Mistral AI, "Introducing Mistral 3", 2 December 2025. https://mistral.ai/news/mistral-3 2 3
-
Mistral AI documentation, Ministral 3 14B model card, version 25.12, released 2 December 2025, prices per million tokens, read 29 September 2026. https://docs.mistral.ai/models/model-cards/ministral-3-14b-25-12 2 3
-
Arithmetic: 13.9B parameters x 1 byte in FP8 = 13.9 GB, weights only. https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512
-
Mistral AI documentation, Mistral Small 4 model card. https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03
Altri modelli di Mistral AI
- Mistral Medium 3.5Totale128BAttivi128B
- Mistral Small 4Totale119BAttivi6,5B
- Mistral Large 3Totale675BAttivi41B