Skip the cover

Mistral AI

Mistral Large 3

Mistral's first mixture of experts since Mixtral: 675B parameters with 41B active, open weights under Apache 2.0.

Total
675B
Active
41B
Experts
128 routed + 1 shared4 per token
Layers
61the first 3 dense
Context
256k tokens
Input
text, image
Output
text
Licence
Apache 2.0

Contents1 to 3

How it is built

Mistral Large 3 activates 41B of its 675B parameters for each token.11Mistral AI documentation, Mistral Large 3 model card, version 25.12, released 2 December 2025, prices per million tokens, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12. Arithmetic: 41 / 675 = 6.1%. Most of the total is the language model, at 673B with 39B active. The vision encoder adds 2.5B.22Mistral AI, Mistral Large 3 675B Instruct 2512 model card, Hugging Face. https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512

The language model has 61 layers. The first 3 are dense, and the other 58 are mixture of experts layers.33Mistral AI, params.json in the official repository: n_layers 61, first_k_dense_replace 3, num_experts 128, num_shared_experts 1, num_experts_per_tok 4, expert_hidden_dim 4096, hidden_dim 16384. https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512/blob/main/params.json. Arithmetic: 61 - 3 = 58 MoE layers; 4,096 / 16,384 = 1/4. Each of those holds 128 routed experts and 1 shared expert, and the router sends every token to 4 of the routed ones.33

Fig. 1
Token

61 layers, the first 3 dense. In each of the other 58, a token goes to 4 of 128 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.
Shared expert
An expert that every token passes through, next to the few the router picks.

Mistral calls the design a "granular Mixture-of-Experts".11 The experts are narrow. Each is 4,096 wide, a quarter of the 16,384 in the dense layers.33 About 6.1% of the parameters work on any one token.11

Shared expert
An expert that every token passes through, next to the few the router picks.

What is new

Mistral had not built a mixture of experts since the Mixtral series.44Mistral AI, "Introducing Mistral 3", 2 December 2025. https://mistral.ai/news/mistral-3 Large 3 was trained from scratch on 3,000 NVIDIA H200 GPUs.44

The base model and the instruction-tuned model both came out under Apache 2.0.44 At launch Large 3 ranked second among open non-reasoning models on LMArena.44 Amazon Bedrock and Azure Foundry carry it as well as Mistral's own platform.44

The model card lists its limits. Large 3 is "not a dedicated reasoning model", and on multimodal tasks it can trail models built for vision first.22

Running it

The instruct weights are FP8, about 675 GB at one byte per parameter (675B x 1 byte).2255Arithmetic: 675B x 1 byte = 675 GB, weights only, before the KV cache. https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512

The card puts the FP8 version on a single node of B200 or H200 GPUs.22 An NVFP4 checkpoint brings it down to a single node of H100 or A100 GPUs. A BF16 version is published too, with an EAGLE draft model for speculative decoding.22

The vLLM command on the card sets a tensor-parallel size of 8 and a maximum model length of 262,144 tokens.22 For production, Mistral wants the temperature under 0.1.22

On the API it is mistral-large-2512, at $0.5 per million input tokens and $1.5 per million output tokens, with a 256k context window.11

Less than five months later Mistral released Mistral Medium 3.5, a dense 128B model. Its output costs five times as much, $7.5 per million tokens.66Mistral AI documentation, Mistral Medium 3.5 model card, released 28 April 2026, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04. Arithmetic: 7.5 / 1.5 = 5; 2 December 2025 to 28 April 2026 is almost five months.

Notes

  1. Mistral AI documentation, Mistral Large 3 model card, version 25.12, released 2 December 2025, prices per million tokens, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12. Arithmetic: 41 / 675 = 6.1%. 2 3 4

  2. Mistral AI, Mistral Large 3 675B Instruct 2512 model card, Hugging Face. https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512 2 3 4 5 6 7

  3. Mistral AI, params.json in the official repository: n_layers 61, first_k_dense_replace 3, num_experts 128, num_shared_experts 1, num_experts_per_tok 4, expert_hidden_dim 4096, hidden_dim 16384. https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512/blob/main/params.json. Arithmetic: 61 - 3 = 58 MoE layers; 4,096 / 16,384 = 1/4. 2 3

  4. Mistral AI, "Introducing Mistral 3", 2 December 2025. https://mistral.ai/news/mistral-3 2 3 4 5

  5. Arithmetic: 675B x 1 byte = 675 GB, weights only, before the KV cache. https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512

  6. Mistral AI documentation, Mistral Medium 3.5 model card, released 28 April 2026, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04. Arithmetic: 7.5 / 1.5 = 5; 2 December 2025 to 28 April 2026 is almost five months.

More from Mistral AI

All models by Mistral AI