Mistral Medium 3.5
Mistral's dense 128B flagship, with open weights under a licence that shuts out companies above $20 million in monthly revenue.
- Total
- 128B
- Active
- 128B
- Experts
- Nonedense model
- Layers
- 88
- Attention
- GQA, 96 query and 8 KV heads
- Context
- 256k tokens
- Input
- text, image
- Output
- text
- Licence
- Modified MIT
- Weights
- Hugging Face
Contents1 to 3
How it is built
Mistral's flagship is dense. All 128B parameters are used on every token.11Mistral AI, Mistral Medium 3.5 128B model card, Hugging Face: "Mistral Medium 3.5 is our first flagship merged model." https://huggingface.co/mistralai/Mistral-Medium-3.5-128B The language model stacks 88 layers with a hidden size of 12,288.22Mistral AI, params.json in the official repository: n_layers 88, dim 12288, n_heads 96, n_kv_heads 8, hidden_dim 28672, vocab_size 131072, moe null, qformat_weight fp8_e4m3, vision encoder with 48 layers. https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/params.json
Attention is grouped-query, with 96 query heads and 8 key-value heads in each layer.22 The feed-forward block is 28,672 wide, and the vocabulary has 131,072 tokens.22
In front of the language model sits a vision encoder of 48 layers.22 Mistral trained it from scratch to handle variable image sizes and aspect ratios.11 The published weights are FP8, in the e4m3 format.22
What is new
Mistral calls Medium 3.5 its first flagship merged model. Reasoning and coding live in the same weights as instruction following.11
It took over from Mistral Medium 3.1 and Magistral in Le Chat, and from Devstral 2 in the Vibe coding agent.11
reasoning_effort takes none or high on each request. With high, Mistral recommends a temperature of 0.7. With none it allows anything from 0.0 to 0.7.11
Mistral reports 77.6% on SWE-Bench Verified and 91.4% on tau3-Telecom.11 The card says the model supersedes Devstral "across all benchmarks".11
The API version went live on 28 April 2026, and the documentation calls it Mistral's frontier-class multimodal model.33Mistral AI documentation, Mistral Medium 3.5 model card, version 26.04, released 28 April 2026, "frontier-class multimodal model optimized for agentic and coding use cases", prices per million tokens, read 29 September 2026. The API release of 28 April is the earliest official date; the launch post is dated 22 May 2026. https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04 The launch post came on 22 May.44Mistral AI, launch post for Mistral Medium 3.5 and remote agents in Vibe, 22 May 2026. https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5/
Running it
The weights are under a Modified MIT License. It grants no rights to anyone whose company, or employer, had global consolidated monthly revenue above $20 million in the preceding month.55Mistral AI, Modified MIT License of Mistral Medium 3.5, condition 2. https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/LICENSE Those companies need a commercial licence from Mistral, or its hosted services.55
In FP8 the weights take about 128 GB (128B x 1 byte).66Arithmetic: 128B parameters x 1 byte in FP8 = 128 GB for the language model weights, before the vision encoder and the KV cache. https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/params.json The launch post says self-hosting is possible "on as few as four GPUs".44 The vLLM command on the model card uses eight.11 An EAGLE draft model is published to speed up local inference.11
Through Mistral's API, mistral-medium-3-5 costs $1.5 per million input tokens and $7.5 per million output tokens. Its context window is 256k.33
The older Mistral Large 3 is much bigger, a mixture of experts of 675B released in December 2025.77Mistral AI documentation, Mistral Large 3 model card, version 25.12, released 2 December 2025, 675B total parameters, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12
Notes
-
Mistral AI, Mistral Medium 3.5 128B model card, Hugging Face: "Mistral Medium 3.5 is our first flagship merged model." https://huggingface.co/mistralai/Mistral-Medium-3.5-128B 2 3 4 5 6 7 8 9
-
Mistral AI, params.json in the official repository: n_layers 88, dim 12288, n_heads 96, n_kv_heads 8, hidden_dim 28672, vocab_size 131072, moe null, qformat_weight fp8_e4m3, vision encoder with 48 layers. https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/params.json 2 3 4 5
-
Mistral AI documentation, Mistral Medium 3.5 model card, version 26.04, released 28 April 2026, "frontier-class multimodal model optimized for agentic and coding use cases", prices per million tokens, read 29 September 2026. The API release of 28 April is the earliest official date; the launch post is dated 22 May 2026. https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04 2
-
Mistral AI, launch post for Mistral Medium 3.5 and remote agents in Vibe, 22 May 2026. https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5/ 2
-
Mistral AI, Modified MIT License of Mistral Medium 3.5, condition 2. https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/LICENSE 2
-
Arithmetic: 128B parameters x 1 byte in FP8 = 128 GB for the language model weights, before the vision encoder and the KV cache. https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/params.json
-
Mistral AI documentation, Mistral Large 3 model card, version 25.12, released 2 December 2025, 675B total parameters, read 29 September 2026. https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12
More from Mistral AI
- Mistral Small 4Total119BActive6.5B
- Ministral 3 14BTotal13.9BActive13.9B
- Mistral Large 3Total675BActive41B