DeepSeek-V4-Pro
DeepSeek's largest model, with 1.6T parameters and 49B active per token. The 0813 snapshot replaced the April preview.
- Totale
- 1,6T
- Attivi
- 49B
- Esperti
- 384 instradati + 1 condiviso6 per token
- Strati
- 61
- Attenzione
- Hybrid: CSA and HCA
- Contesto
- 1M token
- Uscita massima
- 384k token
- Ingresso
- text
- Uscita
- text
- Licenza
- MIT
- Pesi
- Hugging Face
Indiceda 1 a 3
How it is built
At 1.6T parameters, DeepSeek-V4-Pro is the largest model DeepSeek has released. A token runs through 49B of them.11DeepSeek, DeepSeek-V4-Pro (preview) model card, 24 April 2026. FLOPs and KV cache compared with DeepSeek-V3.2 at a 1M-token context; "MoE expert parameters use FP4 precision; most other parameters use FP8". https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro 22DeepSeek, DeepSeek-V4.1-Flash model card, 10 September 2026. Its base-model table lists V4-Flash at 284B, V4-Pro at 1.6T and V4.1-Flash at 552B. Agentic benchmarks at maximum reasoning effort; SimpleQA-Verified is the 25-shot exact match of the base models. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Every one of the 61 layers carries 384 routed experts and 1 shared expert. A token uses 6 of the routed ones.33DeepSeek, DeepSeek-V4-Pro-0813 model card and config.json (num_hidden_layers 61, n_routed_experts 384, n_shared_experts 1, num_experts_per_tok 6, num_hash_layers 3). Benchmarks in the minimal mode of DeepSeek Harness, max reasoning effort, temperature 1.0, top_p 0.95. https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 In the first 3 layers a hash of the token ID fixes the experts.44DeepSeek, DeepSeek-V4 technical report, arXiv 2606.19348: the first dense layers replaced by hash-routed MoE layers, and affinity scores computed with Sqrt(Softplus) instead of Sigmoid. https://arxiv.org/abs/2606.19348 The later layers learn their routing and score experts with Sqrt(Softplus), where DeepSeek-V3 used a sigmoid.44
Attention combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA). At a 1M context DeepSeek puts the model at 27% of the single-token inference FLOPs of DeepSeek-V3.2 and 10% of its KV cache.11 Manifold-Constrained Hyper-Connections (mHC) reinforce the residual connections.11 The routed experts are stored in FP4, most other weights in FP8.11
What is new
The first version, published on 24 April 2026, was a preview with open weights.55DeepSeek API Docs, DeepSeek-V4 preview release note, 24 April 2026. https://api-docs.deepseek.com/news/news260424 DeepSeek-V4-Pro-0813 is the official release. It keeps the preview's structure and adds a DSpark module for speculative decoding.33
The preview scored 12.8 on DeepSWE. The 0813 snapshot scores 62.7.33 On Terminal Bench 2.1 it moved from 72.1 to 87.9, in the minimal mode of DeepSeek Harness at max effort.33 Reasoning effort now has three levels, from low to max.33
In September DeepSeek released the smaller DeepSeek-V4.1-Flash. Its card has it ahead of V4-Pro on every agentic benchmark the two share.22 V4-Pro keeps the lead on knowledge: its base model scores 55.2 on SimpleQA-Verified, against 42.3.22
Running it
The 0813 weights carry the MIT licence.33 The card's vLLM example serves them on a single node of four GB300 GPUs, and one flag turns DSpark on.33 At high and max effort it recommends room for 384K output tokens.33
The API is less settled. The pricing page maps deepseek-v4-pro to DeepSeek-V4-Pro-0813, with 1M tokens of context and 384K of output. Images are not accepted.66DeepSeek API Docs, Models and Pricing, as listed on 29 September 2026. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. https://api-docs.deepseek.com/quick_start/pricing At peak a million input tokens cost $1.32, or $0.044 on a cache hit, and a million output tokens $3.96. Off-peak prices are half.66
The release note of 10 September contradicts it. From 14 September, it states, every deepseek-v4-pro request goes to V4.1-Flash at V4.1-Flash prices, until a V4.1-Pro launches.77DeepSeek API Docs, DeepSeek-V4.1-Flash release note, 10 September 2026. The routing starts on 14 September at 04:00 UTC. https://api-docs.deepseek.com/news/news260910 The pricing page does not mention that routing.66
Note
-
DeepSeek, DeepSeek-V4-Pro (preview) model card, 24 April 2026. FLOPs and KV cache compared with DeepSeek-V3.2 at a 1M-token context; "MoE expert parameters use FP4 precision; most other parameters use FP8". https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro 2 3 4
-
DeepSeek, DeepSeek-V4.1-Flash model card, 10 September 2026. Its base-model table lists V4-Flash at 284B, V4-Pro at 1.6T and V4.1-Flash at 552B. Agentic benchmarks at maximum reasoning effort; SimpleQA-Verified is the 25-shot exact match of the base models. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 2 3
-
DeepSeek, DeepSeek-V4-Pro-0813 model card and config.json (num_hidden_layers 61, n_routed_experts 384, n_shared_experts 1, num_experts_per_tok 6, num_hash_layers 3). Benchmarks in the minimal mode of DeepSeek Harness, max reasoning effort, temperature 1.0, top_p 0.95. https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 2 3 4 5 6 7 8
-
DeepSeek, DeepSeek-V4 technical report, arXiv 2606.19348: the first dense layers replaced by hash-routed MoE layers, and affinity scores computed with Sqrt(Softplus) instead of Sigmoid. https://arxiv.org/abs/2606.19348 2
-
DeepSeek API Docs, DeepSeek-V4 preview release note, 24 April 2026. https://api-docs.deepseek.com/news/news260424
-
DeepSeek API Docs, Models and Pricing, as listed on 29 September 2026. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. https://api-docs.deepseek.com/quick_start/pricing 2 3
-
DeepSeek API Docs, DeepSeek-V4.1-Flash release note, 10 September 2026. The routing starts on 14 September at 04:00 UTC. https://api-docs.deepseek.com/news/news260910
Altri modelli di DeepSeek
- DeepSeek-V4.1-FlashTotale552BAttivi16B