Salta la copertina

DeepSeekIn inglese

DeepSeek-V4.1-Flash

A 552B mixture of experts built as a causal encoder and decoder, with 8B parameters active per token in prefill and 16B in decode.

Totale
552B
Attivi
16B
Esperti
384 instradati + 1 condiviso6 per token
Strati
40
Attenzione
Compressed Sparse Attention 2
Contesto
1M token
Uscita massima
384k token
Ingresso
text, image
Uscita
text
Licenza
MIT

Indiceda 1 a 3

How it is built

Prefill and decode
Prefill reads the prompt. Decode then writes the answer, one token at a time.

DeepSeek-V4.1-Flash splits its 40 layers in two: a causal encoder of 20 layers, then a decoder of 20.11DeepSeek, DeepSeek-V4.1-Flash model card, 10 September 2026. Instruct benchmarks at reasoning_effort=100, temperature 1.0, top_p 0.95; Terminal-Bench 2.1 Pass@1 in the Minimal mode of DeepSeek Harness with a 1M-token context. Base models on 25-shot SimpleQA-Verified (exact match): V4.1-Flash 42.3, V4-Pro 55.2. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash The decoder does not build its global KV cache from its own layers. It projects the cache from the encoder's final hidden states.11 DeepSeek gives this as the reason a token needs only 8B parameters in prefill and 16B in decode.11

Prefill and decode
Prefill reads the prompt. Decode then writes the answer, one token at a time.
Engram
A conditional memory that the model reads sparsely, through lookups keyed on the tokens.

The backbone counts 552B parameters.11 The Engram memory is listed apart at 196B. The card does not say whether the 552B includes it.11

Engram
A conditional memory that the model reads sparsely, through lookups keyed on the tokens.

In each MoE layer a token reaches 6 of 384 routed experts, plus 1 shared expert.22DeepSeek, DeepSeek-V4.1-Flash config.json (n_routed_experts 384, n_shared_experts 1, num_experts_per_tok 6, num_hidden_layers 40, sliding_window 128, fp8 weights in 32 x 32 blocks with fp4 experts, vision encoder of 32 layers). https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/config.json Attention is Compressed Sparse Attention 2 (CSA2). Each attention layer is fixed to one of three modes, and through them layers share their KV cache and reuse the tokens that an earlier layer's sparse attention picked.11 A 128-token sliding window runs beside it.22 Images go through a vision encoder of 32 layers and enter the model with the text.1122

Fig. 1
Token

40 layers. In each one, a token goes to 6 of 384 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.

What is new

Weights and API came out together on 10 September 2026.33DeepSeek API Docs, DeepSeek-V4.1-Flash release note, 10 September 2026. https://api-docs.deepseek.com/news/news260910

The main KV cache is now stored in FP4. At 890 bytes per token the global cache is about a quarter the size of DeepSeek-V4-Flash's.11 A single sequence of 1,048,576 tokens fills about 933 MB.44Arithmetic on the model card figure: 890 bytes x 1,048,576 tokens = 933,232,640 bytes, about 933 MB. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash What stays on SSD shrinks to about an eighth of V4-Flash's.11

Training started from scratch on a multimodal corpus of 45T tokens. The context was stretched to 1M tokens at the 34T mark.11 Reasoning effort takes any integer from 1 to 100.11

DeepSeek's own table puts V4.1-Flash ahead of the larger DeepSeek-V4-Pro on Terminal-Bench 2.1, at 90.6 against 87.9.11 On factual recall its base model trails.11

Running it

The weights are MIT-licensed. They are stored in FP8 in blocks of 32 by 32, with the routed experts in FP4.1122 No Jinja chat template comes with them, and DeepSeek ships a Python reference encoder in its place.11 The card asks for a 1M context window and a max_tokens of at least 256K.11

The API calls the model deepseek-flash, with 1M tokens of context and up to 384K of output.55DeepSeek API Docs, Models and Pricing, as listed on 29 September 2026. Off-peak covers all other hours, including weekends and Chinese public holidays. https://api-docs.deepseek.com/quick_start/pricing The price depends on the hour. From 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, a million input tokens cost $0.30, or $0.006 from the cache, and a million output tokens $1.20.55 Every other hour costs half.55

Note

  1. DeepSeek, DeepSeek-V4.1-Flash model card, 10 September 2026. Instruct benchmarks at reasoning_effort=100, temperature 1.0, top_p 0.95; Terminal-Bench 2.1 Pass@1 in the Minimal mode of DeepSeek Harness with a 1M-token context. Base models on 25-shot SimpleQA-Verified (exact match): V4.1-Flash 42.3, V4-Pro 55.2. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16

  2. DeepSeek, DeepSeek-V4.1-Flash config.json (n_routed_experts 384, n_shared_experts 1, num_experts_per_tok 6, num_hidden_layers 40, sliding_window 128, fp8 weights in 32 x 32 blocks with fp4 experts, vision encoder of 32 layers). https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/config.json 2 3 4

  3. DeepSeek API Docs, DeepSeek-V4.1-Flash release note, 10 September 2026. https://api-docs.deepseek.com/news/news260910

  4. Arithmetic on the model card figure: 890 bytes x 1,048,576 tokens = 933,232,640 bytes, about 933 MB. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

  5. DeepSeek API Docs, Models and Pricing, as listed on 29 September 2026. Off-peak covers all other hours, including weekends and Chinese public holidays. https://api-docs.deepseek.com/quick_start/pricing 2 3

Altri modelli di DeepSeek

Tutti i modelli di DeepSeek