Salta la copertina

ExplainerIn inglese

Mixture of experts, token by token

How frontier models route each token to a few experts out of many, and why a model that large can still answer fast.

Indiceda 1 a 9

Dense

A language model writes one token at a time, and a dense model runs all of its weights to produce each one. A full stop costs as much arithmetic as a rare chemical name. The network has no way to skip the parts a given token does not need.

Non-embedding parameters
All the weights except the lookup table that turns each token into a vector.

Kaplan and colleagues put a number on that cost in 2020. A forward pass over one token takes about 2N floating-point operations, where N is the number of non-embedding parameters, plus a term that grows with the length of the context.11Kaplan et al. (2020), arXiv 2001.08361: forward-pass cost about 2N plus an attention term, 2 n_layer n_ctx d_model, that grows with context; training about 6N per token, times batch and steps for a full run; N is the non-embedding parameter count. https://arxiv.org/abs/2001.08361. Worked examples are arithmetic on that rule: 2 x 70B = 140B; 2 x 37B = 74B for DeepSeek-V3, attention excluded. Training takes about 6N per token, so a whole run costs about 6N times the number of tokens seen.11

Non-embedding parameters
All the weights except the lookup table that turns each token into a vector.

The rule is linear in size. A dense 70B model spends about 140 billion floating-point operations on every token it reads or writes (2 x 70B).11 Ten times the parameters means ten times the operations per token, in training and in use.

On the same hardware the larger model also answers more slowly, because each new token means reading every one of those weights from memory again.22Apple Machine Learning Research, post on LLMs with MLX on the M5, 19 November 2025 (MacBook Pro M5, 24 GB; Qwen3-30B-A3B-MLX-4bit footprint 17.31 GB; 153 GB/s on M5, 120 GB/s on M4; M5 vs M4 generation speedup 1.25x). https://machinelearning.apple.com/research/exploring-llms-mlx-m5. Arithmetic: 153 / 120 = 1.275.

GFLOP
A billion floating-point operations.

DeepSeek's engineers put the training cost of the dense LLaMA-405B at 2,448 GFLOPs per token.33DeepSeek, Insights paper on DeepSeek-V3, arXiv 2505.09343, Table 2 (training cost per token: DeepSeek-V3 250, LLaMA-405B 2448). https://arxiv.org/abs/2505.09343. Arithmetic: 6 x 405B = 2,430 and 2,448 / 2,430 = 1.007; 250 / 2,448 = 0.10; 671 / 405 = 1.66. Kaplan's 6N gives 6 x 405B, or 2,430 GFLOPs, less than 1% away from their figure.

GFLOP
A billion floating-point operations.

Experts and a router

Each layer of a Transformer has an attention block and a feed-forward block, and the feed-forward block works on each position separately. A mixture-of-experts (MoE) layer keeps attention as it is, shared by all tokens, and swaps the single feed-forward block for several independent ones called experts.

In Mixtral the "feedforward blocks are replaced by Mixture-of-Expert layers" in all 32 layers.44Mistral AI, Mixtral paper, arXiv 2401.04088. https://arxiv.org/abs/2401.04088 Each expert is an ordinary feed-forward network; in Mixtral and gpt-oss it is a SwiGLU block.4455OpenAI, gpt-oss model card, arXiv 2508.10925, and the official Hugging Face configs. https://arxiv.org/abs/2508.10925 GShard and the base Switch Transformer models made the same swap on every other layer.66Lepikhin et al. (2020), GShard, arXiv 2006.16668. https://arxiv.org/abs/2006.16668 77Fedus, Zoph, Shazeer (2021), Switch Transformer, arXiv 2101.03961. https://arxiv.org/abs/2101.03961

A small router decides which experts see each token. It is a linear layer: it multiplies the token's hidden vector by a matrix and returns one score per expert.7755

Softmax
A function that turns a list of scores into positive numbers that add up to one.

The router keeps the k highest scores, the top-k, and sets the rest aside. The kept scores then become weights. Mixtral does this with a softmax applied to its two surviving scores.44 Switch Transformer works the other way round, with a softmax over all experts first and only the top one kept.77

Softmax
A function that turns a list of scores into positive numbers that add up to one.

Each selected expert then processes the token, and its output is scaled by its weight. The scaled outputs are summed and added back to the token's vector through the residual connection.88Dai et al. (2024), DeepSeekMoE, arXiv 2401.06066. https://arxiv.org/abs/2401.06066 Experts that were not selected do no work for that token. Routing happens separately in every MoE layer, so one token can meet a different set of experts at each depth.

Switch Transformer sends each token to one expert, GShard and Mixtral to two, gpt-oss to four and DeepSeek-V3 to eight.7766445599DeepSeek (2024), DeepSeek-V3 technical report, arXiv 2412.19437, and its config.json (n_routed_experts 256, num_experts_per_tok 8, n_group 8, topk_group 4). https://arxiv.org/abs/2412.19437

DeepSeek-V3 also scores experts differently. It applies a sigmoid to each score on its own and then normalizes the weights among the eight experts it picked, while gpt-oss takes a softmax over its four.9955

Fig. 1
Token
Top k

Top 2 of 8 for “storm”: experts 3 and 7. The other 6 do no work for this token.

Scores are illustrative, not from a real model.

The router keeps the k highest scores and splits the token between those experts by weight.

Each Mixtral 8x7B layer holds 8 experts, and the router sends every token to 2 of them.44 The model has 46.7B parameters in total and uses 12.9B for each token.1010Mistral AI, Mixtral 8x7B announcement. https://mistral.ai/news/mixtral-of-experts

Eight experts of 7B would make 56B. The real total is lower because only the feed-forward blocks are repeated eight times, while attention and the embeddings exist once.1111Arithmetic: 8 x 7B = 56B, against 46.7B reported; 2 / 8 x 46.7B = 11.7B, against 12.9B active. Only the feed-forward blocks are replaced by experts (arXiv 2401.04088), so attention and embeddings are counted once. https://arxiv.org/abs/2401.04088

Two experts out of eight would be a quarter of 46.7B, or 11.7B. A token uses 12.9B because attention and the embeddings count in full every time.1111 Experts still hold most of the weights; in gpt-oss the MoE weights make up "90+% of the total parameter count".55

An idea from 1991

In 1991 Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton published "Adaptive Mixtures of Local Experts" in Neural Computation.1212Jacobs, Jordan, Nowlan, Hinton (1991), "Adaptive Mixtures of Local Experts", Neural Computation 3, 79-87. https://www.cs.toronto.edu/~hinton/absps/jjnh91.pdf A footnote in the paper dates the idea to 1988, when Jacobs and Hinton presented it at the Connectionist Summer School in Pittsburgh.1212

  1. 1991Adaptive Mixtures of Local Experts
    Several separate networks and a gating network that "decides which of the experts should be used for each training case".1212
  2. 2017Outrageously Large Neural Networks
    Shazeer and colleagues, Hinton among them, placed an MoE layer between stacked LSTM layers and chose experts with noisy top-k gating. The layer held up to 137B parameters.13
  3. 2020GShard
    The layer moved inside a Transformer, in place of every other feed-forward layer. A translation model of "beyond 600 billion parameters" trained "on 2048 TPU v3 accelerators in 4 days".66
  4. 2021Switch Transformer
    A single expert per token. The 64-expert Switch-Base reached the quality of T5-Base in "one-seventh the time".77
  5. 2021GLaM
    1.2T parameters, with 96.6B active per token.14
  6. 2023Mixtral 8x7B
    Each layer holds 8 experts, and the router sends every token to 2 of them.44
  7. 2024DeepSeek-V3
    671B parameters in total, 37B active for each token.99
  8. 2026Kimi K3
    2.8T parameters in total and 104B active.15

The 1991 paper's main change was to the error function. When the outputs are simply blended, the experts cooperate, and each can settle for patching what the others miss. Jacobs and colleagues made them compete instead, so that each expert had to produce "the whole of the output vector rather than a residual".1212 On a vowel discrimination task, the system split the problem into subtasks, each handled by a very simple expert network.1212

In Shazeer's noisy top-k gating, each expert's score gets random noise, scaled by a second learned projection, before the top k are kept and the rest are zeroed.13 A hierarchical version used up to 131,072 experts.13 The authors reported "greater than 1000x improvements in model capacity" at a small cost in efficiency.13

BLEU
Rates a machine translation by how much it overlaps with human reference translations. Higher is better.

On English to French translation the model scored 40.56 BLEU against 39.22 for GNMT.13

BLEU
Rates a machine translation by how much it overlaps with human reference translations. Higher is better.

GShard's model covered 100 languages into English, and each token went to at most two experts.66 Counted in training steps, Switch-Base reached T5-Base quality at step 60,000, where T5-Base needed 450,000.77 At 1.6T parameters and 2,048 experts, Switch-C was the largest version.77

Two numbers on every model card

An MoE model has two sizes. The total counts every parameter, all experts included. The active count covers what a single token uses: the shared parts plus the k experts the router picked. DeepSeek-V3 has "671B total parameters with 37B activated for each token".99

Model cards now print both numbers.1616Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. The report gives 1.04T total; the model card rounds to 1T. https://arxiv.org/abs/2507.20534 1717Alibaba Cloud, Qwen3.5 announcement, and the Qwen3.5-397B-A17B model card. https://www.alibabacloud.com/blog/602894 1818DeepSeek, DeepSeek-V4 preview release note, 24 April 2026, with the DeepSeek-V4-Pro model card and config. https://api-docs.deepseek.com/news/news260424 15

  • Mixtral 8x7B
    Total
    46.7B
    Active per token
    12.9B
    Experts
    8
    Experts per token
    2
  • DeepSeek-V3
    Total
    671B
    Active per token
    37B
    Experts
    256 + 1
    Experts per token
    8
  • Kimi K2
    Total
    1T
    Active per token
    32B
    Experts
    384 + 1
    Experts per token
    8
  • Qwen3.5
    Total
    397B
    Active per token
    17B
    Experts
    512 + 1
    Experts per token
    10
  • DeepSeek-V4-Pro
    Total
    1.6T
    Active per token
    49B
    Experts
    384 + 1
    Experts per token
    6
  • Kimi K3
    Total
    2.8T
    Active per token
    104B
    Experts
    896 + 2
    Experts per token
    16
  • Qwen3-30B-A3B
    Total
    30.5B
    Active per token
    3.3B
    Experts
    128
    Experts per token
    8

Some names carry the pair: Qwen3-30B-A3B has 30.5B parameters, 3.3B of them active.1919Qwen3-30B-A3B model card. https://huggingface.co/Qwen/Qwen3-30B-A3B. Arithmetic: 3.3B x 0.5 byte = 1.65 GB. Against DeepSeek-V3, Kimi K2 has 54% more total parameters and 13% fewer active ones, by its report's own count.1616

Since Mixtral the active share has fallen from about 28% to 3 or 4%.2020Arithmetic on the reported sizes: Mixtral 12.9 / 46.7 = 27.6%; DeepSeek-V3 37 / 671 = 5.5%; Kimi K2 32 / 1,040 = 3.1%; DeepSeek-V4-Pro 49 / 1,600 = 3.1%; Kimi K3 104 / 2,800 = 3.7%. Sources: https://mistral.ai/news/mixtral-of-experts, https://arxiv.org/abs/2412.19437, https://arxiv.org/abs/2507.20534, https://api-docs.deepseek.com/news/news260424, https://kimi.ai/blog/kimi-k3 A Mixtral token uses 12.9B of 46.7B parameters.1010 Grok-1, whose weights xAI released in March 2024, has 314B parameters with "25% of the weights active on a given token".2121xAI, Grok-1 open release, 17 March 2024. https://x.ai/news/grok-os

A DeepSeek-V3 token uses 37B of 671B, about 5.5%. Kimi K2 and DeepSeek-V4-Pro sit near 3%, and Kimi K3 at 3.7%.2020

Compute per token follows the active count. Kaplan's rule applied to DeepSeek-V3's 37B active parameters gives about 74 GFLOPs per token for the forward pass (2 x 37B), before attention.11 DeepSeek's own accounting puts V3's training cost at 250 GFLOPs per token. That is about a tenth of the dense LLaMA-405B's cost, for a model with about 1.7 times as many parameters.33

GLaM showed the same effect in 2021. It had 1.2T parameters with 96.6B active per token and needed 180 GFLOPs per token at inference, against 350 for the dense GPT-3.14

Qwen reports that its MoE base models reach "similar performance to Qwen3 dense base models with only 1/5 activated parameters".2222Qwen (2025), Qwen3 technical report, arXiv 2505.09388, for the MoE results and the removal of shared experts. https://arxiv.org/abs/2505.09388. The April 2025 release date comes from the Qwen3 model cards on Hugging Face, for example https://huggingface.co/Qwen/Qwen3-30B-A3B Google's Gemini 3 Pro model card credits sparse MoE with letting a model "decouple total model capacity from computation and serving cost per token".2323Google DeepMind, Gemini 3 Pro model card. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf

Fig. 2
Each row is one model. The clay dot counts the parameters one token uses, the ink dot all the parameters the model holds. The scale is logarithmic.

Speed follows active, memory follows total

Memory bandwidth
The rate at which a chip reads data from memory, measured in gigabytes per second.

To produce a new token, a model reads the weights it uses from memory, and that reading sets the pace more than the arithmetic does. Apple's machine learning team writes that generating subsequent tokens "is bounded by memory bandwidth, rather than compute ability".22

Memory bandwidth
The rate at which a chip reads data from memory, measured in gigabytes per second.

An MoE model reads only its active weights for each token. It still keeps every expert loaded, because the next token may pick different ones. The Mixtral paper puts the memory cost of serving the model as "proportional to its sparse parameter count, 47B", where sparse parameter count is Mistral's term for the total.44

OpenAI shrank the stored size of gpt-oss by keeping its MoE weights in MXFP4. That lets gpt-oss-120b fit "on a single 80GB GPU" and gpt-oss-20b run "on systems with as little as 16GB memory".55

On a server GPU

SGLang
A serving engine, the software that loads a model and answers requests.

Both effects show up in one benchmark, run by Qwen on an NVIDIA H20 GPU with 96 GB. With SGLang in BF16 and one request at a time, Qwen3-30B-A3B generates 137 tokens per second. The dense Qwen3-32B generates 21, so the MoE model is about 6.6 times faster (137.18 / 20.72).2424Qwen, official speed benchmark, NVIDIA H20 96GB, batch size 1, input length 1, 2,048 generated tokens. Speed is (prompt + generated tokens) / time. Figures by backend, Qwen3-30B-A3B against Qwen3-32B: SGLang BF16, 137.18 against 20.72 tokens per second; SGLang FP8, 155.55 against 46.17; plain Hugging Face Transformers BF16, 1.89 against 26.24, with the GPTQ-INT4 MoE entry marked "MoE Kernel Unsupported"; GPU memory in Transformers BF16, 58,462 MB against 62,751 MB. https://qwen.readthedocs.io/en/latest/getting_started/speed_benchmark.html

SGLang
A serving engine, the software that loads a model and answers requests.
BF16, FP8, MXFP4
BF16 is a 16-bit number format and FP8 an 8-bit one. MXFP4 is a 4-bit format that takes about 4.25 bits per parameter.55

In FP8 the figures are 155.55 and 46.17 tokens per second, and the MoE model is still 3.4 times faster (155.55 / 46.17).2424 The memory the two need is close. Loaded in Hugging Face Transformers on the same GPU, they occupy 58.5 GB and 62.8 GB.2424

BF16, FP8, MXFP4
BF16 is a 16-bit number format and FP8 an 8-bit one. MXFP4 is a 4-bit format that takes about 4.25 bits per parameter.55
Kernels
The low-level GPU routines that carry out each step of a model.

Speed also depends on the software. Fast MoE inference needs kernels written for the expert step. Plain Hugging Face Transformers in BF16, without such kernels, generates only 1.89 tokens per second with Qwen3-30B-A3B, while the dense Qwen3-32B reaches 26.24.2424

Kernels
The low-level GPU routines that carry out each step of a model.

The same table has no speed figure for the 4-bit GPTQ version of the MoE model. Its entry says the MoE kernel is unsupported.2424

On a laptop

Unified memory
The pool of memory the CPU and GPU share.

Apple ran Qwen3-30B-A3B at 4-bit on a MacBook Pro with an M5 chip and 24 GB of unified memory, and the model took 17.31 GB.22 The M5 reads memory at 153 GB/s.22

Unified memory
The pool of memory the CPU and GPU share.

At 4 bits a parameter takes half a byte, so the 3.3B active parameters come to about 1.65 GB per token.1919 Dividing one by the other gives a ceiling of about 93 tokens per second (153 / 1.65). A dense 32B model at 4-bit reads about 16 GB per token, for a ceiling near 10 tokens per second (153 / 16).2525Arithmetic, not a measurement: 153 GB/s / 1.65 GB = 93 tokens per second; 32B x 0.5 byte = 16 GB; 153 / 16 = 9.6 tokens per second. KV cache, quantization scales and router overhead are ignored. Inputs: https://machinelearning.apple.com/research/exploring-llms-mlx-m5 and https://huggingface.co/Qwen/Qwen3-30B-A3B

KV cache
The attention data kept for earlier tokens.

These ceilings are arithmetic, not measurements. Real speeds are lower, because the KV cache also has to be read, along with other overheads. Apple published only relative speedups and memory footprints, with no absolute speeds.2525

KV cache
The attention data kept for earlier tokens.

Apple did compare the M5 with the M4, which reads memory at 120 GB/s.22 The M5 has 1.275 times the bandwidth (153 / 120), and it generated tokens with this model 1.25 times faster.22

Fig. 3
Qwen's benchmark on one NVIDIA H20. Speed comes from SGLang and memory from Hugging Face Transformers, both in BF16.

Keeping every expert busy

Left alone, a router can send most tokens to a few experts while the rest get little. In the DeepSeek-V3 report, an unbalanced expert load leads to "routing collapse".99

Uneven load is also a hardware problem. In a large model the experts sit on different GPUs, and the Mixtral paper stresses that "it is essential to distribute the workload evenly across the GPUs".44

A hard limit

Capacity factor
How many tokens an expert may take, as a multiple of an even share of the batch.

Switch Transformer put a hard limit on each expert. Every expert gets a capacity, computed as (tokens per batch / number of experts) x capacity factor.77 A factor of 1.0 gives each expert exactly an even share of the batch. Tokens beyond the limit skip the expert and are "passed directly to the next layer through the residual connection".77

Capacity factor
How many tokens an expert may take, as a multiple of an even share of the batch.

Switch tried factors from 1.0 to 2.0, and values from 1.0 to 1.25 worked better at scale.77

A penalty

Load-balancing loss
An extra term in the training objective that grows when experts are used unevenly.

Switch also adds a load-balancing loss. It uses two numbers per expert: the fraction of tokens the expert receives and the average probability the router gives it. The loss multiplies them for each expert and sums the products. The sum is smallest when every expert gets an equal share, and the term enters training with a weight of 0.01.77

Load-balancing loss
An extra term in the training objective that grows when experts are used unevenly.
Router z-loss
Routers use exponentials that "exacerbate roundoff errors". The z-loss penalizes large router scores.

ST-MoE later added a router z-loss, with a weight of 0.001.2626Zoph et al. (2022), ST-MoE, arXiv 2202.08906. https://arxiv.org/abs/2202.08906

Router z-loss
Routers use exponentials that "exacerbate roundoff errors". The z-loss penalizes large router scores.

A balancing loss also pulls the model away from its main task, and the DeepSeek-V3 report warns that "Too large an auxiliary loss will impair the model performance".99

A bias

DeepSeek-V3 balances mainly with a bias for each expert instead. The bias is added to the expert's score only when the router picks the top k; the weights that mix the outputs ignore it. When an expert is overloaded its bias goes down by 0.001, and when it is underloaded the bias goes up by the same step.99

That step size applied for the first 14.3T training tokens. After that it was set to zero, which stops the biases changing, but they stay in use. A small sequence-level balancing loss with a weight of 0.0001 runs alongside the biases as a complement.99 V3 drops no tokens, in training or in inference.99

An earlier paper by DeepSeek authors tested the bias method on a 1B model. Their measure of imbalance is the load of the busiest expert minus the mean load, divided by the mean.2727Wang et al. (2024), auxiliary-loss-free load balancing, arXiv 2408.15664. The imbalance measure is MaxVio_global. https://arxiv.org/abs/2408.15664

Perplexity
A measure of how well a model predicts text, where lower is better.

With a balancing loss it came to 0.72, so the busiest expert carried 72% more tokens than the average one. With the bias it came to 0.04.2727 Perplexity improved slightly, from 9.56 to 9.50.2727 A 3B model went from 0.52 to 0.04, with perplexity from 7.97 to 7.92.2727

Perplexity
A measure of how well a model predicts text, where lower is better.
Fig. 4
Above, a schematic of expert load around the mean. Below, the imbalance measured by DeepSeek authors with a balancing loss and with the bias.

Smaller experts, more of them

DeepSeekMoE, published in January 2024, changed the size of the experts. It split each expert into m smaller ones by cutting the feed-forward hidden size to 1/m, then activated m times as many, so compute per token stays the same.88 At 16.4B total and about 2.8B active, DeepSeekMoE 16B performed comparably to LLaMA2 7B with about 40% of its compute.88

Smaller experts give the router far more combinations to choose from. Mixtral's router can form 28 different pairs from its 8 experts. DeepSeek-V3's router works with 256 routed experts and takes 8. Picking 8 of 256 allows about 410 trillion sets in principle.2828Arithmetic: 8 choose 2 = 28; 256 choose 8 = 409,663,695,276,000, before V3's group limit. Expert counts from https://arxiv.org/abs/2401.04088 and https://arxiv.org/abs/2412.19437 V3 trims that by limiting each token to 4 of its 8 expert groups.99

Shared experts
Experts that are "always activated, aiming at capturing and consolidating common knowledge". Every token passes through them, and the router chooses only among the others.

The same paper set some experts apart as shared experts.88

Shared experts
Experts that are "always activated, aiming at capturing and consolidating common knowledge". Every token passes through them, and the router chooses only among the others.

Per layer:

  • Mixtral has 8 experts and uses 2.44
  • DeepSeek-V3 has 256 routed experts plus 1 shared, and uses 8 of the routed ones.99
  • Qwen3.5 has 512 plus 1 shared and uses 10.1717
  • Kimi K3 has 896 plus 2 shared and uses 16.15
  • DeepSeek-V4-Pro moved the other way on k, with 384 routed experts plus 1 shared and 6 active.1818

Kimi K2's technical report tested the trend. It defines sparsity as total experts divided by active experts. With the active count held fixed, more experts gave lower loss.1616

A model at sparsity 8 needed 1.69 times as much compute to reach the same loss as one at sparsity 48, which has 384 experts with 8 active.1616

Qwen3 dropped shared experts in April 2025, and Qwen3.5 has one again.22221717 Mixtral and gpt-oss do not use them.4455

Fig. 5
One square per expert in a single layer, all drawn the same size. Clay marks the experts one token uses; which ones light up is only an example.

What the experts actually learn

In an MoE model, "expert" is the name of a slot the router can choose. No one assigns the slots a subject, and what each one ends up handling is found by measuring.

Mixtral's authors checked how the model routes text from ArXiv, PubMed, PhilPapers, GitHub and other sources. They "do not observe obvious patterns in the assignment of experts" by topic, and only DM Mathematics showed a marginally different distribution.44

Some of it follows syntax. The token "self" in Python and the word "Question" in English often go to the same expert, and "indentation tokens are always assigned to the same experts".44 Consecutive tokens also tend to reuse an expert: at layer 15 the first choice repeats in 24 to 28% of cases, against 12.5% for random routing.44

In 2017 Shazeer's team saw its experts become "highly specialized based on syntax and semantics".13 In ST-MoE, encoder experts split by token type, such as punctuation, verbs, proper names and numbers. In multilingual training they "do not exhibit language specialization".2626

The authors of OLMoE, an open model from AI2, found strong domain experts in it. On arXiv text, the first expert in its layer 0 is "nearly 100% specialized".2929AI2 (2024), OLMoE-1B-7B, arXiv 2409.02060. https://arxiv.org/abs/2409.02060 The same paper measured Mixtral's experts as "activated close to the uniform routing baseline".2929

Other teams tie specialization to balancing. Qwen's team found that balancing within each sequence suppresses it, while balancing over the global batch lets domain experts appear.3030Qwen (2025), global-batch load balancing, arXiv 2501.11873. https://arxiv.org/abs/2501.11873 The DeepSeek-V3 report finds "greater expert specialization patterns" with the bias method than with an auxiliary loss.99

The price

Memory

Every expert has to stay in memory, even though a token uses only a few. DeepSeek-V3 ships its weights in FP8, one byte per parameter, so the main weights alone take about 671 GB (671B x 1 byte), before the KV cache and activations.3131DeepSeek-V3 model card: 671B main model weights, "we only provide FP8 weights". Arithmetic: 671B x 1 byte = 671 GB. https://huggingface.co/deepseek-ai/DeepSeek-V3

Traffic

All-to-all
An exchange in which every GPU in a group sends data to every other one.

A model that size is spread across many GPUs, each holding some of the experts. Tokens travel to the GPU that holds their expert and the results travel back, through all-to-all exchanges at every MoE layer.6699

All-to-all
An exchange in which every GPU in a group sends data to every other one.
Nodes
Servers. Each one holds several GPUs.

Shazeer's paper had already warned that "Network bandwidth can be a bottleneck".13 DeepSeek-V3 limits each token to experts on at most 4 nodes. It also reserves 20 streaming multiprocessors for the exchange, and overlaps the exchange with computation.99

Streaming multiprocessors
The GPU's compute units.
Nodes
Servers. Each one holds several GPUs.
Streaming multiprocessors
The GPU's compute units.

Batches

Large batches are another requirement. When each token picks a few experts out of many, each expert receives only a small part of every batch.13 An expert that gets few tokens still has to read all of its weights from memory to process them, so the GPU spends more time reading than computing.

Arithmetic intensity
Enough computation for every byte read from memory.

Mixtral's authors call MoE layers "more suitable for batched workloads", where the GPU can reach a good "arithmetic intensity".44 DeepSeek's notes on its inference system add that the model's "high sparsity necessitates an extremely large overall batch size".3232DeepSeek, open-infra-index, day 6 (inference system overview, February 2025). https://github.com/deepseek-ai/open-infra-index

Arithmetic intensity
Enough computation for every byte read from memory.

The V3 report describes decoding units of 40 nodes and 320 GPUs. "Each GPU hosts only one expert", and 64 GPUs hold redundant copies of high-load experts along with the shared experts.99 The other 256 GPUs match the 256 routed experts one to one (256 + 64 = 320).

Closed models

Which closed frontier models use MoE is only partly public. Google says so for Gemini. The Gemini 2.5 report calls its models "sparse mixture-of-experts (MoE) transformers", and the Gemini 3 Pro card says the same.3333Google (2025), Gemini 2.5 technical report, arXiv 2507.06261. https://arxiv.org/abs/2507.06261 2323 OpenAI's open-weight gpt-oss models are MoE.55

The GPT-4 technical report gives "no further details about the architecture (including model size)".3434OpenAI (2023), GPT-4 technical report, arXiv 2303.08774. https://arxiv.org/abs/2303.08774 The GPT-5 system card describes a "real-time router that quickly decides which model to use", which is a choice between whole models.3535OpenAI, GPT-5 system card, 13 August 2025. https://cdn.openai.com/gpt-5-system-card.pdf

Note

  1. Kaplan et al. (2020), arXiv 2001.08361: forward-pass cost about 2N plus an attention term, 2 n_layer n_ctx d_model, that grows with context; training about 6N per token, times batch and steps for a full run; N is the non-embedding parameter count. https://arxiv.org/abs/2001.08361. Worked examples are arithmetic on that rule: 2 x 70B = 140B; 2 x 37B = 74B for DeepSeek-V3, attention excluded. 2 3 4

  2. Apple Machine Learning Research, post on LLMs with MLX on the M5, 19 November 2025 (MacBook Pro M5, 24 GB; Qwen3-30B-A3B-MLX-4bit footprint 17.31 GB; 153 GB/s on M5, 120 GB/s on M4; M5 vs M4 generation speedup 1.25x). https://machinelearning.apple.com/research/exploring-llms-mlx-m5. Arithmetic: 153 / 120 = 1.275. 2 3 4 5 6

  3. DeepSeek, Insights paper on DeepSeek-V3, arXiv 2505.09343, Table 2 (training cost per token: DeepSeek-V3 250, LLaMA-405B 2448). https://arxiv.org/abs/2505.09343. Arithmetic: 6 x 405B = 2,430 and 2,448 / 2,430 = 1.007; 250 / 2,448 = 0.10; 671 / 405 = 1.66. 2

  4. Mistral AI, Mixtral paper, arXiv 2401.04088. https://arxiv.org/abs/2401.04088 2 3 4 5 6 7 8 9 10 11 12 13 14

  5. OpenAI, gpt-oss model card, arXiv 2508.10925, and the official Hugging Face configs. https://arxiv.org/abs/2508.10925 2 3 4 5 6 7 8 9

  6. Lepikhin et al. (2020), GShard, arXiv 2006.16668. https://arxiv.org/abs/2006.16668 2 3 4 5

  7. Fedus, Zoph, Shazeer (2021), Switch Transformer, arXiv 2101.03961. https://arxiv.org/abs/2101.03961 2 3 4 5 6 7 8 9 10 11

  8. Dai et al. (2024), DeepSeekMoE, arXiv 2401.06066. https://arxiv.org/abs/2401.06066 2 3 4

  9. DeepSeek (2024), DeepSeek-V3 technical report, arXiv 2412.19437, and its config.json (n_routed_experts 256, num_experts_per_tok 8, n_group 8, topk_group 4). https://arxiv.org/abs/2412.19437 2 3 4 5 6 7 8 9 10 11 12 13 14 15

  10. Mistral AI, Mixtral 8x7B announcement. https://mistral.ai/news/mixtral-of-experts 2

  11. Arithmetic: 8 x 7B = 56B, against 46.7B reported; 2 / 8 x 46.7B = 11.7B, against 12.9B active. Only the feed-forward blocks are replaced by experts (arXiv 2401.04088), so attention and embeddings are counted once. https://arxiv.org/abs/2401.04088 2

  12. Jacobs, Jordan, Nowlan, Hinton (1991), "Adaptive Mixtures of Local Experts", Neural Computation 3, 79-87. https://www.cs.toronto.edu/~hinton/absps/jjnh91.pdf 2 3 4 5

  13. Shazeer et al. (2017), "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer", arXiv 1701.06538. https://arxiv.org/abs/1701.06538 2 3 4 5 6 7 8

  14. Du et al. (2021), GLaM, arXiv 2112.06905. https://arxiv.org/abs/2112.06905 2

  15. Moonshot AI, Kimi K3 announcement and model card. https://kimi.ai/blog/kimi-k3 2 3

  16. Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. The report gives 1.04T total; the model card rounds to 1T. https://arxiv.org/abs/2507.20534 2 3 4

  17. Alibaba Cloud, Qwen3.5 announcement, and the Qwen3.5-397B-A17B model card. https://www.alibabacloud.com/blog/602894 2 3

  18. DeepSeek, DeepSeek-V4 preview release note, 24 April 2026, with the DeepSeek-V4-Pro model card and config. https://api-docs.deepseek.com/news/news260424 2

  19. Qwen3-30B-A3B model card. https://huggingface.co/Qwen/Qwen3-30B-A3B. Arithmetic: 3.3B x 0.5 byte = 1.65 GB. 2

  20. Arithmetic on the reported sizes: Mixtral 12.9 / 46.7 = 27.6%; DeepSeek-V3 37 / 671 = 5.5%; Kimi K2 32 / 1,040 = 3.1%; DeepSeek-V4-Pro 49 / 1,600 = 3.1%; Kimi K3 104 / 2,800 = 3.7%. Sources: https://mistral.ai/news/mixtral-of-experts, https://arxiv.org/abs/2412.19437, https://arxiv.org/abs/2507.20534, https://api-docs.deepseek.com/news/news260424, https://kimi.ai/blog/kimi-k3 2

  21. xAI, Grok-1 open release, 17 March 2024. https://x.ai/news/grok-os

  22. Qwen (2025), Qwen3 technical report, arXiv 2505.09388, for the MoE results and the removal of shared experts. https://arxiv.org/abs/2505.09388. The April 2025 release date comes from the Qwen3 model cards on Hugging Face, for example https://huggingface.co/Qwen/Qwen3-30B-A3B 2

  23. Google DeepMind, Gemini 3 Pro model card. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf 2

  24. Qwen, official speed benchmark, NVIDIA H20 96GB, batch size 1, input length 1, 2,048 generated tokens. Speed is (prompt + generated tokens) / time. Figures by backend, Qwen3-30B-A3B against Qwen3-32B: SGLang BF16, 137.18 against 20.72 tokens per second; SGLang FP8, 155.55 against 46.17; plain Hugging Face Transformers BF16, 1.89 against 26.24, with the GPTQ-INT4 MoE entry marked "MoE Kernel Unsupported"; GPU memory in Transformers BF16, 58,462 MB against 62,751 MB. https://qwen.readthedocs.io/en/latest/getting_started/speed_benchmark.html 2 3 4 5

  25. Arithmetic, not a measurement: 153 GB/s / 1.65 GB = 93 tokens per second; 32B x 0.5 byte = 16 GB; 153 / 16 = 9.6 tokens per second. KV cache, quantization scales and router overhead are ignored. Inputs: https://machinelearning.apple.com/research/exploring-llms-mlx-m5 and https://huggingface.co/Qwen/Qwen3-30B-A3B 2

  26. Zoph et al. (2022), ST-MoE, arXiv 2202.08906. https://arxiv.org/abs/2202.08906 2

  27. Wang et al. (2024), auxiliary-loss-free load balancing, arXiv 2408.15664. The imbalance measure is MaxVio_global. https://arxiv.org/abs/2408.15664 2 3 4

  28. Arithmetic: 8 choose 2 = 28; 256 choose 8 = 409,663,695,276,000, before V3's group limit. Expert counts from https://arxiv.org/abs/2401.04088 and https://arxiv.org/abs/2412.19437

  29. AI2 (2024), OLMoE-1B-7B, arXiv 2409.02060. https://arxiv.org/abs/2409.02060 2

  30. Qwen (2025), global-batch load balancing, arXiv 2501.11873. https://arxiv.org/abs/2501.11873

  31. DeepSeek-V3 model card: 671B main model weights, "we only provide FP8 weights". Arithmetic: 671B x 1 byte = 671 GB. https://huggingface.co/deepseek-ai/DeepSeek-V3

  32. DeepSeek, open-infra-index, day 6 (inference system overview, February 2025). https://github.com/deepseek-ai/open-infra-index

  33. Google (2025), Gemini 2.5 technical report, arXiv 2507.06261. https://arxiv.org/abs/2507.06261

  34. OpenAI (2023), GPT-4 technical report, arXiv 2303.08774. https://arxiv.org/abs/2303.08774

  35. OpenAI, GPT-5 system card, 13 August 2025. https://cdn.openai.com/gpt-5-system-card.pdf