Mixture of experts, token by token
How frontier models route each token to a few experts out of many, and why a model that large can still answer fast.
Contents1 to 9
Dense
A language model writes one token at a time, and a dense model runs all of its weights to produce each one. A full stop costs as much arithmetic as a rare chemical name. The network has no way to skip the parts a given token does not need.
Kaplan and colleagues put a number on that cost in 2020. A forward pass over one token takes about 2N floating-point operations, where N is the number of non-embedding parameters, plus a term that grows with the length of the context.11Kaplan et al. (2020), arXiv 2001.08361: forward-pass cost about 2N plus an attention term, 2 n_layer n_ctx d_model, that grows with context; training about 6N per token, times batch and steps for a full run; N is the non-embedding parameter count. https://arxiv.org/abs/2001.08361. Worked examples are arithmetic on that rule: 2 x 70B = 140B; 2 x 37B = 74B for DeepSeek-V3, attention excluded. Training takes about 6N per token, so a whole run costs about 6N times the number of tokens seen.11
The rule is linear in size. A dense 70B model spends about 140 billion floating-point operations on every token it reads or writes (2 x 70B).11 Ten times the parameters means ten times the operations per token, in training and in use.
On the same hardware the larger model also answers more slowly, because each new token means reading every one of those weights from memory again.22Apple Machine Learning Research, post on LLMs with MLX on the M5, 19 November 2025 (MacBook Pro M5, 24 GB; Qwen3-30B-A3B-MLX-4bit footprint 17.31 GB; 153 GB/s on M5, 120 GB/s on M4; M5 vs M4 generation speedup 1.25x). https://machinelearning.apple.com/research/exploring-llms-mlx-m5. Arithmetic: 153 / 120 = 1.275.
DeepSeek's engineers put the training cost of the dense LLaMA-405B at 2,448 GFLOPs per token.33DeepSeek, Insights paper on DeepSeek-V3, arXiv 2505.09343, Table 2 (training cost per token: DeepSeek-V3 250, LLaMA-405B 2448). https://arxiv.org/abs/2505.09343. Arithmetic: 6 x 405B = 2,430 and 2,448 / 2,430 = 1.007; 250 / 2,448 = 0.10; 671 / 405 = 1.66. Kaplan's 6N gives 6 x 405B, or 2,430 GFLOPs, less than 1% away from their figure.
Experts and a router
Each layer of a Transformer has an attention block and a feed-forward block, and the feed-forward block works on each position separately. A mixture-of-experts (MoE) layer keeps attention as it is, shared by all tokens, and swaps the single feed-forward block for several independent ones called experts.
In Mixtral the "feedforward blocks are replaced by Mixture-of-Expert layers" in all 32 layers.44Mistral AI, Mixtral paper, arXiv 2401.04088. https://arxiv.org/abs/2401.04088 Each expert is an ordinary feed-forward network; in Mixtral and gpt-oss it is a SwiGLU block.4455OpenAI, gpt-oss model card, arXiv 2508.10925, and the official Hugging Face configs. https://arxiv.org/abs/2508.10925 GShard and the base Switch Transformer models made the same swap on every other layer.66Lepikhin et al. (2020), GShard, arXiv 2006.16668. https://arxiv.org/abs/2006.16668 77Fedus, Zoph, Shazeer (2021), Switch Transformer, arXiv 2101.03961. https://arxiv.org/abs/2101.03961
A small router decides which experts see each token. It is a linear layer: it multiplies the token's hidden vector by a matrix and returns one score per expert.7755
The router keeps the k highest scores, the top-k, and sets the rest aside. The kept scores then become weights. Mixtral does this with a softmax applied to its two surviving scores.44 Switch Transformer works the other way round, with a softmax over all experts first and only the top one kept.77
Each selected expert then processes the token, and its output is scaled by its weight. The scaled outputs are summed and added back to the token's vector through the residual connection.88Dai et al. (2024), DeepSeekMoE, arXiv 2401.06066. https://arxiv.org/abs/2401.06066 Experts that were not selected do no work for that token. Routing happens separately in every MoE layer, so one token can meet a different set of experts at each depth.
Switch Transformer sends each token to one expert, GShard and Mixtral to two, gpt-oss to four and DeepSeek-V3 to eight.7766445599DeepSeek (2024), DeepSeek-V3 technical report, arXiv 2412.19437, and its config.json (n_routed_experts 256, num_experts_per_tok 8, n_group 8, topk_group 4). https://arxiv.org/abs/2412.19437
DeepSeek-V3 also scores experts differently. It applies a sigmoid to each score on its own and then normalizes the weights among the eight experts it picked, while gpt-oss takes a softmax over its four.9955
Top 2 of 8 for “storm”: experts 3 and 7. The other 6 do no work for this token.
Scores are illustrative, not from a real model.
Each Mixtral 8x7B layer holds 8 experts, and the router sends every token to 2 of them.44 The model has 46.7B parameters in total and uses 12.9B for each token.1010Mistral AI, Mixtral 8x7B announcement. https://mistral.ai/news/mixtral-of-experts
Eight experts of 7B would make 56B. The real total is lower because only the feed-forward blocks are repeated eight times, while attention and the embeddings exist once.1111Arithmetic: 8 x 7B = 56B, against 46.7B reported; 2 / 8 x 46.7B = 11.7B, against 12.9B active. Only the feed-forward blocks are replaced by experts (arXiv 2401.04088), so attention and embeddings are counted once. https://arxiv.org/abs/2401.04088
Two experts out of eight would be a quarter of 46.7B, or 11.7B. A token uses 12.9B because attention and the embeddings count in full every time.1111 Experts still hold most of the weights; in gpt-oss the MoE weights make up "90+% of the total parameter count".55
An idea from 1991
In 1991 Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton published "Adaptive Mixtures of Local Experts" in Neural Computation.1212Jacobs, Jordan, Nowlan, Hinton (1991), "Adaptive Mixtures of Local Experts", Neural Computation 3, 79-87. https://www.cs.toronto.edu/~hinton/absps/jjnh91.pdf A footnote in the paper dates the idea to 1988, when Jacobs and Hinton presented it at the Connectionist Summer School in Pittsburgh.1212
- 1991Adaptive Mixtures of Local Experts
- 2017Outrageously Large Neural NetworksShazeer and colleagues, Hinton among them, placed an MoE layer between stacked LSTM layers and chose experts with noisy top-k gating. The layer held up to 137B parameters.13
- 2020GShard
- 2021Switch Transformer
- 2021GLaM1.2T parameters, with 96.6B active per token.14
- 2023Mixtral 8x7B
- 2024DeepSeek-V3
- 2026Kimi K32.8T parameters in total and 104B active.15
The 1991 paper's main change was to the error function. When the outputs are simply blended, the experts cooperate, and each can settle for patching what the others miss. Jacobs and colleagues made them compete instead, so that each expert had to produce "the whole of the output vector rather than a residual".1212 On a vowel discrimination task, the system split the problem into subtasks, each handled by a very simple expert network.1212
In Shazeer's noisy top-k gating, each expert's score gets random noise, scaled by a second learned projection, before the top k are kept and the rest are zeroed.13 A hierarchical version used up to 131,072 experts.13 The authors reported "greater than 1000x improvements in model capacity" at a small cost in efficiency.13
On English to French translation the model scored 40.56 BLEU against 39.22 for GNMT.13
GShard's model covered 100 languages into English, and each token went to at most two experts.66 Counted in training steps, Switch-Base reached T5-Base quality at step 60,000, where T5-Base needed 450,000.77 At 1.6T parameters and 2,048 experts, Switch-C was the largest version.77
Two numbers on every model card
An MoE model has two sizes. The total counts every parameter, all experts included. The active count covers what a single token uses: the shared parts plus the k experts the router picked. DeepSeek-V3 has "671B total parameters with 37B activated for each token".99
Model cards now print both numbers.1616Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. The report gives 1.04T total; the model card rounds to 1T. https://arxiv.org/abs/2507.20534 1717Alibaba Cloud, Qwen3.5 announcement, and the Qwen3.5-397B-A17B model card. https://www.alibabacloud.com/blog/602894 1818DeepSeek, DeepSeek-V4 preview release note, 24 April 2026, with the DeepSeek-V4-Pro model card and config. https://api-docs.deepseek.com/news/news260424 15
- Mixtral 8x7B
- Total
- 46.7B
- Active per token
- 12.9B
- Experts
- 8
- Experts per token
- 2
- DeepSeek-V3
- Total
- 671B
- Active per token
- 37B
- Experts
- 256 + 1
- Experts per token
- 8
- Kimi K2
- Total
- 1T
- Active per token
- 32B
- Experts
- 384 + 1
- Experts per token
- 8
- Qwen3.5
- Total
- 397B
- Active per token
- 17B
- Experts
- 512 + 1
- Experts per token
- 10
- DeepSeek-V4-Pro
- Total
- 1.6T
- Active per token
- 49B
- Experts
- 384 + 1
- Experts per token
- 6
- Kimi K3
- Total
- 2.8T
- Active per token
- 104B
- Experts
- 896 + 2
- Experts per token
- 16
- Qwen3-30B-A3B
- Total
- 30.5B
- Active per token
- 3.3B
- Experts
- 128
- Experts per token
- 8
Some names carry the pair: Qwen3-30B-A3B has 30.5B parameters, 3.3B of them active.1919Qwen3-30B-A3B model card. https://huggingface.co/Qwen/Qwen3-30B-A3B. Arithmetic: 3.3B x 0.5 byte = 1.65 GB. Against DeepSeek-V3, Kimi K2 has 54% more total parameters and 13% fewer active ones, by its report's own count.1616
Since Mixtral the active share has fallen from about 28% to 3 or 4%.2020Arithmetic on the reported sizes: Mixtral 12.9 / 46.7 = 27.6%; DeepSeek-V3 37 / 671 = 5.5%; Kimi K2 32 / 1,040 = 3.1%; DeepSeek-V4-Pro 49 / 1,600 = 3.1%; Kimi K3 104 / 2,800 = 3.7%. Sources: https://mistral.ai/news/mixtral-of-experts, https://arxiv.org/abs/2412.19437, https://arxiv.org/abs/2507.20534, https://api-docs.deepseek.com/news/news260424, https://kimi.ai/blog/kimi-k3 A Mixtral token uses 12.9B of 46.7B parameters.1010 Grok-1, whose weights xAI released in March 2024, has 314B parameters with "25% of the weights active on a given token".2121xAI, Grok-1 open release, 17 March 2024. https://x.ai/news/grok-os
A DeepSeek-V3 token uses 37B of 671B, about 5.5%. Kimi K2 and DeepSeek-V4-Pro sit near 3%, and Kimi K3 at 3.7%.2020
Compute per token follows the active count. Kaplan's rule applied to DeepSeek-V3's 37B active parameters gives about 74 GFLOPs per token for the forward pass (2 x 37B), before attention.11 DeepSeek's own accounting puts V3's training cost at 250 GFLOPs per token. That is about a tenth of the dense LLaMA-405B's cost, for a model with about 1.7 times as many parameters.33
GLaM showed the same effect in 2021. It had 1.2T parameters with 96.6B active per token and needed 180 GFLOPs per token at inference, against 350 for the dense GPT-3.14
Qwen reports that its MoE base models reach "similar performance to Qwen3 dense base models with only 1/5 activated parameters".2222Qwen (2025), Qwen3 technical report, arXiv 2505.09388, for the MoE results and the removal of shared experts. https://arxiv.org/abs/2505.09388. The April 2025 release date comes from the Qwen3 model cards on Hugging Face, for example https://huggingface.co/Qwen/Qwen3-30B-A3B Google's Gemini 3 Pro model card credits sparse MoE with letting a model "decouple total model capacity from computation and serving cost per token".2323Google DeepMind, Gemini 3 Pro model card. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf
Speed follows active, memory follows total
To produce a new token, a model reads the weights it uses from memory, and that reading sets the pace more than the arithmetic does. Apple's machine learning team writes that generating subsequent tokens "is bounded by memory bandwidth, rather than compute ability".22
An MoE model reads only its active weights for each token. It still keeps every expert loaded, because the next token may pick different ones. The Mixtral paper puts the memory cost of serving the model as "proportional to its sparse parameter count, 47B", where sparse parameter count is Mistral's term for the total.44
OpenAI shrank the stored size of gpt-oss by keeping its MoE weights in MXFP4. That lets gpt-oss-120b fit "on a single 80GB GPU" and gpt-oss-20b run "on systems with as little as 16GB memory".55
On a server GPU
Both effects show up in one benchmark, run by Qwen on an NVIDIA H20 GPU with 96 GB. With SGLang in BF16 and one request at a time, Qwen3-30B-A3B generates 137 tokens per second. The dense Qwen3-32B generates 21, so the MoE model is about 6.6 times faster (137.18 / 20.72).2424Qwen, official speed benchmark, NVIDIA H20 96GB, batch size 1, input length 1, 2,048 generated tokens. Speed is (prompt + generated tokens) / time. Figures by backend, Qwen3-30B-A3B against Qwen3-32B: SGLang BF16, 137.18 against 20.72 tokens per second; SGLang FP8, 155.55 against 46.17; plain Hugging Face Transformers BF16, 1.89 against 26.24, with the GPTQ-INT4 MoE entry marked "MoE Kernel Unsupported"; GPU memory in Transformers BF16, 58,462 MB against 62,751 MB. https://qwen.readthedocs.io/en/latest/getting_started/speed_benchmark.html
In FP8 the figures are 155.55 and 46.17 tokens per second, and the MoE model is still 3.4 times faster (155.55 / 46.17).2424 The memory the two need is close. Loaded in Hugging Face Transformers on the same GPU, they occupy 58.5 GB and 62.8 GB.2424
Speed also depends on the software. Fast MoE inference needs kernels written for the expert step. Plain Hugging Face Transformers in BF16, without such kernels, generates only 1.89 tokens per second with Qwen3-30B-A3B, while the dense Qwen3-32B reaches 26.24.2424
The same table has no speed figure for the 4-bit GPTQ version of the MoE model. Its entry says the MoE kernel is unsupported.2424
On a laptop
Apple ran Qwen3-30B-A3B at 4-bit on a MacBook Pro with an M5 chip and 24 GB of unified memory, and the model took 17.31 GB.22 The M5 reads memory at 153 GB/s.22
At 4 bits a parameter takes half a byte, so the 3.3B active parameters come to about 1.65 GB per token.1919 Dividing one by the other gives a ceiling of about 93 tokens per second (153 / 1.65). A dense 32B model at 4-bit reads about 16 GB per token, for a ceiling near 10 tokens per second (153 / 16).2525Arithmetic, not a measurement: 153 GB/s / 1.65 GB = 93 tokens per second; 32B x 0.5 byte = 16 GB; 153 / 16 = 9.6 tokens per second. KV cache, quantization scales and router overhead are ignored. Inputs: https://machinelearning.apple.com/research/exploring-llms-mlx-m5 and https://huggingface.co/Qwen/Qwen3-30B-A3B
These ceilings are arithmetic, not measurements. Real speeds are lower, because the KV cache also has to be read, along with other overheads. Apple published only relative speedups and memory footprints, with no absolute speeds.2525
Apple did compare the M5 with the M4, which reads memory at 120 GB/s.22 The M5 has 1.275 times the bandwidth (153 / 120), and it generated tokens with this model 1.25 times faster.22
Keeping every expert busy
Left alone, a router can send most tokens to a few experts while the rest get little. In the DeepSeek-V3 report, an unbalanced expert load leads to "routing collapse".99
Uneven load is also a hardware problem. In a large model the experts sit on different GPUs, and the Mixtral paper stresses that "it is essential to distribute the workload evenly across the GPUs".44
A hard limit
Switch Transformer put a hard limit on each expert. Every expert gets a capacity, computed as (tokens per batch / number of experts) x capacity factor.77 A factor of 1.0 gives each expert exactly an even share of the batch. Tokens beyond the limit skip the expert and are "passed directly to the next layer through the residual connection".77
Switch tried factors from 1.0 to 2.0, and values from 1.0 to 1.25 worked better at scale.77
A penalty
Switch also adds a load-balancing loss. It uses two numbers per expert: the fraction of tokens the expert receives and the average probability the router gives it. The loss multiplies them for each expert and sums the products. The sum is smallest when every expert gets an equal share, and the term enters training with a weight of 0.01.77
ST-MoE later added a router z-loss, with a weight of 0.001.2626Zoph et al. (2022), ST-MoE, arXiv 2202.08906. https://arxiv.org/abs/2202.08906
A balancing loss also pulls the model away from its main task, and the DeepSeek-V3 report warns that "Too large an auxiliary loss will impair the model performance".99
A bias
DeepSeek-V3 balances mainly with a bias for each expert instead. The bias is added to the expert's score only when the router picks the top k; the weights that mix the outputs ignore it. When an expert is overloaded its bias goes down by 0.001, and when it is underloaded the bias goes up by the same step.99
That step size applied for the first 14.3T training tokens. After that it was set to zero, which stops the biases changing, but they stay in use. A small sequence-level balancing loss with a weight of 0.0001 runs alongside the biases as a complement.99 V3 drops no tokens, in training or in inference.99
An earlier paper by DeepSeek authors tested the bias method on a 1B model. Their measure of imbalance is the load of the busiest expert minus the mean load, divided by the mean.2727Wang et al. (2024), auxiliary-loss-free load balancing, arXiv 2408.15664. The imbalance measure is MaxVio_global. https://arxiv.org/abs/2408.15664
With a balancing loss it came to 0.72, so the busiest expert carried 72% more tokens than the average one. With the bias it came to 0.04.2727 Perplexity improved slightly, from 9.56 to 9.50.2727 A 3B model went from 0.52 to 0.04, with perplexity from 7.97 to 7.92.2727
Smaller experts, more of them
DeepSeekMoE, published in January 2024, changed the size of the experts. It split each expert into m smaller ones by cutting the feed-forward hidden size to 1/m, then activated m times as many, so compute per token stays the same.88 At 16.4B total and about 2.8B active, DeepSeekMoE 16B performed comparably to LLaMA2 7B with about 40% of its compute.88
Smaller experts give the router far more combinations to choose from. Mixtral's router can form 28 different pairs from its 8 experts. DeepSeek-V3's router works with 256 routed experts and takes 8. Picking 8 of 256 allows about 410 trillion sets in principle.2828Arithmetic: 8 choose 2 = 28; 256 choose 8 = 409,663,695,276,000, before V3's group limit. Expert counts from https://arxiv.org/abs/2401.04088 and https://arxiv.org/abs/2412.19437 V3 trims that by limiting each token to 4 of its 8 expert groups.99
The same paper set some experts apart as shared experts.88
Per layer:
- Mixtral has 8 experts and uses 2.44
- DeepSeek-V3 has 256 routed experts plus 1 shared, and uses 8 of the routed ones.99
- Qwen3.5 has 512 plus 1 shared and uses 10.1717
- Kimi K3 has 896 plus 2 shared and uses 16.15
- DeepSeek-V4-Pro moved the other way on k, with 384 routed experts plus 1 shared and 6 active.1818
Kimi K2's technical report tested the trend. It defines sparsity as total experts divided by active experts. With the active count held fixed, more experts gave lower loss.1616
A model at sparsity 8 needed 1.69 times as much compute to reach the same loss as one at sparsity 48, which has 384 experts with 8 active.1616
Qwen3 dropped shared experts in April 2025, and Qwen3.5 has one again.22221717 Mixtral and gpt-oss do not use them.4455
What the experts actually learn
In an MoE model, "expert" is the name of a slot the router can choose. No one assigns the slots a subject, and what each one ends up handling is found by measuring.
Mixtral's authors checked how the model routes text from ArXiv, PubMed, PhilPapers, GitHub and other sources. They "do not observe obvious patterns in the assignment of experts" by topic, and only DM Mathematics showed a marginally different distribution.44
Some of it follows syntax. The token "self" in Python and the word "Question" in English often go to the same expert, and "indentation tokens are always assigned to the same experts".44 Consecutive tokens also tend to reuse an expert: at layer 15 the first choice repeats in 24 to 28% of cases, against 12.5% for random routing.44
In 2017 Shazeer's team saw its experts become "highly specialized based on syntax and semantics".13 In ST-MoE, encoder experts split by token type, such as punctuation, verbs, proper names and numbers. In multilingual training they "do not exhibit language specialization".2626
The authors of OLMoE, an open model from AI2, found strong domain experts in it. On arXiv text, the first expert in its layer 0 is "nearly 100% specialized".2929AI2 (2024), OLMoE-1B-7B, arXiv 2409.02060. https://arxiv.org/abs/2409.02060 The same paper measured Mixtral's experts as "activated close to the uniform routing baseline".2929
Other teams tie specialization to balancing. Qwen's team found that balancing within each sequence suppresses it, while balancing over the global batch lets domain experts appear.3030Qwen (2025), global-batch load balancing, arXiv 2501.11873. https://arxiv.org/abs/2501.11873 The DeepSeek-V3 report finds "greater expert specialization patterns" with the bias method than with an auxiliary loss.99
The price
Memory
Every expert has to stay in memory, even though a token uses only a few. DeepSeek-V3 ships its weights in FP8, one byte per parameter, so the main weights alone take about 671 GB (671B x 1 byte), before the KV cache and activations.3131DeepSeek-V3 model card: 671B main model weights, "we only provide FP8 weights". Arithmetic: 671B x 1 byte = 671 GB. https://huggingface.co/deepseek-ai/DeepSeek-V3
Traffic
A model that size is spread across many GPUs, each holding some of the experts. Tokens travel to the GPU that holds their expert and the results travel back, through all-to-all exchanges at every MoE layer.6699
Shazeer's paper had already warned that "Network bandwidth can be a bottleneck".13 DeepSeek-V3 limits each token to experts on at most 4 nodes. It also reserves 20 streaming multiprocessors for the exchange, and overlaps the exchange with computation.99
Batches
Large batches are another requirement. When each token picks a few experts out of many, each expert receives only a small part of every batch.13 An expert that gets few tokens still has to read all of its weights from memory to process them, so the GPU spends more time reading than computing.
Mixtral's authors call MoE layers "more suitable for batched workloads", where the GPU can reach a good "arithmetic intensity".44 DeepSeek's notes on its inference system add that the model's "high sparsity necessitates an extremely large overall batch size".3232DeepSeek, open-infra-index, day 6 (inference system overview, February 2025). https://github.com/deepseek-ai/open-infra-index
The V3 report describes decoding units of 40 nodes and 320 GPUs. "Each GPU hosts only one expert", and 64 GPUs hold redundant copies of high-load experts along with the shared experts.99 The other 256 GPUs match the 256 routed experts one to one (256 + 64 = 320).
Closed models
Which closed frontier models use MoE is only partly public. Google says so for Gemini. The Gemini 2.5 report calls its models "sparse mixture-of-experts (MoE) transformers", and the Gemini 3 Pro card says the same.3333Google (2025), Gemini 2.5 technical report, arXiv 2507.06261. https://arxiv.org/abs/2507.06261 2323 OpenAI's open-weight gpt-oss models are MoE.55
The GPT-4 technical report gives "no further details about the architecture (including model size)".3434OpenAI (2023), GPT-4 technical report, arXiv 2303.08774. https://arxiv.org/abs/2303.08774 The GPT-5 system card describes a "real-time router that quickly decides which model to use", which is a choice between whole models.3535OpenAI, GPT-5 system card, 13 August 2025. https://cdn.openai.com/gpt-5-system-card.pdf
Notes
-
Kaplan et al. (2020), arXiv 2001.08361: forward-pass cost about 2N plus an attention term, 2 n_layer n_ctx d_model, that grows with context; training about 6N per token, times batch and steps for a full run; N is the non-embedding parameter count. https://arxiv.org/abs/2001.08361. Worked examples are arithmetic on that rule: 2 x 70B = 140B; 2 x 37B = 74B for DeepSeek-V3, attention excluded. 2 3 4
-
Apple Machine Learning Research, post on LLMs with MLX on the M5, 19 November 2025 (MacBook Pro M5, 24 GB; Qwen3-30B-A3B-MLX-4bit footprint 17.31 GB; 153 GB/s on M5, 120 GB/s on M4; M5 vs M4 generation speedup 1.25x). https://machinelearning.apple.com/research/exploring-llms-mlx-m5. Arithmetic: 153 / 120 = 1.275. 2 3 4 5 6
-
DeepSeek, Insights paper on DeepSeek-V3, arXiv 2505.09343, Table 2 (training cost per token: DeepSeek-V3 250, LLaMA-405B 2448). https://arxiv.org/abs/2505.09343. Arithmetic: 6 x 405B = 2,430 and 2,448 / 2,430 = 1.007; 250 / 2,448 = 0.10; 671 / 405 = 1.66. 2
-
Mistral AI, Mixtral paper, arXiv 2401.04088. https://arxiv.org/abs/2401.04088 2 3 4 5 6 7 8 9 10 11 12 13 14
-
OpenAI, gpt-oss model card, arXiv 2508.10925, and the official Hugging Face configs. https://arxiv.org/abs/2508.10925 2 3 4 5 6 7 8 9
-
Lepikhin et al. (2020), GShard, arXiv 2006.16668. https://arxiv.org/abs/2006.16668 2 3 4 5
-
Fedus, Zoph, Shazeer (2021), Switch Transformer, arXiv 2101.03961. https://arxiv.org/abs/2101.03961 2 3 4 5 6 7 8 9 10 11
-
Dai et al. (2024), DeepSeekMoE, arXiv 2401.06066. https://arxiv.org/abs/2401.06066 2 3 4
-
DeepSeek (2024), DeepSeek-V3 technical report, arXiv 2412.19437, and its config.json (n_routed_experts 256, num_experts_per_tok 8, n_group 8, topk_group 4). https://arxiv.org/abs/2412.19437 2 3 4 5 6 7 8 9 10 11 12 13 14 15
-
Mistral AI, Mixtral 8x7B announcement. https://mistral.ai/news/mixtral-of-experts 2
-
Arithmetic: 8 x 7B = 56B, against 46.7B reported; 2 / 8 x 46.7B = 11.7B, against 12.9B active. Only the feed-forward blocks are replaced by experts (arXiv 2401.04088), so attention and embeddings are counted once. https://arxiv.org/abs/2401.04088 2
-
Jacobs, Jordan, Nowlan, Hinton (1991), "Adaptive Mixtures of Local Experts", Neural Computation 3, 79-87. https://www.cs.toronto.edu/~hinton/absps/jjnh91.pdf 2 3 4 5
-
Shazeer et al. (2017), "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer", arXiv 1701.06538. https://arxiv.org/abs/1701.06538 2 3 4 5 6 7 8
-
Du et al. (2021), GLaM, arXiv 2112.06905. https://arxiv.org/abs/2112.06905 2
-
Moonshot AI, Kimi K3 announcement and model card. https://kimi.ai/blog/kimi-k3 2 3
-
Moonshot AI, Kimi K2 technical report, arXiv 2507.20534. The report gives 1.04T total; the model card rounds to 1T. https://arxiv.org/abs/2507.20534 2 3 4
-
Alibaba Cloud, Qwen3.5 announcement, and the Qwen3.5-397B-A17B model card. https://www.alibabacloud.com/blog/602894 2 3
-
DeepSeek, DeepSeek-V4 preview release note, 24 April 2026, with the DeepSeek-V4-Pro model card and config. https://api-docs.deepseek.com/news/news260424 2
-
Qwen3-30B-A3B model card. https://huggingface.co/Qwen/Qwen3-30B-A3B. Arithmetic: 3.3B x 0.5 byte = 1.65 GB. 2
-
Arithmetic on the reported sizes: Mixtral 12.9 / 46.7 = 27.6%; DeepSeek-V3 37 / 671 = 5.5%; Kimi K2 32 / 1,040 = 3.1%; DeepSeek-V4-Pro 49 / 1,600 = 3.1%; Kimi K3 104 / 2,800 = 3.7%. Sources: https://mistral.ai/news/mixtral-of-experts, https://arxiv.org/abs/2412.19437, https://arxiv.org/abs/2507.20534, https://api-docs.deepseek.com/news/news260424, https://kimi.ai/blog/kimi-k3 2
-
xAI, Grok-1 open release, 17 March 2024. https://x.ai/news/grok-os
-
Qwen (2025), Qwen3 technical report, arXiv 2505.09388, for the MoE results and the removal of shared experts. https://arxiv.org/abs/2505.09388. The April 2025 release date comes from the Qwen3 model cards on Hugging Face, for example https://huggingface.co/Qwen/Qwen3-30B-A3B 2
-
Google DeepMind, Gemini 3 Pro model card. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf 2
-
Qwen, official speed benchmark, NVIDIA H20 96GB, batch size 1, input length 1, 2,048 generated tokens. Speed is (prompt + generated tokens) / time. Figures by backend, Qwen3-30B-A3B against Qwen3-32B: SGLang BF16, 137.18 against 20.72 tokens per second; SGLang FP8, 155.55 against 46.17; plain Hugging Face Transformers BF16, 1.89 against 26.24, with the GPTQ-INT4 MoE entry marked "MoE Kernel Unsupported"; GPU memory in Transformers BF16, 58,462 MB against 62,751 MB. https://qwen.readthedocs.io/en/latest/getting_started/speed_benchmark.html 2 3 4 5
-
Arithmetic, not a measurement: 153 GB/s / 1.65 GB = 93 tokens per second; 32B x 0.5 byte = 16 GB; 153 / 16 = 9.6 tokens per second. KV cache, quantization scales and router overhead are ignored. Inputs: https://machinelearning.apple.com/research/exploring-llms-mlx-m5 and https://huggingface.co/Qwen/Qwen3-30B-A3B 2
-
Zoph et al. (2022), ST-MoE, arXiv 2202.08906. https://arxiv.org/abs/2202.08906 2
-
Wang et al. (2024), auxiliary-loss-free load balancing, arXiv 2408.15664. The imbalance measure is MaxVio_global. https://arxiv.org/abs/2408.15664 2 3 4
-
Arithmetic: 8 choose 2 = 28; 256 choose 8 = 409,663,695,276,000, before V3's group limit. Expert counts from https://arxiv.org/abs/2401.04088 and https://arxiv.org/abs/2412.19437
-
AI2 (2024), OLMoE-1B-7B, arXiv 2409.02060. https://arxiv.org/abs/2409.02060 2
-
Qwen (2025), global-batch load balancing, arXiv 2501.11873. https://arxiv.org/abs/2501.11873
-
DeepSeek-V3 model card: 671B main model weights, "we only provide FP8 weights". Arithmetic: 671B x 1 byte = 671 GB. https://huggingface.co/deepseek-ai/DeepSeek-V3
-
DeepSeek, open-infra-index, day 6 (inference system overview, February 2025). https://github.com/deepseek-ai/open-infra-index
-
Google (2025), Gemini 2.5 technical report, arXiv 2507.06261. https://arxiv.org/abs/2507.06261
-
OpenAI (2023), GPT-4 technical report, arXiv 2303.08774. https://arxiv.org/abs/2303.08774
-
OpenAI, GPT-5 system card, 13 August 2025. https://cdn.openai.com/gpt-5-system-card.pdf