Skip the cover

OpenAI

gpt-oss-20b

The smaller of OpenAI's two open-weight models, a mixture of experts with 20.9B parameters and 3.6B active per token, built to run in 16GB of memory.

Total
20.9B
Active
3.6B
Experts
32 routed4 per token
Layers
24
Attention
Alternating 128-token window and full
Context
131k tokens
Max output
131k tokens
Input
text
Output
text
Licence
Apache 2.0

Contents1 to 3

How it is built

The smaller gpt-oss is a mixture of experts transformer with 24 layers.11OpenAI, gpt-oss-120b and gpt-oss-20b model card, arXiv 2508.10925, August 2025, sections 2.1 to 2.6 and Tables 1 and 3. https://arxiv.org/abs/2508.10925. Arithmetic: 3.6 / 20.9 = 17%; 4 / 32 = 1/8. Of its 20.9B parameters, 3.6B work on each token, about 17%.11

Each layer has 32 experts and no shared expert. The router sends every token to 4 of them, an eighth of the layer.11

Fig. 1
Token

24 layers. In each one, a token goes to 4 of 32 routed experts. The experts it meets here are simulated, because the real router depends on the trained weights.

The experts hold 19.12B parameters, and a token reaches an eighth of them, about 2.39B.22Arithmetic on Table 1 of the model card, which counts the unembedding as active and the embedding as not: experts 19.12B / 8 = 2.39B; unembedding 201,088 x 2,880 = 0.58B, from vocab_size and hidden_size in the config; 2.39 + 0.64 + 0.58 = 3.61B. Embedding and unembedding together are 1.16B in both models, 5.5% of gpt-oss-20b (1.16 / 20.91) and 1.0% of gpt-oss-120b (1.16 / 116.83). https://arxiv.org/abs/2508.10925 and https://huggingface.co/openai/gpt-oss-20b/blob/main/config.json Attention adds 0.64B and the unembedding about 0.58B, for 3.61B in all.22

Unembedding
The last layer, which turns the final hidden vector into one score for every token in the vocabulary.

The embedding and unembedding did not shrink with the model. Together they hold 1.16B parameters in both sizes, which is 5.5% of gpt-oss-20b and 1.0% of the larger model.22

Unembedding
The last layer, which turns the final hidden vector into one score for every token in the vocabulary.

Everything else matches gpt-oss-120b. Attention alternates between a banded window of 128 tokens and full attention, with 64 query heads and 8 key and value heads.11 The residual stream has 2,880 dimensions, and YaRN extends the context to 131,072 tokens.11

What is new

Harmony
OpenAI's chat format for gpt-oss. Special tokens mark each message, and channels keep the chain of thought apart from the answer.

It was released on the same day as the larger model, under the same Apache 2.0 licence.33OpenAI, gpt-oss-20b model card on Hugging Face. https://huggingface.co/openai/gpt-oss-20b Both use the harmony chat format, with one of three reasoning levels set in the system prompt.11

Harmony
OpenAI's chat format for gpt-oss. Special tokens mark each message, and channels keep the chain of thought apart from the answer.

At high effort OpenAI reports 60.7% on SWE-bench Verified, against 62.4% for gpt-oss-120b.11 On AIME 2025 with tools the smaller model comes out slightly ahead, 98.7% to 97.9%.11 It falls behind on GPQA Diamond without tools, at 71.5%, and the card puts the lag on knowledge tasks down to its smaller size.11

For each AIME problem it writes more than 20,000 chain-of-thought tokens on average.11 Its training took almost ten times fewer H100 hours than the 2.1 million of the larger model.11

A safety reasoning model built on it, gpt-oss-safeguard-20b, followed in October 2025.44OpenAI, API changelog, entry of 29 October 2025. https://developers.openai.com/api/docs/changelog

Running it

The expert weights are stored in MXFP4, and the checkpoint is 12.8 GiB.11 OpenAI says the model runs within 16GB of memory.33

Apple ran it with MLX on a MacBook Pro with an M5 chip and 24GB of unified memory. With a prompt of 4,096 tokens it used 12.08 GB, and it generated tokens 1.24 times as fast as on an M4.55Apple Machine Learning Research, post on LLMs with MLX on the M5, 19 November 2025 (MacBook Pro M5, 24 GB; prompt of 4,096 tokens; gpt-oss-20b in MXFP4 at 12.08 GB; generation speedup over the M4 1.24x). https://machinelearning.apple.com/research/exploring-llms-mlx-m5

The official repository has a reference Metal implementation for Apple silicon. OpenAI says it is not production-ready.66OpenAI, gpt-oss repository, README. https://github.com/openai/gpt-oss The README also gives commands for Ollama and vLLM.66

The weights must be used with the harmony format.33 OpenAI's model page gives the same limits as for the larger model, 131,072 tokens of context and of output.77OpenAI, gpt-oss-20b model page, read 29 September 2026. https://developers.openai.com/api/docs/models/gpt-oss-20b

Notes

  1. OpenAI, gpt-oss-120b and gpt-oss-20b model card, arXiv 2508.10925, August 2025, sections 2.1 to 2.6 and Tables 1 and 3. https://arxiv.org/abs/2508.10925. Arithmetic: 3.6 / 20.9 = 17%; 4 / 32 = 1/8. 2 3 4 5 6 7 8 9 10 11 12

  2. Arithmetic on Table 1 of the model card, which counts the unembedding as active and the embedding as not: experts 19.12B / 8 = 2.39B; unembedding 201,088 x 2,880 = 0.58B, from vocab_size and hidden_size in the config; 2.39 + 0.64 + 0.58 = 3.61B. Embedding and unembedding together are 1.16B in both models, 5.5% of gpt-oss-20b (1.16 / 20.91) and 1.0% of gpt-oss-120b (1.16 / 116.83). https://arxiv.org/abs/2508.10925 and https://huggingface.co/openai/gpt-oss-20b/blob/main/config.json 2 3

  3. OpenAI, gpt-oss-20b model card on Hugging Face. https://huggingface.co/openai/gpt-oss-20b 2 3

  4. OpenAI, API changelog, entry of 29 October 2025. https://developers.openai.com/api/docs/changelog

  5. Apple Machine Learning Research, post on LLMs with MLX on the M5, 19 November 2025 (MacBook Pro M5, 24 GB; prompt of 4,096 tokens; gpt-oss-20b in MXFP4 at 12.08 GB; generation speedup over the M4 1.24x). https://machinelearning.apple.com/research/exploring-llms-mlx-m5

  6. OpenAI, gpt-oss repository, README. https://github.com/openai/gpt-oss 2

  7. OpenAI, gpt-oss-20b model page, read 29 September 2026. https://developers.openai.com/api/docs/models/gpt-oss-20b

More from OpenAI

All models by OpenAI