Skip the cover

OpenAI

gpt-oss-120b

The larger of OpenAI's two open-weight models, a mixture of experts with 116.8B parameters and 5.1B active per token, sized for one 80GB GPU.

Total
116.8B
Active
5.1B
Experts
128 routed4 per token
Layers
36
Attention
Alternating 128-token window and full
Context
131k tokens
Max output
131k tokens
Input
text
Output
text
Licence
Apache 2.0

Contents1 to 3

How it is built

gpt-oss-120b has 36 layers and 116.8B parameters. Each token uses 5.1B of them, about 4.4%.11OpenAI, gpt-oss-120b and gpt-oss-20b model card, arXiv 2508.10925, August 2025, sections 2.1 to 2.6 and Tables 1 and 3. https://arxiv.org/abs/2508.10925. Arithmetic: 5.1 / 116.8 = 4.4%; 114.71 / 116.83 = 98%.

It is a mixture of experts transformer. Every layer holds 128 experts and no shared one. A linear router picks 4 for each token, and a softmax over those 4 alone sets their weights.11

Fig. 1
Token

36 layers. In each one, a token goes to 4 of 128 routed experts. The experts it meets here are simulated, because the real router depends on the trained weights.

The experts hold 114.71B parameters, 98% of the model.11 A token reaches 4 of the 128, about 3.58B. Attention adds 0.96B and the unembedding about 0.58B, which comes to about 5.12B against the 5.13B reported.22OpenAI, gpt-oss-120b config.json on Hugging Face (num_local_experts 128, num_experts_per_tok 4, hidden_size 2880, vocab_size 201088, quantization modules_to_not_convert). https://huggingface.co/openai/gpt-oss-120b/blob/main/config.json. Arithmetic on Table 1 of the model card, which counts the unembedding as active and the embedding as not: experts 114.71B x 4 / 128 = 3.58B; unembedding 201,088 x 2,880 = 0.58B; 3.58 + 0.96 + 0.58 = 5.12B.

Banded window
Attention that looks back only over a fixed number of recent tokens.

From one layer to the next, attention alternates between a banded window of 128 tokens and full attention.11 Each layer has 64 query heads sharing 8 key and value heads.11 A learned bias in each head's softmax lets the head attend to no token at all.11

Banded window
Attention that looks back only over a fixed number of recent tokens.

The residual stream is 2,880 wide. The tokenizer, o200k_harmony, has 201,088 tokens, and YaRN stretches the dense layers to a context of 131,072.11

What is new

By the model card's account, gpt-oss builds on the GPT-2 and GPT-3 architectures.11 Post-training used chain-of-thought reinforcement learning, with techniques similar to those of o3.11

It was trained on harmony, a chat format whose channels keep the chain of thought apart from tool calls and the final answer.11 Effort has three levels, picked by a line in the system prompt such as "Reasoning: low".11

At high effort OpenAI reports 62.4% on SWE-bench Verified and 97.9% on AIME 2025 with tools.11 GPQA Diamond without tools comes in at 80.1%, and the Codeforces Elo with tools at 2622.11 OpenAI's own summary is that the model surpasses o3-mini and approaches o4-mini.11

Training took 2.1 million H100 hours, and the knowledge cutoff is June 2024.11 In October 2025 OpenAI released gpt-oss-safeguard-120b, a safety reasoning model built on gpt-oss.33OpenAI, API changelog, entry of 29 October 2025. https://developers.openai.com/api/docs/changelog

Running it

The expert weights were post-trained in MXFP4, at 4.25 bits per parameter.11 The config keeps attention and the router out of MXFP4, along with the embedding and the output layer.22

In that format the model fits on a single 80GB GPU, such as an NVIDIA H100.44OpenAI, gpt-oss-120b model card on Hugging Face. https://huggingface.co/openai/gpt-oss-120b The checkpoint is 60.8 GiB.11 The weights are on Hugging Face under Apache 2.0, and the card warns that without the harmony format the model will not work correctly.44

The official repository gives commands for vLLM and Ollama.55OpenAI, gpt-oss repository, README, section on the reference PyTorch implementation. https://github.com/openai/gpt-oss OpenAI calls its reference PyTorch implementation inefficient. It upcasts every weight to BF16, and the larger model needs something like four H100s or two H200s to run it.55

OpenAI's model page lists 131,072 tokens for the context window and the same for output.66OpenAI, gpt-oss-120b model page, read 29 September 2026. https://developers.openai.com/api/docs/models/gpt-oss-120b gpt-oss-20b is the same design with 24 layers and 32 experts in each.11

Notes

  1. OpenAI, gpt-oss-120b and gpt-oss-20b model card, arXiv 2508.10925, August 2025, sections 2.1 to 2.6 and Tables 1 and 3. https://arxiv.org/abs/2508.10925. Arithmetic: 5.1 / 116.8 = 4.4%; 114.71 / 116.83 = 98%. 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18

  2. OpenAI, gpt-oss-120b config.json on Hugging Face (num_local_experts 128, num_experts_per_tok 4, hidden_size 2880, vocab_size 201088, quantization modules_to_not_convert). https://huggingface.co/openai/gpt-oss-120b/blob/main/config.json. Arithmetic on Table 1 of the model card, which counts the unembedding as active and the embedding as not: experts 114.71B x 4 / 128 = 3.58B; unembedding 201,088 x 2,880 = 0.58B; 3.58 + 0.96 + 0.58 = 5.12B. 2

  3. OpenAI, API changelog, entry of 29 October 2025. https://developers.openai.com/api/docs/changelog

  4. OpenAI, gpt-oss-120b model card on Hugging Face. https://huggingface.co/openai/gpt-oss-120b 2

  5. OpenAI, gpt-oss repository, README, section on the reference PyTorch implementation. https://github.com/openai/gpt-oss 2

  6. OpenAI, gpt-oss-120b model page, read 29 September 2026. https://developers.openai.com/api/docs/models/gpt-oss-120b

More from OpenAI

All models by OpenAI