gpt-oss-120b
The larger of OpenAI's two open-weight models, a mixture of experts with 116.8B parameters and 5.1B active per token, sized for one 80GB GPU.
- Totale
- 116,8B
- Attivi
- 5,1B
- Esperti
- 128 instradati4 per token
- Strati
- 36
- Attenzione
- Alternating 128-token window and full
- Contesto
- 131k token
- Uscita massima
- 131k token
- Ingresso
- text
- Uscita
- text
- Licenza
- Apache 2.0
- Pesi
- Hugging Face
Indiceda 1 a 3
How it is built
gpt-oss-120b has 36 layers and 116.8B parameters. Each token uses 5.1B of them, about 4.4%.11OpenAI, gpt-oss-120b and gpt-oss-20b model card, arXiv 2508.10925, August 2025, sections 2.1 to 2.6 and Tables 1 and 3. https://arxiv.org/abs/2508.10925. Arithmetic: 5.1 / 116.8 = 4.4%; 114.71 / 116.83 = 98%.
It is a mixture of experts transformer. Every layer holds 128 experts and no shared one. A linear router picks 4 for each token, and a softmax over those 4 alone sets their weights.11
The experts hold 114.71B parameters, 98% of the model.11 A token reaches 4 of the 128, about 3.58B. Attention adds 0.96B and the unembedding about 0.58B, which comes to about 5.12B against the 5.13B reported.22OpenAI, gpt-oss-120b config.json on Hugging Face (num_local_experts 128, num_experts_per_tok 4, hidden_size 2880, vocab_size 201088, quantization modules_to_not_convert). https://huggingface.co/openai/gpt-oss-120b/blob/main/config.json. Arithmetic on Table 1 of the model card, which counts the unembedding as active and the embedding as not: experts 114.71B x 4 / 128 = 3.58B; unembedding 201,088 x 2,880 = 0.58B; 3.58 + 0.96 + 0.58 = 5.12B.
From one layer to the next, attention alternates between a banded window of 128 tokens and full attention.11 Each layer has 64 query heads sharing 8 key and value heads.11 A learned bias in each head's softmax lets the head attend to no token at all.11
The residual stream is 2,880 wide. The tokenizer, o200k_harmony, has 201,088 tokens, and YaRN stretches the dense layers to a context of 131,072.11
What is new
By the model card's account, gpt-oss builds on the GPT-2 and GPT-3 architectures.11 Post-training used chain-of-thought reinforcement learning, with techniques similar to those of o3.11
It was trained on harmony, a chat format whose channels keep the chain of thought apart from tool calls and the final answer.11 Effort has three levels, picked by a line in the system prompt such as "Reasoning: low".11
At high effort OpenAI reports 62.4% on SWE-bench Verified and 97.9% on AIME 2025 with tools.11 GPQA Diamond without tools comes in at 80.1%, and the Codeforces Elo with tools at 2622.11 OpenAI's own summary is that the model surpasses o3-mini and approaches o4-mini.11
Training took 2.1 million H100 hours, and the knowledge cutoff is June 2024.11 In October 2025 OpenAI released gpt-oss-safeguard-120b, a safety reasoning model built on gpt-oss.33OpenAI, API changelog, entry of 29 October 2025. https://developers.openai.com/api/docs/changelog
Running it
The expert weights were post-trained in MXFP4, at 4.25 bits per parameter.11 The config keeps attention and the router out of MXFP4, along with the embedding and the output layer.22
In that format the model fits on a single 80GB GPU, such as an NVIDIA H100.44OpenAI, gpt-oss-120b model card on Hugging Face. https://huggingface.co/openai/gpt-oss-120b The checkpoint is 60.8 GiB.11 The weights are on Hugging Face under Apache 2.0, and the card warns that without the harmony format the model will not work correctly.44
The official repository gives commands for vLLM and Ollama.55OpenAI, gpt-oss repository, README, section on the reference PyTorch implementation. https://github.com/openai/gpt-oss OpenAI calls its reference PyTorch implementation inefficient. It upcasts every weight to BF16, and the larger model needs something like four H100s or two H200s to run it.55
OpenAI's model page lists 131,072 tokens for the context window and the same for output.66OpenAI, gpt-oss-120b model page, read 29 September 2026. https://developers.openai.com/api/docs/models/gpt-oss-120b gpt-oss-20b is the same design with 24 layers and 32 experts in each.11
Note
-
OpenAI, gpt-oss-120b and gpt-oss-20b model card, arXiv 2508.10925, August 2025, sections 2.1 to 2.6 and Tables 1 and 3. https://arxiv.org/abs/2508.10925. Arithmetic: 5.1 / 116.8 = 4.4%; 114.71 / 116.83 = 98%. 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18
-
OpenAI, gpt-oss-120b config.json on Hugging Face (num_local_experts 128, num_experts_per_tok 4, hidden_size 2880, vocab_size 201088, quantization modules_to_not_convert). https://huggingface.co/openai/gpt-oss-120b/blob/main/config.json. Arithmetic on Table 1 of the model card, which counts the unembedding as active and the embedding as not: experts 114.71B x 4 / 128 = 3.58B; unembedding 201,088 x 2,880 = 0.58B; 3.58 + 0.96 + 0.58 = 5.12B. 2
-
OpenAI, API changelog, entry of 29 October 2025. https://developers.openai.com/api/docs/changelog
-
OpenAI, gpt-oss-120b model card on Hugging Face. https://huggingface.co/openai/gpt-oss-120b 2
-
OpenAI, gpt-oss repository, README, section on the reference PyTorch implementation. https://github.com/openai/gpt-oss 2
-
OpenAI, gpt-oss-120b model page, read 29 September 2026. https://developers.openai.com/api/docs/models/gpt-oss-120b
Altri modelli di OpenAI
- GPT-6 AstraTotaleNon dichiaratoAttiviNon dichiarato
- GPT-6.1 SolTotaleNon dichiaratoAttiviNon dichiarato
- GPT-6 LunaTotaleNon dichiaratoAttiviNon dichiarato
- gpt-oss-20bTotale20,9BAttivi3,6B