Salta la copertina

xAIIn inglese

Grok 2

xAI's 2024 mixture of experts, with 8 experts in each of 64 layers and 2 used per token, released as open weights in August 2025.

Totale
Non dichiarato
Attivi
Non dichiarato
Esperti
8 instradati2 per token
Strati
64
Attenzione
Full, 64 query and 8 KV heads
Contesto
131k token
Ingresso
text
Uscita
text
Licenza
xAI Community License

Non dichiarato da xAI

Indiceda 1 a 3

How it is built

Its Hugging Face card calls Grok 2 "a model trained and used at xAI in 2024".11xAI, Grok 2 model card on Hugging Face. https://huggingface.co/xai-org/grok-2 xAI announced Grok-2 on 13 August 2024.22xAI, Grok-2 Beta Release, 13 August 2024, read in the Internet Archive copy of 28 September 2026. https://x.ai/news/grok-2 and https://web.archive.org/web/20260928085448/https://x.ai/news/grok-2 The weights of Grok 2 followed on 22 August 2025.33Hugging Face, metadata of the xai-org/grok-2 repository, created 22 August 2025. https://huggingface.co/api/models/xai-org/grok-2

It is a mixture of experts. Each of its 64 layers holds 8 experts, and the router sends every token to 2 of them.44xAI, Grok 2 config.json: num_hidden_layers 64, num_local_experts 8, num_experts_per_tok 2, intermediate_size 32768, moe_intermediate_size 16384, residual_moe true, num_attention_heads 64, num_key_value_heads 8, hidden_size 8192, vocab_size 131072, max_position_embeddings 131072, original_max_position_embeddings 8192, scaling_factor 16. https://huggingface.co/xai-org/grok-2/blob/main/config.json. Arithmetic: 32,768 / 16,384 = 2; 8,192 x 16 = 131,072. The config also sets a dense feed-forward width of 32,768, twice the 16,384 of one expert, and turns on residual_moe.44

Grouped-query attention
Several query heads share one set of keys and values, which shrinks the memory kept for earlier tokens.

Every layer attends over the full context, with 64 query heads sharing 8 key and value heads.44 The hidden size is 8,192. The vocabulary holds 131,072 tokens.44

Grouped-query attention
Several query heads share one set of keys and values, which shrinks the memory kept for earlier tokens.

xAI does not state the parameter count. The checkpoint holds 539,032,697,512 bytes of bfloat16 weights at two bytes each, about 269.5B parameters.55Hugging Face file listing of xai-org/grok-2: 39 safetensors files with 539,032,697,512 bytes in all; the config gives torch_dtype bfloat16. https://huggingface.co/api/models/xai-org/grok-2/tree/main. Arithmetic: 539,032,697,512 / 2 = 269.52B. From the config, per layer: attention 0.151B, eight experts 8 x 3 x 8,192 x 16,384 = 3.221B, dense block 3 x 8,192 x 32,768 = 0.805B, so 4.178B; 64 x 4.178B = 267.37B, plus 2 x 131,072 x 8,192 = 2.15B for the embeddings and the output head, gives 269.51B. Counting from the config gives the same total if every layer carries that dense block beside its experts.55

Fig. 1
Token

64 layers. In each one, a token goes to 2 of 8 routed experts. The experts it meets here are simulated, because the real router depends on the trained weights.

What is new

Grok-1, whose weights came out in March 2024, had the same 64 layers and the same 8 experts with 2 per token.66xAI, Grok-1 repository README: 314B parameters, 8 experts with 2 per token, 64 layers, 48 query and 8 key and value heads, embedding size 6,144, context 8,192 tokens, Apache 2.0. https://github.com/xai-org/grok-1. The release date, 17 March 2024, is on the xAI company timeline, read in the Internet Archive copy of 21 September 2026: https://web.archive.org/web/20260921183023/https://x.ai/company Grok 2 is wider. The embedding grows from 6,144 to 8,192 and the query heads from 48 to 64.6644 The context goes from 8,192 tokens to 131,072, with the rotary position embeddings stretched by a factor of 16.6644

At its announcement xAI reported 87.5% on MMLU, zero-shot with chain of thought, and 56.0% on GPQA. Grok-1.5 had scored 81.3% and 35.9%.22 xAI does not say whether the released checkpoint is the one it tested.

The licence is new too. Grok-1 came under Apache 2.0.66 Grok 2 comes under the xAI Community License, which allows commercial use within xAI's Acceptable Use Policy.77xAI, xAI Community License Agreement for Grok 2, last updated 4 November 2025. https://huggingface.co/xai-org/grok-2/blob/main/LICENSE It forbids using the model or its outputs to train other foundation or general-purpose models, apart from fine-tuning xAI's own materials.77

Running it

The download is about 500 GB in 42 files. The card warns that it may fail, and says to retry until it succeeds.11 Serving needs SGLang 0.5.1 or later.11 The checkpoint is split for tensor parallelism across 8 GPUs, each with more than 40 GB of memory, and the launch command passes an FP8 quantization flag.11

Tensor parallelism
Splitting each layer's weight matrices across several GPUs that work on every token together.

At one byte per parameter the weights come to about 270 GB. Eight GPUs of 40 GB hold 320 GB, before the KV cache takes its share.88Arithmetic, not a measurement, on the card and the file listing: 269.5B parameters x 1 byte in FP8 = about 270 GB; 8 GPUs x 40 GB = 320 GB. Embeddings kept at higher precision, the KV cache and activations are ignored. https://huggingface.co/xai-org/grok-2 and https://huggingface.co/api/models/xai-org/grok-2/tree/main

Tensor parallelism
Splitting each layer's weight matrices across several GPUs that work on every token together.

The model is post-trained for chat and expects its own chat template.11 Anyone who distributes it, or a product built on it, has to display "Powered by xAI" and include the licence.77

Note

  1. xAI, Grok 2 model card on Hugging Face. https://huggingface.co/xai-org/grok-2 2 3 4 5

  2. xAI, Grok-2 Beta Release, 13 August 2024, read in the Internet Archive copy of 28 September 2026. https://x.ai/news/grok-2 and https://web.archive.org/web/20260928085448/https://x.ai/news/grok-2 2

  3. Hugging Face, metadata of the xai-org/grok-2 repository, created 22 August 2025. https://huggingface.co/api/models/xai-org/grok-2

  4. xAI, Grok 2 config.json: num_hidden_layers 64, num_local_experts 8, num_experts_per_tok 2, intermediate_size 32768, moe_intermediate_size 16384, residual_moe true, num_attention_heads 64, num_key_value_heads 8, hidden_size 8192, vocab_size 131072, max_position_embeddings 131072, original_max_position_embeddings 8192, scaling_factor 16. https://huggingface.co/xai-org/grok-2/blob/main/config.json. Arithmetic: 32,768 / 16,384 = 2; 8,192 x 16 = 131,072. 2 3 4 5 6

  5. Hugging Face file listing of xai-org/grok-2: 39 safetensors files with 539,032,697,512 bytes in all; the config gives torch_dtype bfloat16. https://huggingface.co/api/models/xai-org/grok-2/tree/main. Arithmetic: 539,032,697,512 / 2 = 269.52B. From the config, per layer: attention 0.151B, eight experts 8 x 3 x 8,192 x 16,384 = 3.221B, dense block 3 x 8,192 x 32,768 = 0.805B, so 4.178B; 64 x 4.178B = 267.37B, plus 2 x 131,072 x 8,192 = 2.15B for the embeddings and the output head, gives 269.51B. 2

  6. xAI, Grok-1 repository README: 314B parameters, 8 experts with 2 per token, 64 layers, 48 query and 8 key and value heads, embedding size 6,144, context 8,192 tokens, Apache 2.0. https://github.com/xai-org/grok-1. The release date, 17 March 2024, is on the xAI company timeline, read in the Internet Archive copy of 21 September 2026: https://web.archive.org/web/20260921183023/https://x.ai/company 2 3 4

  7. xAI, xAI Community License Agreement for Grok 2, last updated 4 November 2025. https://huggingface.co/xai-org/grok-2/blob/main/LICENSE 2 3

  8. Arithmetic, not a measurement, on the card and the file listing: 269.5B parameters x 1 byte in FP8 = about 270 GB; 8 GPUs x 40 GB = 320 GB. Embeddings kept at higher precision, the KV cache and activations are ignored. https://huggingface.co/xai-org/grok-2 and https://huggingface.co/api/models/xai-org/grok-2/tree/main

Altri modelli di xAI

  • Grok 4.7TotaleNon dichiaratoAttiviNon dichiarato
  • Grok Build 0.1TotaleNon dichiaratoAttiviNon dichiarato
  • Grok 4.3TotaleNon dichiaratoAttiviNon dichiarato
Tutti i modelli di xAI