Skip the cover

Google DeepMind

Gemini 3.1 Pro

Google's Pro model of the Gemini 3.1 generation, served in the Gemini API as a preview since 19 February 2026.

Total
Undisclosed
Active
Undisclosed
Experts
Mixture of expertsExpert counts not disclosed
Context
1M tokens
Max output
64k tokens
Input
text, image, video, audio, PDF
Output
text
Licence
Proprietary
Weights
API only

Not disclosed by Google DeepMind

Contents1 to 3

What is disclosed

Gemini 3.1 Pro went into the Gemini API on 19 February 2026, the day its model card was published.11Google DeepMind, Gemini 3.1 Pro model card, 19 February 2026. Architecture: "Gemini 3.1 Pro is based on Gemini 3 Pro." Inputs up to 1M tokens, output 64K. Benchmark table, columns "Gemini 3.1 Pro Thinking (High)" and "Gemini 3 Pro Thinking (High)": ARC-AGI-2, ARC Prize Verified, 77.1% and 31.1%; Humanity's Last Exam, full set, no tools, 44.4% and 37.5%; SWE-Bench Verified, single attempt, 80.6% and 76.2%; MRCR v2 (8-needle), 1M pointwise, 26.3% and 26.3%. https://deepmind.google/models/model-cards/gemini-3-1-pro/ 22Google, Gemini API release notes, entries of 19 February 2026 (Gemini 3.1 Pro Preview released) and 9 March 2026 (Gemini 3 Pro Preview shut down; gemini-3-pro-preview now points to gemini-3.1-pro-preview). https://ai.google.dev/gemini-api/docs/changelog

On the design, the card says only that Gemini 3.1 Pro "is based on Gemini 3 Pro", and leaves the architecture to the Gemini 3 Pro card.11

Sparse mixture of experts
Each token runs through "a subset of model parameters", picked by a learned router. The rest of the weights do no work for that token.

That card calls Gemini 3 Pro a sparse mixture of experts transformer.33Google DeepMind, Gemini 3 Pro model card: "a sparse mixture-of-experts (MoE)" transformer that activates "a subset of model parameters per input token". No parameter, expert or layer counts. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf It prints no parameter count and no expert count. The layers are not counted either.

Sparse mixture of experts
Each token runs through "a subset of model parameters", picked by a learned router. The rest of the weights do no work for that token.

A request can carry 1,048,576 input tokens, and the reply stops at 65,536.44Google, Gemini API model page for gemini-3.1-pro-preview, read 29 September 2026: input limit 1,048,576 tokens, output limit 65,536, status preview, function calling and search grounding supported, the customtools variant "optimized for agentic workflows that use custom tools and bash", text output only. https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview The card rounds the two limits to 1M and 64K.11

What is new

The card's comparison is with Gemini 3 Pro, both run at the Thinking (High) setting.11

The biggest jump is on ARC-AGI-2, a set of abstract reasoning puzzles. In the ARC Prize's verified run the score went from 31.1% to 77.1%, about two and a half times as high.1155Arithmetic on the card's ARC-AGI-2 scores: 77.1 / 31.1 = 2.48. https://deepmind.google/models/model-cards/gemini-3-1-pro/

Humanity's Last Exam, without tools, went from 37.5% to 44.4%.11 The SWE-Bench Verified score, with a single attempt per task, is 80.6% against 76.2%.11

MRCR v2 with eight needles, read at the full 1M tokens, gives 26.3% for both models.11

Gemini 3 Pro is gone from the API. Its preview shut down on 9 March 2026, and the old id now leads to 3.1 Pro.22

Using it

The model is served as a preview, under the id gemini-3.1-pro-preview.44 A second id, gemini-3.1-pro-preview-customtools, is tuned for agents that call their own tools and bash.44

Function calling and search grounding are supported. Output is text only.44

Thinking tokens
The tokens a model writes while it reasons, before its answer.

Input costs $2.00 per million tokens and output $12.00, for prompts up to 200,000 tokens.66Google, Gemini API pricing, paid tier, standard rates per 1M tokens, page last updated 24 September 2026, read on 29 September 2026. https://ai.google.dev/gemini-api/docs/pricing Past that length the rates rise to $4.00 and $18.00.66 Thinking tokens are billed as output.66

Thinking tokens
The tokens a model writes while it reasons, before its answer.

Context caching costs $0.20 per million tokens on the shorter prompts and $0.40 on the longer ones.66 Keeping the cache stored adds $4.50 per million tokens for every hour.66

Notes

  1. Google DeepMind, Gemini 3.1 Pro model card, 19 February 2026. Architecture: "Gemini 3.1 Pro is based on Gemini 3 Pro." Inputs up to 1M tokens, output 64K. Benchmark table, columns "Gemini 3.1 Pro Thinking (High)" and "Gemini 3 Pro Thinking (High)": ARC-AGI-2, ARC Prize Verified, 77.1% and 31.1%; Humanity's Last Exam, full set, no tools, 44.4% and 37.5%; SWE-Bench Verified, single attempt, 80.6% and 76.2%; MRCR v2 (8-needle), 1M pointwise, 26.3% and 26.3%. https://deepmind.google/models/model-cards/gemini-3-1-pro/ 2 3 4 5 6 7 8

  2. Google, Gemini API release notes, entries of 19 February 2026 (Gemini 3.1 Pro Preview released) and 9 March 2026 (Gemini 3 Pro Preview shut down; gemini-3-pro-preview now points to gemini-3.1-pro-preview). https://ai.google.dev/gemini-api/docs/changelog 2

  3. Google DeepMind, Gemini 3 Pro model card: "a sparse mixture-of-experts (MoE)" transformer that activates "a subset of model parameters per input token". No parameter, expert or layer counts. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf

  4. Google, Gemini API model page for gemini-3.1-pro-preview, read 29 September 2026: input limit 1,048,576 tokens, output limit 65,536, status preview, function calling and search grounding supported, the customtools variant "optimized for agentic workflows that use custom tools and bash", text output only. https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview 2 3 4

  5. Arithmetic on the card's ARC-AGI-2 scores: 77.1 / 31.1 = 2.48. https://deepmind.google/models/model-cards/gemini-3-1-pro/

  6. Google, Gemini API pricing, paid tier, standard rates per 1M tokens, page last updated 24 September 2026, read on 29 September 2026. https://ai.google.dev/gemini-api/docs/pricing 2 3 4 5

More from Google DeepMind

All models by Google DeepMind