Skip the cover

Google DeepMind

Gemini 3.8 Flash

Google's Flash model of September 2026, built on Gemini 3.7 Flash and meant for long-horizon software engineering and agents.

Total
Undisclosed
Active
Undisclosed
Experts
Undisclosed
Context
1M tokens
Max output
64k tokens
Input
text, image, video, audio, PDF
Output
text
Licence
Proprietary
Weights
API only

Not disclosed by Google DeepMind

Contents1 to 3

What is disclosed

Google's release notes of 2 September 2026 call Gemini 3.8 Flash its "most intelligent Flash model". It went into the Gemini API that day as a stable model.11Google, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber", 2 September 2026. Flash Cyber is "our most capable cybersecurity model", "available to trusted defenders through our new Fairwind Program"; 3.8 Flash comes "at the same introductory price" as 3.7 Flash. https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ 22Google, Gemini API model page for gemini-3.8-flash (input limit 1,048,576 tokens, output limit 65,536, stable, thinking levels low, medium and high, "minimal" not supported, computer use in preview, no image or audio generation), and the release notes entry of 2 September 2026 ("our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows"). https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash and https://ai.google.dev/gemini-api/docs/changelog

Its card says the model "is based on Gemini 3.7 Flash" and sends the reader there for the architecture.33Google DeepMind, Gemini 3.8 Flash model card, September 2026. Architecture: "Gemini 3.8 Flash is based on Gemini 3.7 Flash." Knowledge cutoff March 2026, with some domains limited to January 2025. DeepSWE v1.1: 73.7% against 65.3% for Gemini 3.7 Flash. Terminal-bench 2.1: 89.4% against 85.8%. Price rows: $0.75 input and $3.75 output per 1M tokens for both models ($1.50 and $7.50 regular). https://deepmind.google/models/model-cards/gemini-3-8-flash/ The 3.7 Flash card does the same with 3.6 Flash.44Google DeepMind, model cards for Gemini 3.7 Flash (13 August 2026), Gemini 3.6 Flash (21 July 2026) and Gemini 3.5 Flash (19 May 2026), each "based on" the one before it, and the Gemini 3 Flash card (December 2025): "Gemini 3 Flash is based on Gemini 3 Pro." Counting back from 3.8 Flash, Gemini 3 Pro is the fifth card. https://deepmind.google/models/model-cards/gemini-3-7-flash/, https://deepmind.google/models/model-cards/gemini-3-6-flash/, https://deepmind.google/models/model-cards/gemini-3-5-flash/, https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Flash-Model-Card.pdf The chain carries on through Gemini 3.5 Flash to Gemini 3 Flash, whose card names Gemini 3 Pro as its base.44

The only description of the design is five cards back. It is on the Gemini 3 Pro card, which describes a sparse mixture of experts.4455Google DeepMind, Gemini 3 Pro model card: "a sparse mixture-of-experts (MoE)" transformer, with no parameter or expert counts. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf None of the Flash cards in the chain says so about itself, and none gives a parameter count.

The context window is 1,048,576 tokens and the output limit 65,536.22 The card gives March 2026 as the knowledge cutoff. In some domains it is January 2025.33

What is new

Harness
The agent program around the model. It hands the model its tools and runs each step of the task.

On DeepSWE v1.1, a software engineering test, the card gives 73.7%, against 65.3% for Gemini 3.7 Flash. Google computed its own scores there with a mini-swe agent harness at high thinking.3366Google DeepMind, Gemini 3.8 Flash model evaluation, September 2026: Gemini scores are pass@1, run through the Gemini API for gemini-3.8-flash with default sampling; DeepSWE v1.1 self computed with "a mini-swe agent harness with high thinking"; Terminal-Bench 2.1 with the default agent harness, Terminus 2. https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_flash_model_evaluation.pdf

Harness
The agent program around the model. It hands the model its tools and runs each step of the task.
Pass@1
The share of tasks solved on the first attempt, with no voting across several tries.

Terminal-bench 2.1, run with the Terminus 2 harness, rose less, from 85.8% to 89.4%.3366 Every Gemini score in the table is pass@1 through the API, with default sampling.66

Pass@1
The share of tasks solved on the first attempt, with no voting across several tries.

Google kept the introductory price of 3.7 Flash.1133

Gemini 3.8 Flash Cyber launched the same day. Google calls it its most capable cybersecurity model and gives access only to "trusted defenders", through a new Fairwind Program.11

Using it

In the API the model is gemini-3.8-flash.22 Thinking comes in three levels, from low to high. The minimal level is not supported.22

Input costs $0.75 per million tokens and output $3.75 until 31 December 2026, with thinking tokens counted as output.77Google, Gemini API pricing, paid tier, standard rates per 1M tokens, page last updated 24 September 2026, read on 29 September 2026. https://ai.google.dev/gemini-api/docs/pricing. Arithmetic: 1.50 / 0.75 = 2; 7.50 / 3.75 = 2. On 1 January 2027 both rates double, to $1.50 and $7.50.77 Until then cached input costs $0.075 per million tokens.77

Computer use is listed as a preview feature. Image and audio generation are not supported.22

Notes

  1. Google, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber", 2 September 2026. Flash Cyber is "our most capable cybersecurity model", "available to trusted defenders through our new Fairwind Program"; 3.8 Flash comes "at the same introductory price" as 3.7 Flash. https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ 2 3

  2. Google, Gemini API model page for gemini-3.8-flash (input limit 1,048,576 tokens, output limit 65,536, stable, thinking levels low, medium and high, "minimal" not supported, computer use in preview, no image or audio generation), and the release notes entry of 2 September 2026 ("our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows"). https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash and https://ai.google.dev/gemini-api/docs/changelog 2 3 4 5

  3. Google DeepMind, Gemini 3.8 Flash model card, September 2026. Architecture: "Gemini 3.8 Flash is based on Gemini 3.7 Flash." Knowledge cutoff March 2026, with some domains limited to January 2025. DeepSWE v1.1: 73.7% against 65.3% for Gemini 3.7 Flash. Terminal-bench 2.1: 89.4% against 85.8%. Price rows: $0.75 input and $3.75 output per 1M tokens for both models ($1.50 and $7.50 regular). https://deepmind.google/models/model-cards/gemini-3-8-flash/ 2 3 4 5

  4. Google DeepMind, model cards for Gemini 3.7 Flash (13 August 2026), Gemini 3.6 Flash (21 July 2026) and Gemini 3.5 Flash (19 May 2026), each "based on" the one before it, and the Gemini 3 Flash card (December 2025): "Gemini 3 Flash is based on Gemini 3 Pro." Counting back from 3.8 Flash, Gemini 3 Pro is the fifth card. https://deepmind.google/models/model-cards/gemini-3-7-flash/, https://deepmind.google/models/model-cards/gemini-3-6-flash/, https://deepmind.google/models/model-cards/gemini-3-5-flash/, https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Flash-Model-Card.pdf 2 3

  5. Google DeepMind, Gemini 3 Pro model card: "a sparse mixture-of-experts (MoE)" transformer, with no parameter or expert counts. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf

  6. Google DeepMind, Gemini 3.8 Flash model evaluation, September 2026: Gemini scores are pass@1, run through the Gemini API for gemini-3.8-flash with default sampling; DeepSWE v1.1 self computed with "a mini-swe agent harness with high thinking"; Terminal-Bench 2.1 with the default agent harness, Terminus 2. https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_flash_model_evaluation.pdf 2 3

  7. Google, Gemini API pricing, paid tier, standard rates per 1M tokens, page last updated 24 September 2026, read on 29 September 2026. https://ai.google.dev/gemini-api/docs/pricing. Arithmetic: 1.50 / 0.75 = 2; 7.50 / 3.75 = 2. 2 3

More from Google DeepMind

All models by Google DeepMind