Salta la copertina

Google DeepMindIn inglese

Gemini 3.5 Flash-Lite

Google's Flash-Lite model for subagent tasks and document parsing at volume, priced above Gemini 3.1 Flash-Lite.

Totale
Non dichiarato
Attivi
Non dichiarato
Esperti
Non dichiarato
Contesto
1M token
Uscita massima
64k token
Ingresso
text, image, video, audio, PDF
Uscita
text
Licenza
Proprietary
Pesi
Solo API

Non dichiarato da Google DeepMind

Indiceda 1 a 3

What is disclosed

Google's card for Gemini 3.5 Flash-Lite, published on 21 July 2026, points one step back for the architecture. The model "is based on Gemini 3.1 Flash-Lite".11Google DeepMind, Gemini 3.5 Flash-Lite model card, 21 July 2026. Architecture: "Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite." Knowledge cutoff March 2026, with some domains limited to January 2025. Benchmark table against Gemini 3.1 Flash-Lite: SWE-Bench Pro (Public) 54.2% and 38.3%; Terminal-bench 2.1, Terminus-2 harness, 54.0% and 31.0%; OSWorld-Verified 74.0% and 54.3%; GDM-MRCR v2 (8-needle), 128k average, 72.2% and 60.1%; 1M pointwise, 21.3% and 12.3%. Price rows per 1M tokens: $0.30 and $2.50 against $0.25 and $1.50. https://deepmind.google/models/model-cards/gemini-3-5-flash-lite/ It went into the Gemini API as a stable release the same day.22Google, Gemini API model page for gemini-3.5-flash-lite (input limit 1,048,576 tokens, output limit 65,536, stable, computer use in preview, Live API not supported), and the release notes entry of 21 July 2026 ("a low-latency, highly cost-effective subagent option designed for high-volume automation"). https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite and https://ai.google.dev/gemini-api/docs/changelog

TPU
Tensor Processing Unit, the chip Google designs for machine learning.

The Gemini 3.1 Flash-Lite card in turn names Gemini 3 Pro as its base. It adds that the model was trained on Google's TPUs.33Google DeepMind, Gemini 3.1 Flash-Lite model card, 3 March 2026: "Gemini 3.1 Flash-Lite is based on Gemini 3 Pro"; "trained using Google's Tensor Processing Units (TPUs)". https://deepmind.google/models/model-cards/gemini-3-1-flash-lite/

TPU
Tensor Processing Unit, the chip Google designs for machine learning.

Only the Gemini 3 Pro card describes a sparse mixture of experts.44Google DeepMind, Gemini 3 Pro model card: "a sparse mixture-of-experts (MoE)" transformer, with no parameter or expert counts. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf Neither Flash-Lite card uses the term. There is no parameter count on either.

Knowledge runs to March 2026, though the card warns that some domains stop at January 2025.11 The context window is 1,048,576 tokens, and a reply can be up to 65,536 tokens long.22

What is new

Google set it against Gemini 3.1 Flash-Lite, scoring both through the API at high thinking, on the first attempt.1155Google DeepMind, Gemini 3.5 Flash-Lite model evaluation, July 2026: Gemini scores are pass@1, run through the Gemini API for gemini-3.5-flash-lite "with high thinking and default sampling settings"; SWE-Bench Pro with "an internal version of the Antigravity harness"; Terminal-Bench 2.1 with the default agent harness, Terminus 2; OSWorld-Verified "averaged over 5 runs with a single attempt per run". https://storage.googleapis.com/deepmind-media/gemini/gemini_3-5_flash-lite_model_evaluation.pdf

The biggest jump is on Terminal-bench 2.1, from 31.0% to 54.0% under the Terminus 2 harness.1155 SWE-Bench Pro went from 38.3% to 54.2% on the public set. That run used an internal version of Google's Antigravity harness.1155 OSWorld-Verified, a computer use test, was averaged over five runs and reached 74.0%, against 54.3%.1155

At 128k tokens, MRCR v2 with eight needles averages 72.2%, up from 60.1%.11 At the full 1M the score is 21.3%. The older model managed 12.3%.11

The price went up. The card lists Gemini 3.1 Flash-Lite at $0.25 per million input tokens and $1.50 per million output tokens.11

Using it

Subagent
A model that a larger agent calls to handle one part of a task.

The id is gemini-3.5-flash-lite.22 The API page describes it as "optimized for high-throughput, low-cost execution for subagent tasks and document parsing".22 In the release notes it is a subagent option "designed for high-volume automation".22

Subagent
A model that a larger agent calls to handle one part of a task.

Input costs $0.30 per million tokens, at one rate for every type of input.66Google, Gemini API pricing, paid tier, standard rates per 1M tokens, page last updated 24 September 2026, read on 29 September 2026. The input row reads "$0.30 (text / image / video / audio)". https://ai.google.dev/gemini-api/docs/pricing Output costs $2.50 per million, thinking tokens included.66

Cached input is $0.03 per million tokens, plus $1.00 per million tokens for each hour the cache stays stored.66

The Live API is not supported. Computer use is, as a preview.22

Note

  1. Google DeepMind, Gemini 3.5 Flash-Lite model card, 21 July 2026. Architecture: "Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite." Knowledge cutoff March 2026, with some domains limited to January 2025. Benchmark table against Gemini 3.1 Flash-Lite: SWE-Bench Pro (Public) 54.2% and 38.3%; Terminal-bench 2.1, Terminus-2 harness, 54.0% and 31.0%; OSWorld-Verified 74.0% and 54.3%; GDM-MRCR v2 (8-needle), 128k average, 72.2% and 60.1%; 1M pointwise, 21.3% and 12.3%. Price rows per 1M tokens: $0.30 and $2.50 against $0.25 and $1.50. https://deepmind.google/models/model-cards/gemini-3-5-flash-lite/ 2 3 4 5 6 7 8 9

  2. Google, Gemini API model page for gemini-3.5-flash-lite (input limit 1,048,576 tokens, output limit 65,536, stable, computer use in preview, Live API not supported), and the release notes entry of 21 July 2026 ("a low-latency, highly cost-effective subagent option designed for high-volume automation"). https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite and https://ai.google.dev/gemini-api/docs/changelog 2 3 4 5 6

  3. Google DeepMind, Gemini 3.1 Flash-Lite model card, 3 March 2026: "Gemini 3.1 Flash-Lite is based on Gemini 3 Pro"; "trained using Google's Tensor Processing Units (TPUs)". https://deepmind.google/models/model-cards/gemini-3-1-flash-lite/

  4. Google DeepMind, Gemini 3 Pro model card: "a sparse mixture-of-experts (MoE)" transformer, with no parameter or expert counts. https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf

  5. Google DeepMind, Gemini 3.5 Flash-Lite model evaluation, July 2026: Gemini scores are pass@1, run through the Gemini API for gemini-3.5-flash-lite "with high thinking and default sampling settings"; SWE-Bench Pro with "an internal version of the Antigravity harness"; Terminal-Bench 2.1 with the default agent harness, Terminus 2; OSWorld-Verified "averaged over 5 runs with a single attempt per run". https://storage.googleapis.com/deepmind-media/gemini/gemini_3-5_flash-lite_model_evaluation.pdf 2 3 4

  6. Google, Gemini API pricing, paid tier, standard rates per 1M tokens, page last updated 24 September 2026, read on 29 September 2026. The input row reads "$0.30 (text / image / video / audio)". https://ai.google.dev/gemini-api/docs/pricing 2 3

Altri modelli di Google DeepMind

Tutti i modelli di Google DeepMind