AI-Economy

Google launches Gemini 3.6 Flash and cheaper Flash-Lite

3 min read
Rows of servers in a Google data center with a monitor displaying the Gemini logo and a pricing table, symbolizing the launch of Gemini 3.6 Flash and 3.5 Flash-Lite Image generated with GPT Image 2
Rows of servers in a Google data center with a monitor displaying the Gemini logo and a pricing table, symbolizing the launch of Gemini 3.6 Flash and 3.5 Flash-Lite

TL;DR Too Long; Didn’t read

Google released two new AI models on July 21, 2026: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The price of Flash drops from nine to $7.50 per million tokens, Flash-Lite starts at 30 cents. An independent analysis estimates the real cost per task falls by up to 31 percent thanks to added efficiency gains.

Key takeaways

  • Gemini 3.6 Flash costs $7.50 instead of the previous nine dollars per million output tokens.
  • Flash-Lite targets high-volume workloads and processes 350 tokens per second, according to Google.
  • Flash's knowledge cutoff moves to March 2026, up from January 2025 previously.
  • Flagship model Gemini 3.5 Pro remains in testing with select partners only.
  • Google says it has begun training its next model, Gemini 4.
  • Both new models run in the Gemini app, in AI Studio, and via the Gemini API.

Google presented two new AI models on July 21, 2026: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite replace the previous Flash variants. The output price of Flash drops from nine to $7.50 per million tokens, while Flash-Lite starts at just 30 cents. Both models are said by Google to need fewer computation steps for the same task.

Flash gets faster and cheaper for programming tasks

Gemini 3.6 Flash replaces the previous Flash model as Google DeepMind’s fast all-round AI for programming tasks, research, and image processing. As Google announces in its blog post, the price for output tokens drops from nine to $7.50 per million, while the input price stays at $1.50. According to market researcher Artificial Analysis Index, the model needs 17 percent fewer output tokens for equivalent tasks than its predecessor - so the price cut and the efficiency gain add up to a larger overall cost saving.

In programming tasks, the progress shows up in the DeepSWE benchmark, which measures software-engineering tasks drawn from real GitHub projects: the score rises from 37 to 49 percent of tasks solved. In the OSWorld-Verified benchmark, which tests autonomous operation of desktop programs, Flash reaches 83 percent, up from 78.4 percent. The model’s knowledge cutoff moves from January 2025 to March 2026. Developers reach Gemini 3.6 Flash via Google AI Studio, the Google Antigravity coding environment, and the Gemini Enterprise Agent Platform, while consumers get it through the Gemini app.

Flash-Lite targets high-volume workloads

Gemini 3.5 Flash-Lite is built for tasks that need large amounts of text processed quickly and cheaply, such as automated document review or search agents. The model costs 30 cents per million input tokens and $2.50 per million output tokens. According to the Artificial Analysis Index, Flash-Lite delivers up to 350 tokens per second - enough to generate a full page of text in roughly a second and a half.

In the Terminal-Bench 2.1 benchmark, which tests command-line tasks performed by AI agents, the success rate climbs from 31 to 54 percent compared with the March predecessor, 3.1 Flash-Lite. On the SWE-Bench Pro coding benchmark, Flash-Lite reaches 54.2 percent, edging out the larger, regular Flash model’s 49.6 percent - an unusual result for a lightweight model variant. Beyond the familiar access routes through AI Studio and the Gemini API, Google says it is also rolling the model out in Google Search. For users in Germany and the rest of the EU, both new models are available without restriction, unlike the simultaneously unveiled security variant Flash Cyber, which for now remains limited to governments.

Flagship Gemini 3.5 Pro stays absent as Gemini 4 training begins

One thing was notably missing from the announcement: Google’s flagship Gemini 3.5 Pro, promised for the second quarter of 2026, remains delayed and is currently only in testing with select partners, according to the company. Google gives no new launch date. Instead, the company said it has begun training its next major model, Gemini 4 - describing it as its “most ambitious pre-training run” yet.

The new Flash generation lands amid intensifying price competition among major AI providers: earlier in July, OpenAI introduced three new pricing tiers with GPT-5.6, with the cheapest tier, Luna, priced at one dollar per million input tokens. Google also adds tightened safeguards to Flash 3.6 against misuse for chemical, biological, or nuclear purposes as well as cyberattacks, according to its model report.

What matters now is whether Google can show similar efficiency gains once the overdue flagship Gemini 3.5 Pro ships - its repeated delays increasingly look like a structural problem with Google’s internal coding benchmarks. Until a new date is set, the leaner Flash line remains the only visible progress on Google’s current model roadmap.

Frequently asked questions

What does Gemini 3.6 Flash cost via the API?

Usage costs $1.50 per million input tokens and $7.50 per million output tokens, billed through Google AI Studio or the Gemini API.

Where is Gemini 3.5 Flash-Lite available?

Developers can access it via AI Studio, Android Studio, and the Gemini API; consumers through the Gemini app; Google Search access follows in the coming weeks, according to the company.

Can users in Germany or the EU already use the new models?

Yes, both models are regularly accessible in Germany and the rest of the EU, unlike the simultaneously announced security variant Flash Cyber, which is currently limited to selected governments.

When will Gemini 3.5 Pro launch?

Google still gives no date; the flagship is reportedly in tests with partners after internal coding results fell short of the target.

What is known about Gemini 4?

Google has only confirmed that training has begun and calls it its "most ambitious pre-training run" to date; the company gives no details on capabilities or timing.


← Back to the blog