Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late

Ongoing story : Fatigue hype 2026 : le tri entre modèle et harness· Part 8/9

Models & ToolsSubscribers only 3 h ago7Add to bookmarks

Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late
Illustration : Léa Fontaine

Google publishes three new Gemini models with a focus on "fewer output tokens, cheaper", tests Gemini 4 in the waiting room - and leaves Gemini 3.5 Pro at the starting line.

In plain terms - Google shipped three new Gemini variants on 21 July 2026, headlined by Gemini 3.6 Flash. The pitch is straightforward: fewer output tokens, lower prices, and a cybersecurity-tuned model in the mix. But Gemini 3.5 Pro - the flagship reasoning model teased for weeks - is still in testing. Google is instead pre-announcing Gemini 4.

Contexte

Google's Gemini cadence in 2026 has been rough: 3.5 Pro slipped multiple times, with Google privately citing coding-benchmark regressions. The launch-what-is-ready-now approach (Flash, plus adjacent specialised variants) is a way to keep the release drumbeat going while the reasoning flagship is retooled.

Les données rapportées

  • Three models announced: Gemini 3.6 Flash (general purpose, faster), a cybersecurity-oriented variant, and a third undisclosed variant on the way to public preview.
  • Google claims reduced output tokens for equivalent tasks vs 3.5 Flash, translating to lower per-request cost. Exact benchmark deltas not published in the launch post.
  • Gemini 3.5 Pro is confirmed still in testing.
  • Gemini 4 is teased - no date, no capability spec, no pricing.

Analyse

Two things are happening at once. First, Google is committing to a two-tier model reality: Flash-class models optimised for cost and latency, Pro/Ultra-class for hardest reasoning. That mirrors OpenAI's GPT-5.6 vs Codex/Work split and Anthropic's Fable vs Haiku split. The three frontier labs are converging on the same product topology.

Second, the Gemini 4 tease while 3.5 Pro isn't out is a signal, not a leak. It is Google telling markets and enterprise buyers: the reasoning ceiling is still moving, don't lock in on competitors. The bet is that "there's more coming" is enough to keep the pipeline warm even when the flagship is late.

For engineers on the buy-side: 3.6 Flash's real value is measurable - output-token reductions compound in production. Test it against your worst latency tasks. For the roadmap conversation with a CFO: don't sign multi-year Gemini contracts until 3.5 Pro is actually GA.

Scénarios

  • Base case (55%) : Gemini 3.5 Pro ships within 60 days at parity or slight lead vs Fable 5 on reasoning; Gemini 4 lands Q1 2027.
  • Slip (30%) : 3.5 Pro slips again; Google leans harder on Flash tier while Anthropic and OpenAI extend their reasoning lead.
  • Leapfrog (15%) : Google skips 3.5 Pro entirely and jumps to 4 - the "Windows 9" playbook.

Implications

For a builder: don't wait - 3.6 Flash is worth benchmarking now. For a decision-maker: the two-tier reality is here; buy the tier you actually need, not the flagship you'd like to name-drop.

Content reserved for members

Create a free account to access all our content and the weekly review.

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

7 people liked this article

Like
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
Share:
Comments (7)

Sign in to join the discussion.

HistoryBuff 22 Jul 2026 · 09:21

I hope Gemini 3.6 Flash will handle creative tasks well. Shorter outputs might limit artistic expression.

sandrine.b 22 Jul 2026 · 09:14

I wonder if Gemini 3.6 Flash will be suitable for detailed, nuanced discussions. Shorter outputs might not capture the depth needed for complex topics.

ArtLover99 22 Jul 2026 · 08:52

I'm interested in seeing how Gemini 4 will compare to the Flash and Pro models. Will it offer a balanced mix of cost and quality?

J.P.R. 2 22 Jul 2026 · 11:11

Gemini 4 might focus on advanced features rather than cost, setting it apart from Flash and Pro.

J.P.R. 22 Jul 2026 · 08:47

I'm concerned about the potential lack of depth in Gemini 3.6 Flash. Will it sacrifice quality for brevity?

CriticAtHeart 22 Jul 2026 · 08:44

I wonder how the shorter outputs will impact complex queries. Will Gemini 3.6 Flash still deliver the depth needed for detailed analysis?

Emma_London 22 Jul 2026 · 08:05

I'm curious about the balance between cost and quality in these new models. Will the shorter outputs still provide meaningful insights?

Dr. J. 22 Jul 2026 · 07:58

Google's new Gemini models sound promising, but I wonder how the reduced token output will affect the quality of responses.

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information