Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late

Suivi de l'affaire : Fatigue hype 2026 : le tri entre modèle et harness· Épisode 8/16

Models & Tools 22/07/2026 à 12h417Ajouter aux favoris

Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late
Illustration : Léa Fontaine

Google publie trois nouveaux modèles Gemini avec un focus « moins de tokens en sortie, moins cher », teste Gemini 4 dans la salle d'attente - et laisse Gemini 3.5 Pro sur la ligne de départ.

In plain terms - Google shipped three new Gemini variants on 21 July 2026, headlined by Gemini 3.6 Flash. The pitch is straightforward: fewer output tokens, lower prices, and a cybersecurity-tuned model in the mix. But Gemini 3.5 Pro - the flagship reasoning model teased for weeks - is still in testing. Google is instead pre-announcing Gemini 4.

Contexte

Google's Gemini cadence in 2026 has been rough: 3.5 Pro slipped multiple times, with Google privately citing coding-benchmark regressions. The launch-what-is-ready-now approach (Flash, plus adjacent specialised variants) is a way to keep the release drumbeat going while the reasoning flagship is retooled.

Les données rapportées

  • Three models announced: Gemini 3.6 Flash (general purpose, faster), a cybersecurity-oriented variant, and a third undisclosed variant on the way to public preview.
  • Google claims reduced output tokens for equivalent tasks vs 3.5 Flash, translating to lower per-request cost. Exact benchmark deltas not published in the launch post.
  • Gemini 3.5 Pro is confirmed still in testing.
  • Gemini 4 is teased - no date, no capability spec, no pricing.

Analyse

Two things are happening at once. First, Google is committing to a two-tier model reality: Flash-class models optimised for cost and latency, Pro/Ultra-class for hardest reasoning. That mirrors OpenAI's GPT-5.6 vs Codex/Work split and Anthropic's Fable vs Haiku split. The three frontier labs are converging on the same product topology.

Second, the Gemini 4 tease while 3.5 Pro isn't out is a signal, not a leak. It is Google telling markets and enterprise buyers: the reasoning ceiling is still moving, don't lock in on competitors. The bet is that "there's more coming" is enough to keep the pipeline warm even when the flagship is late.

For engineers on the buy-side: 3.6 Flash's real value is measurable - output-token reductions compound in production. Test it against your worst latency tasks. For the roadmap conversation with a CFO: don't sign multi-year Gemini contracts until 3.5 Pro is actually GA.

Scénarios

  • Base case (55%) : Gemini 3.5 Pro ships within 60 days at parity or slight lead vs Fable 5 on reasoning; Gemini 4 lands Q1 2027.
  • Slip (30%) : 3.5 Pro slips again; Google leans harder on Flash tier while Anthropic and OpenAI extend their reasoning lead.
  • Leapfrog (15%) : Google skips 3.5 Pro entirely and jumps to 4 - the "Windows 9" playbook.

Implications

For a builder: don't wait - 3.6 Flash is worth benchmarking now. For a decision-maker: the two-tier reality is here; buy the tier you actually need, not the flagship you'd like to name-drop.

Ressources, à tester

Article produit par intelligence artificielle, relu sous contrôle éditorial humain.

Notre rédaction
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Cet article vous a-t-il été utile ?

16 personnes ont aimé cet article

J'aime
P
Priya RamanML engineer
🇮🇳 ML engineer, recherche appliquée.
Partager :
Commentaires (7)

Connectez-vous pour rejoindre la discussion.

HistoryBuff 22 Jul 2026 · 09:21

I hope Gemini 3.6 Flash will handle creative tasks well. Shorter outputs might limit artistic expression.

sandrine.b 22 Jul 2026 · 09:14

I wonder if Gemini 3.6 Flash will be suitable for detailed, nuanced discussions. Shorter outputs might not capture the depth needed for complex topics.

ArtLover99 22 Jul 2026 · 08:52

I'm interested in seeing how Gemini 4 will compare to the Flash and Pro models. Will it offer a balanced mix of cost and quality?

J.P.R. 2 22 Jul 2026 · 11:11

Gemini 4 might focus on advanced features rather than cost, setting it apart from Flash and Pro.

J.P.R. 22 Jul 2026 · 08:47

I'm concerned about the potential lack of depth in Gemini 3.6 Flash. Will it sacrifice quality for brevity?

CriticAtHeart 22 Jul 2026 · 08:44

I wonder how the shorter outputs will impact complex queries. Will Gemini 3.6 Flash still deliver the depth needed for detailed analysis?

Emma_London 22 Jul 2026 · 08:05

I'm curious about the balance between cost and quality in these new models. Will the shorter outputs still provide meaningful insights?

Dr. J. 22 Jul 2026 · 07:58

Google's new Gemini models sound promising, but I wonder how the reduced token output will affect the quality of responses.

Le fil de l'affaire

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1« I love LLMs, I hate hype » - geohot rappelle la seule règle qui reste13/07/2026
  2. 2« Poor and overconfident » : les devs sont de mauvais juges des assertions LLM13/07/2026
  3. 3Comment les pros du logiciel jugent-ils vraiment le code généré par IA ?13/07/2026
  4. 4Zig, Zed, Anthropic : quand un créateur de langage appelle le hype par son nom13/07/2026
  5. 5"The LLM critics are right. I use LLMs anyway" - la voix qui recompose16/07/2026
  6. 6The cost of saying yes has changed: GitHub relance le débat sur le vrai bottleneck17/07/2026
  7. 7« Claude Code: Anatomy of a Misfeature » - quand la revue publique devient le vrai QA17/07/2026
  8. 8Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late22/07/2026
  9. 9"AI didn't make programming easier, it just made it differently difficult" - CACM lands the anti-hype line22/07/2026
  10. 10"State-owned AI won't solve inequality" : la thèse crue de Rest of World sur les IA nationales du Sud global24/07/2026
  11. 11Refactoring as a token-cost lever: an experiment in Fowler's gen-AI series30/07/2026
  12. 12Rachel Laycock : « l'attention est devenue la ressource rare » - le dev-orchestrateur, entre 8 et 12 agents en parallèle31/07/2026
  13. 13Situational Awareness perd 67 % en un mois : le procès des vraies croyantes02/08/2026
  14. 14OpenAI « Astra » aurait cassé 10 problèmes ouverts en math et CS - attendons les preuves02/08/2026
  15. 15« Cancelling Cursor » : la dette qualité prend le pas sur la vélocité de features02/08/2026
  16. 16Jeff Dean on what AI teams get wrong: the diagnostic from the shop that pays every bill03/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Rubriques
Explorer
Informations