Skip to content

Google releases three Gemini models at once, cheaper and faster, but still not the one we're waiting for.

Gemini 3.6 Flash, Flash-Lite and Flash Cyber arrive today. Better, more economical, and for once cheaper than the previous version. But the flagship itself remains nowhere to be found.

Advertisement
The essentials in 30 seconds ⚡
On 21 July 2026, Google launched three Gemini models from its fast range at once: 3.6 Flash, the new go-to model for everyday work, 3.5 Flash-Lite, ultra cost-effective, and 3.5 Flash Cyber, specialised in cybersecurity and reserved for governments. Notable point: the new flagship model is better AND cheaper than the one it replaces. But the true flagship, Gemini 3.5 Pro, remains nowhere to be found, and Google is taking the opportunity to tease Gemini 4.

We have reported twice on the repeated delays of Google's flagship model, in a first article and then by noting that it had missed its 17 July date again. Today, Google responds in its own way: rather than delivering the expected model, it releases three other models, more modest but very real. This is the logical continuation of a strategy that has become clear. Let's break down what just happened.

The "Flash" range, what exactly is it

At Google, Gemini models fall into two families. There's the Pro range, the most powerful, built for the most advanced reasoning, and that's the one that's late. And there's the Flash range, optimised not for maximum power but for speed, reduced cost and handling large volumes. It's exactly this second family that Google has just expanded.

This positioning is strategic. Most real-world uses of AI, especially autonomous agents that chain tasks together, don't need the smartest model in the world. They need a decent, fast model, and above all a cheap one, because they run in loops over millions of requests. The Flash range targets precisely this market, the one of volume rather than prowess.

The three new models

Gemini 3.6 Flash is the star of the day, the new default model for coding, knowledge work and multimodal tasks. Its main selling point is not so much raw power as efficiency. On the Artificial Analysis index, it consumes 17% fewer output tokens than the 3.5 Flash model it replaces, and up to 65% fewer on certain coding tests. It also completes multi-step tasks with fewer round-trips and tool calls.

Gemini 3.5 Flash-Lite is the ultra cost-effective model of the family, built for very high-volume tasks such as translation, document analysis or agentic search. It runs at 350 output tokens per second, a high speed, and offers adjustable thinking levels to trade off cost against quality, the mechanism we detailed in our article on reasoning models.

Gemini 3.5 Flash Cyber is the most specialised, and the most closely watched. Built on 3.5 Flash and then fine-tuned to find, validate and fix software security flaws, it is only accessible to governments and trusted partners via a pilot programme. We'll come back to it below, because it illustrates an interesting technical idea.

Model Input / output price (per million tokens) What for
Gemini 3.6 Flash $1.50 / $7.50 Everyday work, coding, multimodal
Gemini 3.5 Flash-Lite $0.30 / $2.50 High volumes, translation, documents
Gemini 3.5 Flash Cyber Restricted access (pilot) Cybersecurity, governments and partners

The surprising detail: cheaper than the old one

There's a small, welcome anomaly in this release. Usually, a new, more capable model costs at least as much as its predecessor. Here, it's the opposite. While input tokens remain at $1.50 per million, output tokens drop to $7.50, versus $9 on the 3.5 Flash. Better and cheaper at the same time.

This isn't generosity, it's competitive strategy. The market for fast, affordable models has become a fierce battleground, where Chinese labs like Moonshot with Kimi K3 or Zhipu with GLM-5.2 keep constant pressure on prices. Google has no choice: to keep developers, it must make its models both better and cheaper with each generation. It's the consumer who wins from this war.

The Cyber model, a clever technical idea

Gemini 3.5 Flash Cyber deserves a closer look, because it illustrates an interesting reversal of logic. To find deep security flaws in code, you need to explore a gigantic space of possibilities. The classic approach would be to throw a single very powerful model at the problem. Google is betting the opposite: use a small and cheap model, but call it a very large number of times in parallel.

Concretely, inside Google's CodeMender security tool, several Flash Cyber agents work simultaneously and merge their results into a single report. It's the same philosophy that made it possible to solve a mathematical conjecture with 64 sub-agents: rather than a single, expensive brain, a swarm of coordinated small brains. The announced results are solid. Tested on the V8 JavaScript engine, the model found 55 confirmed unique vulnerabilities, versus 47 for standard Flash and 36 for Claude Opus 4.6, including 10 that no other model had identified.

Why Google keeps this model under lock and key 🔐
A tool that finds security flaws to fix them is also a tool that finds flaws to exploit them. Google explicitly acknowledges this: the model is as useful for attack as for defence. That's why access is limited to governments and trusted partners. This is exactly the dilemma we explored in our article on jailbreaks: the more capable a cybersecurity tool is, the more it must be controlled, because the same skill serves the best and the worst.

The elephant in the room, and the Gemini 4 teaser

It's impossible to ignore the context. These three models arrive precisely because the real flagship, Gemini 3.5 Pro, still isn't here. According to Bloomberg reports, the internal delays reportedly stem from the Pro model underperforming on coding tests, forcing teams back into retraining. Google confirms it remains in testing with partners, with no firm date for the public.

And that's where Google plays a clever communications move. Rather than apologising for the delayed Pro, the company announces it has already launched "its most ambitious training run to date" for the next generation, Gemini 4. Deflecting attention from the late model by waving the next one around: the manoeuvre is transparent, but it sends a signal, one of a company that wants to show it's already looking beyond its current problem.

What to take away

This trio of models is concrete good news for developers and businesses: better tools, cheaper, available immediately in the Gemini app, Google Search and development platforms. For real everyday use, especially agents running at scale, these fast models often matter more than the long-awaited flagship.

But in terms of narrative, the situation remains uncomfortable for Google. Releasing three excellent secondary models doesn't make the absence of the main model go away, promised for June and still absent at the end of July. Google is moving forward, but sideways: it delivers a lot, except the one thing everyone is waiting for. The real answer will come the day Gemini 3.5 Pro finally ships, and after such a delay, it won't have the right to be merely good. In the meantime, the price war this release fuels mainly benefits one party: you.

Advertisement