Google DeepMind unveiled three new Gemini models on July 21, a "workhorse" lineup built for efficiency, latency, and reliability aimed at customers running AI agents at scale — squarely targeting coding and cost. Yet the most talked-about part of the launch was not what shipped, but what didn't: Google's top-tier flagship, Gemini 3.5 Pro, was a no-show once again.
Gemini 3.6 Flash — 17% Fewer Tokens, Lower Price
The centerpiece is Gemini 3.6 Flash, the successor to the 3.5 Flash unveiled at I/O 2026 in May. Reflecting developer and customer feedback, Google says the model is "more token efficient across tasks." Per the Artificial Analysis Index, it consumes 17% fewer output tokens than 3.5 Flash and takes fewer reasoning steps and tool calls to complete multi-step workflows.
The price came down, too. Output dropped from $9 to $7.50 per million tokens, with input at $1.50 per million. Higher performance at a lower price is the familiar arc of the recently intensifying AI price war.
Models Gemini 3.6 Flash / 3.5 Flash-Lite / 3.5 Flash Cyber
3.6 Flash pricing $1.50 input · $7.50 output (per 1M tokens; was $9 output)
Output tokens 17% fewer vs 3.5 Flash (Artificial Analysis Index)
Knowledge cutoff January 2025 → March 2026
Still missing Flagship Gemini 3.5 Pro (unchanged since February)
Reading the Benchmarks
By Google's own figures, 3.6 Flash improves clearly on coding and research tasks, delivering "higher precision with fewer unwanted code edits and reduced execution loops."
| Area | Benchmark | 3.5 Flash | 3.6 Flash |
|---|---|---|---|
| Coding | DeepSWE | 37% | 49% |
| ML research | MLE Bench | 49.7% | 63.9% |
| Knowledge work | GDPval-AA | 1349 | 1421 |
| Computer use | OSWorld-Verified | 78.4% | 83% |
Flash-Lite and Flash Cyber
Gemini 3.5 Flash-Lite is built for high-throughput, low-latency work such as agentic search and document processing. Google says it offers "significantly better quality" than March's 3.1 Flash-Lite, while pricing sits at $0.30 input and $2.50 output per million tokens — among the cheapest in its class. It sharply outpaces its predecessor on Terminal-Bench 2.1 (54% vs 31%) and long-context handling (GDM-MRCR v2, 72.2% vs 60.1%).
The third model, 3.5 Flash Cyber, is tuned to detect, validate, and patch code-security vulnerabilities at scale. Google's automated CodeMender tool uses agents built on it. Because of misuse concerns, it is available first only to governments and trusted partners as a limited-access pilot.
So Why Is There Still No 'Pro'?
The real story sits in the gap. Google's top-tier flagship, Gemini Pro, has not been updated since February. In the meantime, OpenAI shipped GPT-5.5 and began rolling out its GPT-5.6 family, while Anthropic launched Claude Opus 4.8 and Sonnet 5 and widened access to Fable 5 — underscoring how fast rival labs have been shipping.
Last week, Bloomberg reported that Google was delaying 3.5 Pro after falling short of internal performance goals (see our July 19 report). On July 21, DeepMind product lead Logan Kilpatrick said 3.5 Pro is testing with partners and that the team hopes it will "land soon."
What It Means
With this lineup, Google banks the practical wins of cheaper, faster, more reliable production models. Lower prices and token savings translate directly into cost cuts for enterprises running agents in bulk. At the same time, DeepMind signaled where the next fight is headed, saying it has "already started" its most ambitious pre-training run for Gemini 4. The open question is when the Pro gap closes. If Google holds the base with value-tier models but fails to finish the flagship, it risks continuing to cede the highest-end coding and reasoning workloads to its rivals.
· Google (The Keyword) — Official announcement: Gemini 3.6 Flash, 3.5 Flash-Lite, Flash Cyber
· TechCrunch — Google releases three new Gemini models — but no 3.5 Pro (7/21)
· 9to5Google — Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4 (7/21)
· Bloomberg — Gemini launch delayed as tech falls short of internal goals (7/16, background)
- Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21
- 3.6 Flash: 17% fewer output tokens, output price cut from $9 to $7.50, cutoff now March 2026
- Broad gains — coding (DeepSWE 49% vs 37%) and ML research (MLE Bench 63.9% vs 49.7%)
- Flash Cyber targets vulnerability detection/patching, limited to governments and trusted partners
- Flagship 3.5 Pro absent again — DeepMind says pre-training for Gemini 4 has begun