Gemini 3.6 Flash is here — and 3.5 Pro is still waiting
Google shipped a sharper Flash workhorse, a faster Lite tier, and 3.5 Flash Cyber inside CodeMender. Refresh Gemini web today; update the mobile apps if the new models are missing — and stop holding your breath for 3.5 Pro just yet.
Read time: ~7 minutes
Gemini 3.6 Flash and 3.5 Flash-Lite are live for everyday use. Google's July 21 announcement also introduces 3.5 Flash Cyber inside CodeMender — a multi-agent security stack for finding and patching vulnerabilities. The consumer Flash models are the efficiency layer for production agents: better coding and knowledge work, fewer tokens, lower latency — without waiting for the Pro tier a lot of people expected next.
If you use Gemini in the browser, a refresh is usually enough to see the new models. On phone, open the App Store or Google Play, update the Gemini app to the latest build, then reopen — older installs can lag a release cycle behind web.
What actually shipped
- 3.6 Flash — workhorse Flash. Google cites ~17% fewer output tokens vs 3.5 Flash on the Artificial Analysis Index (and larger savings on some agent benches like DeepSWE), with a lower output price than 3.5 Flash: $1.50 / $7.50 per million input/output tokens.
- 3.5 Flash-Lite — fastest, cheapest 3.5-class option for high-throughput agent and document workflows. Artificial Analysis clocks ~350 output tokens/s. Price: $0.30 / $2.50 per million tokens.
- 3.5 Flash Cyber — a specialized, cyber-focused Flash tuned for finding and fixing software vulnerabilities at a lower price-per-token than larger models. It does not show up as a normal pick in the Gemini consumer app.
3.5 Flash Cyber inside CodeMender
CodeMender is Google's code-security agent system. Multiple 3.5 Flash Cyber agents work together to discover vulnerabilities, validate them, and write patches for complex software issues — then merge into a single combined report. Google positions the stack as competitive on CyberGym-style benches while staying efficient enough to run at scale.
Because dual-use risk is real, Flash Cyber + CodeMender is rolling out as a limited-access pilot for governments and trusted partners — a head start for defenders, not a public model picker option for vibe coders. Worth knowing it exists; don't expect it in your Gemini mobile app update.
Google also notes Gemini 3.5 Pro is still testing with partners, with broader availability "as soon as it's ready," while the lab has already started its most ambitious pre-training run yet for Gemini 4.
Where 3.5 Pro went (and why the wait is rational)
Plenty of builders were waiting for Gemini 3.5 Pro as the next headline drop. Instead Google pushed Flash efficiency first. That reads less like "Pro is cancelled" and more like "Pro isn't ready to sit at the top of the table."
Look at the last few months alone: Opus 4.8, Fable 5, GPT-5.6, Kimi K3, and Cursor Grok 4.5 all landed in the same window. Google cannot casually ship a mid-tier Pro into that pile and call it a win. The AI race is on — and as users we benefit: better outputs now, and real pressure toward lower cost over time.
Pricing: where Flash sits
Flash is not trying to beat Opus on price-per-flagship-token theater. It's the high-volume agent layer — cheaper than prior Flash, and in the same neighborhood as budget coding models when you care about loops, not single-shot brilliance.
| Model | Input / 1M | Output / 1M | Notes |
|---|---|---|---|
| Gemini 3.6 Flash | $1.50 | $7.50 | Workhorse; fewer tokens than 3.5 Flash |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | ~350 tok/s (Artificial Analysis) |
| Grok 4.5 | $2 | $6 | Cursor / SpaceXAI API |
| GPT-5.6 Luna | $1 | $6 | Budget OpenAI tier |
| Claude Opus 4.8 | $5 | $25 | Prior-gen flagship |
Benchmarks: honest placement, not a clean sweep
Google published gains for 3.6 Flash vs its own 3.5 Flash (DeepSWE, OSWorld-Verified, MLE Bench, GDPval) and strong Flash-Lite gains vs older Lite / Flash baselines. Cross-vendor tables are messier — different harnesses, thinking levels, and "max" settings. Below we mix Google's cited scores with numbers we already published for Grok, Opus, and Fable. Treat dashes as "not in that source," not as zeros.
| Benchmark | 3.6 Flash | 3.5 Flash | Flash-Lite | Grok 4.5 | Opus 4.8 | Fable 5 |
|---|---|---|---|---|---|---|
| DeepSWE | 49%vs 3.5 Flash | 37% | — | 62%high | 55.8%max | 66.1%max |
| OSWorld-Verified | 83.0% | 78.4% | 74.0% | — | — | — |
| Terminal-Bench 2.1 | — | — | 54%vs 3.1 Lite | 83.3% | 78.9% | 84.3% |
| SWE-Bench Pro | — | — | 54.2%vs Gemini 3 Flash | 64.7%high | 69.2%max | 80.3% |
| MLE Bench | 63.9% | 49.7% | — | — | — | — |
Read: 3.6 Flash is a clear step up inside Google's Flash line — especially token efficiency and computer-use / multimodal knowledge work. It is not automatically the new SWE-Bench Pro king; Fable 5 and Opus still own several coding peaks we track, and Grok 4.5 still wins on the "near-frontier + cheap loops" story for Cursor users. Flash-Lite is the scale dial: slower brains, absurd throughput and price for agent swarms.
How this fits the recent model posts
- Opus 4.8 — still the Claude prior-gen reference many of us ship against.
- GPT-5.6 Sol / Terra / Luna — peak Terminal-Bench numbers, messy access story.
- Grok 4.5 in Cursor — co-trained SpaceXAI model priced to win agent loops.
Kimi K3 showed up in our last newsletter digest — open-weights noise in the same race, no standalone post yet.
Who should use what
- Default to 3.6 Flash if you want Google's best widely available agent workhorse today — coding, docs, multimodal, computer use.
- Use Flash-Lite when latency and volume dominate (search agents, receipt/doc pipelines, cheap subagents under a smarter orchestrator).
- Keep waiting on 3.5 Pro only if you specifically need the Pro tier story; Google has not timed a public drop. Meanwhile Flash is what you can actually refresh into.
- Stay multi-model in Cursor (or your editor of choice) — pick Gemini when Google wins the task, Grok/Claude/GPT when they don't.
- Flash Cyber / CodeMender — watch the announcement if you work in security; access is partner/government pilot only for now.
Bottom line
Google shipped the Flash upgrades users can run today, held Pro until it can compete at the top, and already started the Gemini 4 pre-train. Refresh Gemini web, update the mobile app, try 3.6 Flash on a real agent loop, and keep following the race — cheaper, sharper models are the consumer upside of everyone fighting for the leaderboard.
Sources: Google Gemini blog (Jul 21, 2026); cross-vendor benches from our prior Grok / GPT-5.6 coverage. Harnesses differ — don't treat the grid as a single official league table.
Disclosure: Some links above are referral or partner links (marked on our Tools page).
Questions? Get in touch — or subscribe for the next AI news post.