
Every major Gemini update through July 2026: the 3.5/3.6 Flash generation, Gemini Omni multimodal creation, Nano Banana 2, and the agentic-era strategy.
Google shipped three new Gemini models on July 21: 3.6 Flash (the new workhorse โ better coding, knowledge work and multimodal performance while cutting token usage up to 17%), 3.5 Flash-Lite (cheapest in class) and 3.5 Flash Cyber (fine-tuned for finding and fixing security vulnerabilities). Notably absent: a 3.5 Pro. The year's theme is Google's "agentic era" framing โ Gemini 3.5 + Gemini Omni headlined I/O in May, with Omni positioned as "create anything from any input, starting with video."
| Date | Release | What It Is |
|---|---|---|
| Mar 2026 | Gemini 3.1 Flash-Lite ยท 3.1 Flash Live | Speed/cost-optimized pair |
| May 2026 (I/O) | Gemini 3.5 ยท Gemini Omni | Flagship reasoning + any-input-to-any-output creation |
| Jun 2026 | Nano Banana 2 Lite ยท Omni Flash (API preview) | Fastest/cheapest image model; multimodal video workflows for devs |
| Jul 21, 2026 | 3.6 Flash ยท 3.5 Flash-Lite ยท 3.5 Flash Cyber | New workhorse (โ17% tokens) ยท cost floor ยท security specialist |
Flash is the franchise โ Google keeps pushing frontier-adjacent capability into its cheap tier rather than reserving it for a Pro SKU (no 3.5 Pro shipped).
Flash Cyber is Google's first mainline model fine-tuned for a single professional domain. If security teams adopt it, "which Gemini" becomes a question about your workload's shape rather than only your budget.
| Item | Price |
|---|---|
| Input | $1.50 / M tokens |
| Output | $7.50 / M tokens (down from $9 on 3.5 Flash) |
| Cached input | $0.15 / M tokens |
| Context window | 1,048,576 in / 65,536 out |
Google cut output pricing while shipping a capability update โ the clearest possible statement that the Flash tier is where volume is expected to live.
Here is the honest complication. On the Artificial Analysis Intelligence Index, 3.6 Flash and 3.5 Flash both score 50 โ no measured improvement. On Google's own benchmarks, 3.6 posts substantial gains:
| Benchmark | 3.5 Flash | 3.6 Flash |
|---|---|---|
| DeepSWE | 37% | 49% |
| OSWorld-Verified | 78.4% | 83.0% |
| MLE-Bench | 49.7% | 63.9% |
| GDPval-AA v2 | 1349 Elo | 1421 Elo |
Third-party aggregate indices and vendor-selected benchmarks measure different things. The ~17% token-efficiency gain is independently observable; the capability jumps are on suites Google chose. Treat both as partial and neither as the whole picture.
3.6 Flash launched into a three-week crush: Claude Sonnet 5 (June 30), GPT-5.6 GA (July 9), Kimi K3 (July 17). On public API rates GPT-5.6 Luna is cheaper โ especially after OpenAI's July 30 cut โ while Claude Sonnet 5 costs roughly twice as much per output token at standard pricing.
Introduced at I/O as a leap in "world understanding, multimodality and editing," Omni generates and edits across modalities from any input, with video first. Omni Flash hit API public preview in June for enterprises building dynamic video workflows โ putting Google in direct competition with the post-Sora video market it now partially leads through Veo 3.1.
Nano Banana 2 Lite (June) became the fastest, most cost-efficient Gemini image model โ the economy tier under the main Nano Banana 2 line, and a key ingredient in Gemini's uncommonly generous free tier, which bundles image generation alongside voice, Deep Research, NotebookLM and a monthly Veo allowance.
3.6 Flash's ~17% token reduction is the underrated number: the same workloads simply cost less to run, independent of any benchmark movement. For production traffic this frequently matters more than a few points of index position.
Flash Cyber opens a template. If it lands, domain-tuned Flash models become a product line rather than an experiment โ and model selection gets a new axis.
The API changelog cadence (monthly, sometimes weekly) makes pinning capabilities over model names as relevant on Gemini as on GPT. Documentation written around a specific model string has a short shelf life here.
Google's free bundle โ image generation, voice, Deep Research, NotebookLM and a monthly Veo allowance โ remains the most generous in the category and is the mechanism by which capability turns into share (free-stack comparison).
The 2.5 generation โ 2025's "thinking model" breakthrough โ is now legacy. Our original deep-dive is preserved as an archived Gemini 2.5 review with a 2026 retrospective.
Sources: TechCrunch on the July releases ยท Google AI updates โ May 2026 ยท Google AI updates โ June 2026 ยท Gemini API changelog
Last updated: July 29, 2026
Anthropic's 2026 in business terms: the $65B Series H at a $965B valuation, confidential IPO filing, $47B run-rate, and enterprise adoption of Claude.
Every major ChatGPT update through July 2026: the GPT-5.5/5.6 transition, Dreaming V3 memory, Lockdown Mode, and what free vs paid gets you now.
Every major Claude update through July 2026: the Claude 5 family, Reflect dashboard, MCP upgrade, Claude Code milestones, pricing and access.
Copilot's 2026: Agent Mode general availability, agentic code review that ships fix PRs, the Copilot CLI, desktop app, and much faster indexing.
Midjourney's 2026 in full: V8 alpha to the V8.2 default, native 2K without upscaling, 50% faster and 25% cheaper renders, and 21-second video.
Notion's 2026: autonomous Custom Agents, the credit billing switch on May 4, agents that read your files, 35โ50% cheaper runs, and what it costs now.
OpenAI's 2026 from the platform side: the GPT-5.x release-and-retire cadence, the Sora shutdown, API changes, and where the company stands mid-year.
Perplexity's browser year: Comet free on every platform since March, iPad and iOS upgrades, Deep Research generating decks, and enterprise controls.