
Gemini 2.5 revisited: the generation where Google's reasoning caught up, and how it reads now that 3.6 Flash sells that capability at $1.50/M input.
📌 2026 update: This is an archived review of Gemini 2.5, Google's 2025 "thinking model" generation. The line has since crossed into the 3.x era — 3.6 Flash is the current workhorse (July 2026), with Gemini Omni leading the multimodal push. For the current state, see our Gemini updates tracker. We keep this review live because 2.5 was the release where Google stopped chasing and started trading blows.
Gemini 2.5 — led by 2.5 Pro — was the generation where Google's reasoning became genuinely competitive: a "thinking model" that deliberated before answering, massive context, and the deepest ecosystem integration in the industry.
It aged as a turning point rather than a peak. The 2.5 generation proved Google could match frontier reasoning; the 3.x strategy that followed — pushing frontier-adjacent capability into the cheap Flash tier rather than defending a Pro flagship — is what actually changed Google's market position.
No 3.5 Pro ever shipped. For a company that had defended a Pro flagship for every prior generation, quietly not releasing one is the loudest possible statement about where the strategy went. The Flash-first lesson came straight from 2.5's economics.
2.5 Pro reasoned through problems visibly, trading latency for reliability on hard tasks. The "thinking budget" concept let developers tune that trade — the first mainstream instance of exposing the compute/quality dial to the caller rather than hiding it behind a model name.
Million-token-class context windows made "load the whole codebase or corpus" workflows practical. This wasn't just a bigger number; it changed which problems were addressable without a retrieval pipeline.
Search grounding, Workspace integration, and the free-tier generosity that became Google's signature competitive weapon. Distribution was always Google's structural advantage; 2.5 was the first generation where the model was good enough for that advantage to matter.
Against its 2025 rivals it traded blows: competitive on reasoning benchmarks, ahead on context length and price-performance, behind the leaders on agentic coding depth — a split that remained true of the line into 2026.
2.5 was the generation where Google stopped shipping one model and started shipping a tier. That structure — a reasoning flagship, a cost-optimized workhorse, and a minimal-latency option — is the shape every vendor converged on, and it originated as a response to 2.5 Pro being too expensive to use as a default.
| Tier | Intended job | What it traded away |
|---|---|---|
| 2.5 Pro | Hard reasoning, long-context analysis | Cost and latency |
| 2.5 Flash | High-volume default | Peak reasoning depth |
| 2.5 Flash-Lite | Latency-critical, simple tasks | Most reasoning capability |
The interesting part is what happened next: the workhorse tier absorbed the flagship's capability faster than anyone forecast, which is precisely why no 3.5 Pro was needed.
2.5 exposed a compute-per-request dial to the caller. Set it low for classification and extraction; set it high for multi-step analysis. Before this, the only lever was picking a different model — a coarse, expensive choice made at integration time rather than per-request.
This is the design decision from 2.5 with the longest tail. Every subsequent reasoning model from every vendor exposes some version of it.
2.5's real innovation was making strong reasoning cheap enough to default to. The 3.x Flash line industrialized exactly that. Current pricing for 3.6 Flash:
| Item | Price |
|---|---|
| Input | $1.50 / M tokens |
| Output | $7.50 / M tokens (down from $9 on 3.5 Flash) |
| Cached input | $0.15 / M tokens |
| Context window | 1,048,576 in / 65,536 out |
Google cut output pricing while improving capability — the exact move 2.5's economics predicted.
Here is the honest complication. On the Artificial Analysis Intelligence Index, 3.6 Flash and 3.5 Flash both score 50 — no measured improvement. But on Google's own benchmarks, 3.6 posts substantial gains:
| Benchmark | 3.5 Flash | 3.6 Flash |
|---|---|---|
| DeepSWE | 37% | 49% |
| OSWorld-Verified | 78.4% | 83.0% |
| MLE-Bench | 49.7% | 63.9% |
| GDPval-AA v2 | 1349 Elo | 1421 Elo |
It also uses ~17% fewer output tokens on the Artificial Analysis Index — meaning the same score costs less to reach.
How to read this: third-party aggregate indices and vendor-selected benchmarks measure different things. The token-efficiency gain is real and independently observable; the capability jumps are on suites Google chose. Treat both as partial.
3.6 Flash launched July 21, 2026 into a three-week crush: Claude Sonnet 5 (June 30), GPT-5.6 GA (July 9), Kimi K3 (July 17). On public API rates, GPT-5.6 Luna remains cheaper; Claude Sonnet 5 costs roughly twice as much per output token.
The price-performance thesis, completely. Every subsequent Google release has been an elaboration of "make good reasoning cheap enough to be the default."
Launch-era coverage (including this article's original version) leaned on benchmark-war framing — "champion," "best thinking model." The honest 2026 read: benchmark leads in that era rotated monthly among the big three. What compounded was distribution and unit economics, where Google's advantages were structural, not momentary.
"Best model" framing itself. In a market where four frontier releases land in three weeks, no single model holds a defensible "best" position long enough for the label to be useful to a buyer.
The 2.5 generation is legacy — succeeded by 3.1 Flash (March), 3.5 + Omni (May's "agentic era" I/O), and the July trio of 3.6 Flash, 3.5 Flash-Lite and the security-specialized Flash Cyber.
The through-line from 2.5 is unmistakable: thinking made cheap, then made ubiquitous, then made specialized.
| Month (2026) | Release |
|---|---|
| March | 3.1 Flash |
| May | 3.5 + Gemini Omni — the "agentic era" I/O |
| July | 3.6 Flash, 3.5 Flash-Lite, Flash Cyber |
Three Flash-tier releases in five months, and no Pro among them. The cadence itself is the argument.
Flash Cyber — a security-shaped variant — is the first sign that the tier structure is splitting along domain rather than only along cost. If that pattern holds, "which Gemini" becomes a question about your workload's shape, not just your budget.
The May release pushed multimodality into the flagship slot that a Pro model would otherwise occupy. Read that as Google answering "what replaces Pro?" with "a different axis entirely" rather than with a bigger reasoning model.
No — start from the current lineup. 2.5 is retired and its successors are cheaper and better.
Yes. Gemini 2.5 is the clearest case study available of a company converting a capability parity moment into a durable structural advantage, by choosing distribution economics over flagship prestige.
Migration to 3.6 Flash is a pricing improvement, not just a capability one — output dropped from $9 to $7.50 per million while token efficiency improved ~17%. The migration pays for itself.
The 3.6 Flash case is instructive precisely because the two disagree. When a vendor posts large gains on suites it selected while an independent aggregate index shows none, the truthful summary is "improved on these specific tasks," not "improved."
A 17% reduction in output tokens at equal score is a real cost saving that no leaderboard position captures. For production workloads this often matters more than a few points of benchmark movement.
Google lowering output pricing from $9 to $7.50 while shipping a capability update is not a promotion — it's a statement that the Flash tier is where the volume is expected to live.
Sources: Google AI updates archive · Gemini API changelog · Memeburn — Gemini 3.6 Flash benchmarks and pricing · Digital Applied — 3.6 Flash per-task price analysis · OpenRouter — Gemini 3.6 Flash API pricing · TechCrunch on the 3.x releases
Originally published July 2025 · Substantially revised July 30, 2026
Seedance 2.5 brings 30-second audio-video generation, 50 mixed references and timed editing. Here is what is official, priced and still worth testing.
MiniMax H3 generates 2K video with native stereo audio, publishes per-second pricing, and offers H3-Base weights under a Community License.
Hermes Agent blends persistent memory with self-improving skills, but its security and outcome depend on the backend, approvals, and review you choose.
OpenClaw is a high-capability agent runtime whose value depends on deliberate gateway security, skill review, and operating discipline.
AI music in 2026: Suno v5.5's capabilities and legal defiance, Udio's settlement path, the Warner deal retiring unlicensed models, and creator rules.
The AI video market after Sora's exit — with real pricing: $0.50 to $2.50 per 10-second clip, who leads the rankings now, and how to choose per shot.
The honest 2026 free-tier comparison: Gemini's generous bundle, Claude's quality-first plan, ChatGPT's capped breadth, and when free stops being enough.
Archived review of GPT-5's August 2025 launch: the unified-router design, the 4o backlash, what improved — and how it reads from mid-2026.