TL;DR: Gemini has two live generations in August 2026. Gemini 3 is the frontier family (3.1 Pro, 3.7 Flash, 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite). Gemini 2.5 (Pro, Flash, Flash-Lite) is the cheaper, proven fallback. Prices run from $0.10 to $12 per million tokens.
Google ships Gemini in two live generations. Gemini 3 is the current frontier family. Gemini 2.5 is the older but still useful family. Inside each, the tiers follow the same names. Pro is the most capable; Flash is the fast, balanced pick; Flash-Lite is the cheapest (Check out the full comparison here). The flagship is Gemini 3.1 Pro, and the newest model is Gemini 3.7 Flash, which Google launched on August 13, 2026, just 23 days after Gemini 3.6 Flash. The long-promised Gemini 3.5 Pro has slipped again and is still not out.
Key takeaways
- Two live families: Gemini 3 (frontier) and Gemini 2.5 (cheaper, proven). Each has Pro, Flash, and Flash-Lite.
- Gemini 3.7 Flash, released August 13, 2026, is the newest and strongest Flash. It is the model to start on for coding and agents.
- 3.7 and 3.6 Flash both run at an introductory $0.75 in and $3.75 out per million tokens through December 31, 2026. Both double to $1.50 and $7.50 on January 1, 2027.
- Every Gemini 3 model has the same 1M token context window and a 64K output cap. There is no 2M tier.
- Output prices range from $0.40 per million on 2.5 Flash-Lite to $12 per million on 3.1 Pro, or $18 above a 200K prompt.
- The names are out of sync. Flash has run ahead to 3.7 while Pro is still 3.1, so read the version and the tier together.
The current Gemini models at a glance
Here are the main text and coding models you can call today, with the exact model ID your engineers will type into the API. Google also ships image, video, audio, and embedding models, which we cover further down. Every Gemini 3 model reads up to 1M tokens and writes at most 64K, so the table below varies on price and depth, not on window size.
| Model | Model ID | Context | Price in / out (per 1M) | Best for |
|---|---|---|---|---|
| Gemini 3.1 Pro | gemini-3.1-pro-preview | 1M | $2 / $12 | The hardest reasoning, agents, and creative work |
| Gemini 3.7 Flash | gemini-3.7-flash | 1M | $0.75 / $3.75 | The newest model and the place to start for coding and agents |
| Gemini 3.6 Flash | gemini-3.6-flash | 1M | $0.75 / $3.75 | Same price as 3.7, kept for pipelines already tuned to it |
| Gemini 3.5 Flash | gemini-3.5-flash | 1M | $1.50 / $9 | Legacy Flash, now the most expensive of the three |
| Gemini 3 Flash | gemini-3-flash-preview | 1M | $0.50 / $3 | Frontier-class work at a lower rate, still in preview |
| Gemini 3.5 Flash-Lite | gemini-3.5-flash-lite | 1M | $0.30 / $2.50 | High throughput that still needs some reasoning |
| Gemini 3.1 Flash-Lite | gemini-3.1-flash-lite | 1M | $0.25 / $1.50 | The cheapest Gemini 3, simple high-volume tasks |
| Gemini 2.5 Pro | gemini-2.5-pro | 1M | $1.25 / $10 | Proven deep reasoning, a solid fallback |
| Gemini 2.5 Flash | gemini-2.5-flash | 1M | $0.30 / $2.50 | Low-cost reasoning at high volume |
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite | 1M | $0.10 / $0.40 | The cheapest and fastest multimodal model |
Two caveats on those numbers. The $0.75 and $3.75 rates on 3.7 and 3.6 Flash are introductory and run only to December 31, 2026, after which both models list at $1.50 and $7.50. And on Gemini 3.1 Pro and 2.5 Pro the price is tiered by prompt size. Cross 200K tokens in a single request and 3.1 Pro moves to $4 in and $18 out, while 2.5 Pro moves to $2.50 and $15.
Which Gemini model fits your team?
Pick your main use case to see the right model.
For chatbots, tagging, and routing at scale, Gemini 3.1 Flash-Lite, 3.5 Flash-Lite, or 2.5 Flash-Lite keep costs low. You only pay for Pro when a task truly needs it. See the cost of building this in-house on our developer rate card.
For everyday coding, reviews, and agent tasks, 3.7 Flash is fast and strong at an introductory $0.75/$3.75. Pair it with engineers who build with the Gemini API. Hire AI agent developers who use it daily.
For deep reasoning, long context, and creative work, the flagship earns its price. For the hardest science and research, Deep Think goes further. Strong AI and machine learning engineers get the most from it.
The model is the easy part. The hard part is the team that builds and ships with it. We match you with pre-vetted engineers in about 24 hours. Tell us what you need.
Two generations, explained in plain English
Google keeps two families live at once. Gemini 3 is the new frontier family. Gemini 2.5 is the older family that still works well and costs less. Each family uses the same three tier names. Pro is the most capable. Flash is the balanced everyday pick. Flash-Lite is the cheapest. Once you know the tier names, the lineup is simpler than it looks.

One thing trips up most teams. The version numbers are out of sync across tiers. The newest Flash is Gemini 3.7 Flash, but the top Pro is still Gemini 3.1 Pro. Flash has now shipped three releases in the time Pro shipped none. A higher number on Flash does not mean it sits above Pro on every task. Read the tier and the version together. When in doubt, Pro is the deepest thinker and Flash is the fast worker.
The Gemini 3 family
Gemini 3.1 Pro
Gemini 3.1 Pro is the flagship. It is built for complex problems, agent work, and creative tasks. It reads up to 1,048,576 tokens, roughly 750,000 words, and writes at most 65,536, which is the same window every other Gemini 3 model gets. Pricing is $2 input and $12 output per million tokens for prompts up to 200K tokens. Above 200K, both rates rise to $4 and $18, so very large prompts cost roughly twice as much per token. Note that it is still a preview model, published as gemini-3.1-pro-preview, so pin it deliberately and expect the ID to change when it reaches general availability.

Gemini 3.1 Deep Think
Deep Think is a high-reasoning mode built on the 3.1 base. Google aims it at the hardest problems in science, research, and engineering. It thinks longer before it answers, which raises both quality and cost on tough tasks. Most teams will not need it for daily work. Reach for it when a normal Pro answer is not deep enough.
Gemini 3.7 Flash
Gemini 3.7 Flash landed on August 13, 2026 and is where most teams should now start. Google calls it its most intelligent workhorse model for coding and agents, and the benchmark jumps over 3.6 Flash are large rather than incremental. DeepSWE v1.1 went from 49.0% to 65.3%, AutomationBench from 17.0% to 30.4%, and OSWorld-2.0 computer use from 33.8% to 47.9%. Long-context recall improved too, from 91.8% to 97.0% on Google’s GDM-MRCR v2 test at 128K tokens. It keeps the 1M token window and the 64K output cap, and it carries the same introductory $0.75 and $3.75 pricing as 3.6 Flash through the end of 2026.
One knob matters here. Gemini 3.7 Flash exposes three thinking levels, low, medium, and high, with medium as the default. Low cuts latency for real-time work, high spends more tokens for harder reasoning and tool use. Setting it deliberately is the single cheapest performance lever on this model, and it is the kind of detail that separates an engineer who has shipped on Gemini from one who has only read about it.
Gemini 3.6 Flash
Gemini 3.6 Flash launched on July 21, 2026 and held the top Flash spot for just 23 days. Google built it for coding, knowledge work, and multimodal tasks, and it beats Gemini 3.5 Flash across agentic coding and computer-use benchmarks while using about 17% fewer output tokens for the same job. It has a 1M token context window and now costs $0.75 input and $3.75 output per million tokens, the same introductory rate as 3.7 Flash. Its knowledge cutoff moved up to March 2026, and computer use is built in rather than a separate model call. Since 3.7 Flash costs the same and scores higher across the board, 3.6 is worth keeping only for pipelines you have already tuned against it.
That is the pattern to plan around on Gemini. Google shipped three Flash models in four months, each cheaper or stronger than the last, while Pro sat still. Pin your model IDs, and put a calendar reminder on the December 31 price change rather than discovering it in January’s bill.
Gemini 3.5 Flash
Gemini 3.5 Flash launched in May 2026 and is still live. Google positions it for sustained frontier performance on long coding and agent runs. It has a 1M token context window and costs $1.50 input and $9 output per million tokens, which now makes it the most expensive Flash in the family and the weakest of the three on coding. Unless you have a tuned pipeline you cannot move yet, there is no reason to start anything new on it.
The Flash-Lite tiers
Gemini 3.5 Flash-Lite is the newer of the two cheap tiers, at $0.30 input and $2.50 output per million tokens, and Google aims it at high-throughput work that still needs a little reasoning. Gemini 3.1 Flash-Lite is cheaper again at $0.25 input and $1.50 output. Both run a 1M token context window. Use them for high-volume, simple tasks where speed and cost matter more than depth. Gemini 3 Flash, the original mid-tier of this family, is still live as a preview model at $0.50 input and $3 output, which makes it the cheapest way to reach frontier-class Flash quality if you can accept preview status. Gemini 3.5 Pro is still not out. Google says it is testing with partners after delays on coding quality, and work on Gemini 4 has already started.
The Gemini 2.5 family
The Gemini 2.5 family is older but still supported and widely used. It is a strong fallback when you want lower cost or a stable model you have already tested. Gemini 2.5 Pro offers deep reasoning at $1.25 input and $10 output per million tokens, rising to $2.50 and $15 once a prompt passes 200K tokens. Gemini 2.5 Flash is a low-cost reasoning model at $0.30 input and $2.50 output. Gemini 2.5 Flash-Lite is the cheapest and fastest model on this list, at $0.10 input and $0.40 output per million tokens.
All three have a 1M token context window. For many teams, Gemini 2.5 Flash and Flash-Lite still handle support bots, tagging, and simple extraction at a tiny cost. There is no rush to move every workload to Gemini 3. Move when the new family gives you a clear quality gain on that specific task.
Image, video, audio, and other models
Gemini is more than text. Google ships separate models for other media, billed their own way. The Nano Banana family now handles images, from Nano Banana 2 Lite for fast edits up to Nano Banana Pro for the highest quality, and the older Imagen 4 is deprecated. Veo 3.1 makes cinematic video and is priced per second, with a Lite tier for cheaper runs. Lyria 3 makes music, now split into a Clip model for loops up to 30 seconds at $0.04 a track and a Pro model for full songs at $0.08. Gemini also has in-model image generation through its Flash Image and Pro Image variants, which output pictures inside a normal API call. For video-heavy workflows, Gemini Omni Flash works more like a conversational video model, taking text, images, audio, or sketches as input while keeping characters, physics, and scene context consistent across edits. It reached developers through the Gemini API on June 30, 2026 and is billed at $1.50 per million input tokens and $17.50 per million video output tokens.
There are tool-focused models too. Gemini Embedding 2 turns text, images, audio, and video into vectors for search and retrieval. Computer Use lets an agent click and type in a browser. Gemini Robotics-ER 2 targets robots and physical tasks, and Deep Research, now with a Max tier, runs long multi-step research jobs on its own. Google has also added an Antigravity agent for code, files, and browsing, and a Live Translate model covering 70 or more languages in real time. The original Nano Banana, gemini-2.5-flash-image, shuts down on October 2, 2026, with Nano Banana 2 as its replacement. You do not need most of these for a normal app. But it helps to know they exist so your engineers pick the right tool instead of forcing the text model to do everything.
Pricing compared across the lineup
Output tokens drive most of the cost. Most apps read a prompt and write a longer answer, so output volume adds up fast. The gap is wide. Gemini 2.5 Flash-Lite output is $0.40 per million. Gemini 3.1 Pro output is $12 per million, which is thirty times more. The chart below shows output price for each main model.

Watch the long-context rule on both Pro models. Prompts over 200K tokens jump to a higher rate for input and output alike, on Gemini 3.1 Pro and on 2.5 Pro. If you regularly send huge prompts, that can nearly double your bill without warning. Audio is the other quiet multiplier. It is billed above the headline text rate on several models, at $0.50 per million on 3.1 Flash-Lite against $0.25 for text, and $1 on 2.5 Flash against $0.30. One fix is to route easy requests to Flash-Lite and only escalate hard ones to Pro. Teams that build this routing layer well often cut their model bill by half or more.

Context windows and what they mean
The context window is how much text the model can read at once, counting your prompt and the conversation so far. Every current Gemini model, from 2.5 Flash-Lite up to 3.1 Pro, has the same 1M token window, which is roughly 750,000 words. That is enough to drop a whole codebase or a stack of long documents into one request.
Because the window is now uniform, context size is no longer a reason to pick one Gemini over another. What separates them is the output cap and the recall. Every Gemini 3 model writes at most 64K tokens in one reply, so plan to stream or chunk anything longer. And recall matters as much as size. Gemini 3.7 Flash reads a 128K prompt at 97.0% on Google’s GDM-MRCR v2 recall test, up from 91.8% on 3.6 Flash and 77.3% on 3.5 Flash. That figure falls sharply as you approach the full million on every model, so a tight, well-built prompt still beats a giant one.

Which model should you use
Start with Gemini 3.7 Flash for coding and agent work. It is fast, strong, and while the introductory price holds it is the best value in the family. Drop to Flash-Lite when the task is simple and you run high volume. Move up to Gemini 3.1 Pro when the task needs deeper reasoning or creative range, though no longer for a bigger window, since every model now shares the same 1M. Use Deep Think only for the hardest research-grade problems. The table below maps common jobs to a model.

| Job | Recommended model | Why |
|---|---|---|
| Support chatbot, FAQ | 2.5 Flash-Lite or 3.1 Flash-Lite | Fast and cheap at high volume |
| Data tagging, routing | Gemini 3.1 Flash-Lite | Simple, repeated, low risk |
| Daily coding and code review | Gemini 3.7 Flash | The strongest coding and agent scores in the Flash tier |
| Document summaries, extraction | Gemini 3.7 Flash or 2.5 Flash | 1M context, good accuracy, low cost |
| Deep reasoning, creative range | Gemini 3.1 Pro | Strongest reasoning, though the window matches the rest |
| Hardest research and science | Gemini 3.1 Deep Think | Thinks longest on the toughest problems |

A simple rule helps. If a cheaper model gives the right answer most of the time, use it and add a check. Escalate only the cases that fail. This keeps quality high and cost low. We worked with an analytics team that ran every request on 3.1 Pro. Moving routine checks to Flash and tagging to Flash-Lite cut their bill by more than half with no drop in quality.
Deprecated and older models
Google retires older models over time, and publishes the dates in advance as the earliest a model might be pulled. Gemini 2.0 Flash and 2.0 Flash-Lite shut down on June 1, 2026, the 1.5 family before them is gone, and the first Gemini 3 Pro preview has been retired in favor of 3.1 Pro. Preview models get only two weeks of notice, which is worth knowing before you build on gemini-3.1-pro-preview. If your code still calls one of these, plan a move. A switch is usually small. Change the model name, then test the output, since the newer models can format and reason a little differently.
| Older model | Status | Move to |
|---|---|---|
| Gemini 2.0 Flash | Shut down June 1, 2026 | Gemini 3.7 Flash |
| Gemini 2.0 Flash-Lite | Shut down June 1, 2026 | Gemini 3.1 Flash-Lite |
| Gemini 1.5 Pro, 1.5 Flash | Retired | Gemini 2.5 Pro or 3.1 Pro |
| Gemini 3 Pro Preview | Shut down | Gemini 3.1 Pro |
| Nano Banana (gemini-2.5-flash-image) | Shuts down October 2, 2026 | Nano Banana 2 |
| Imagen 4 | Deprecated | Nano Banana 2 or Nano Banana Pro |
| Veo 2.0, Veo 3.0 | Shut down June 30, 2026 | Veo 3.1 or Veo 3.1 Lite |
Test before you trust the swap. A newer model may change how it returns JSON or how long its answers run. Check your parsing and your token limits after any move. Google’s model and pricing pages list every active model in one place, and the deprecations page carries the shutdown dates, so confirm the current names before you migrate.

What this means for hiring
Knowing the lineup is step one. Getting value from it is the harder step. The teams that win are not the ones using the biggest model. They route work to the right tier, cache context, set thinking levels on purpose, watch the long-context rule, and ship a clean product around it. That is engineering, not just a model choice.
Strong AI engineers are scarce and costly in the United States. A senior AI engineer there runs $12,000 to $18,000 per month all in. The same skill from a pre-vetted engineer in Vietnam or the Philippines often costs $3,000 to $6,000 per month. The work is the same. The bill is not. Many of our clients build their whole Gemini stack with engineers from our network.
If you also use other models, it helps to compare. Our guides on Claude versus ChatGPT and Claude versus Gemini for coding shows how to judge models on real developer work, not just benchmarks. Read it next to round out your view before you commit a workload to one provider.
Compare other model families
Build your Gemini stack with the right team
Gemini gives you two families and a clear set of tiers. The model is the easy part. The team that builds with it is what ships value. We match you with pre-vetted senior engineers across APAC in about 24 hours, with no upfront cost and payroll handled.



![Singapore AI Companies Leading Southeast Asia. Top 10 Singapore AI Companies Leading Southeast Asia [2026], by Second Talent.](https://www.secondtalent.com/wp-content/uploads/2026/09/singapore-ai-companies-featured-v2-768x403.jpg)
![AI Recruiting Tools. Top 7 AI Recruiting Tools in 2026 [Tried & Tested], by Second Talent.](https://www.secondtalent.com/wp-content/uploads/2026/09/top-ai-recruiting-tools-featured-v2-768x403.jpg)
