Gemini 3.1 Pro vs 3.5 Flash: Full Comparison, Benchmarks And Pricing

/ Google’s AI Flagships Face Off in 2026

Published: June 3, 2026 at 2:00 PM EDT | Updated: September 4, 2026 at 2:53 AM EDT
Gemini 3.1 pro vs gemini 3.5 Flash - TheTweaks
Image: Alison Parker / TheTweaks

Gemini 3.5 Flash is a surprising AI model. It has outperformed the flagship model (Gemini 3.1 Pro). According to this, does that mean that Gemini 3.1 Pro no longer has a place in serious workloads or does it still have an edge in certain scenarios or conditions.

At Google I/O 2026, Google announced Gemini 3.5 Flash, it is a faster and cheaper model than Gemini 3.1 Pro, based on many benchmarks in the industry. This claim is not based primarily on marketing. It also includes BenchLM’s independent rankings, Google’s official evaluation and real world testing by enterprise partners like Box and JetBrains and all support the same conclusion.

There is no guarantee that better performance according to their benchmarks will automatically mean the same for everyone. It all depends on how one uses the language model. This thorough comparison of Gemini Pro 3.1 with Gemini Flash 3.5 is collected using information from Google’s official documentation on the model architecture, BenchLM’s leaderboards, NxCode’s five workload analysis and MindStudio’s actual testing. 

What are Gemini Pro 3.1 and Gemini Flash 3.5 ?

It is important to understand what each model is designed to do before jump into benchmarks. This context will influence the numbers that follow.

Gemini 3.1 Pro

Google released the Gemini 3.1 Pro earlier than 2026, making it Google’s most capable model at that time. The design philosophy of the Gemini 3.1 Pro is based on deep reasoning, multi step logic and long document form. It also focuses on high stakes precision, where getting an answer wrong can have real consequences. The Gemini 3.1 Pro is available via Google AI Studio and Vertex AI and targets enterprise workflows such as law, medicine and finance.

Gemini 3.1 Pro’s benchmarks reflect this design intention. It consistently tops abstract reasoning tests and long context retrieval.

Gemini 3.1 pro - TheTweaks
Image: Gemini 3.1 Pro / Google

Gemini 3.5 Flash

Google I/O announced that the release date for Gemini 3.5 Flash was May 19, 2026. Google DeepMind describes its Flash models as best in frontier performance for agents and coding. This is a major shift from the description Flash models had even a generation ago. Gemini 3.5 Flash is now available in Google AI Studio and the Gemini App. It’s also the default model on Gemini.google.com.

Flash models have traditionally traded reasoning depth for performance. Gemini 3.5 Flash breaks this pattern. It is 2x faster than 3.1 Pro in 8 out of 12 major benchmark categories and generates approximately 2x as many output tokens per second. Gemini 3.5’s 284 tokens/second speed advantage over Pro’s 109, is not marginal. It changes the possibilities of real time applications.

Gemini 3.1 Pro vs 3.5 Flash: Full Benchmark Comparison

This table collects the data from Google DeepMind’s official model page and BenchLM’s independent testing. These are the benchmarks that matter the most for production decisions.

Category Benchmark Gemini 3.5 Flash Gemini3.1 Pro Winner

Coding

Terminal Bench 2.1 76.2% 70.3% Flash

Coding

SWE-Bench Pro (Public) 55.1% 54.2%

Flash

Agentic

MCP Atlas 83.6% 78.2%

Flash

Agentic

Toolathlon 56.5% 49.4%

Flash

Expert Tasks

Finance Agent v2 57.9% 43.0%

Flash

Multimodal

CharXiv 84.2% 83.3%

Flash

Multimodal

MMMU-Pro 83.6% 80.5%

Flash

Long Context

MRCR v2 @ 128k 77.3% 84.9%

Pro

Reasoning

ARC-AGI-2 72.1% 77.1% Pro

Reasoning

Humanity’s Last Exam 40.2% 44.4% Pro

Overall (BenchLM)

Aggregate Score 87 92 Pro

Flash is the winner in 8 out of 12 categories. Pro is the leader in reasoning depth and in long context data. BenchLM’s aggregate score favors Pro by 92 points to 87, but this headline number masks the dominance of Flash in categories that are key for most AI applications.

The largest delta is the 342 point Elo difference on GDPval-AA, which measures the economic value of real world knowledge. This is not a close race.

Gemini 3.5 Benchmarks – Where it Dominates

Finance Agent v2 is the most impressive result of the benchmark table. Flash scored 57.9 per cent against Pro’s 43.0 per cent. The 14.9 point difference in financial analysis and decisions is significant, and signals that Flash’s post training did not just improve performance but also the multi step tool using workflows which business applications depend on.

Flash is ahead 83.6% versus 78.2% on MCP Atlas (multistep workflows utilizing Model Context Protocol). Toolathlon (real world general tool usage) is where Flash leads with 56.5%. These are not narrow margins, but they do reflect a model which was clearly trained using agent pipelines as an example of first class usage.

Box’s enterprise evaluation set, designed to reflect real world multi step tasks that our customers do every day, showed that Gemini 3.5 Flash was 19.6% faster than Gemini 3 Flash. It extracts data more accurately with 96.4% for Life Sciences customers. Ben Kus, CTO, Box

Gemini 3.5’s Flash has a speed advantage that is not captured by benchmarks for single calls. Flash can complete an eight step agent’s loop in about 25 seconds. Pro takes about 100 seconds. This difference is what determines if the interactive developer tools, or applications that are used by customers feel responsive or slow.

Gemini 3.5 flash also manages impressive multimodal input volume: up to 3,000 images per prompt; 45 minutes of video (with audio) or one hour without audio and up to 8.4 hours audio per prompt. Flash now has the same input capacity as Pro for teams creating pipelines with a high visual component or workflows that require document understanding.

Gemini 3.5 Flash - TheTweaks
Image: Gemini 3.5 Flash / Gemini

Gemini 3.1 Pro Benchmarks: Where It Still Leads

Gemini 3.5 Flash is better in various benchmarks, Gemini 3.1 Pro still performs better with long context retrieval and complex reasoning. When accuracy is more important than speed, as in the extraction of specific information from a large contract, MRCR v2 (128K context) scored 84.9% as opposed to Flash’s 77.3%.

Pro also has an edge in the pure reasoning tasks. It scored 77.1% and Flash’s 72.1% on ARC-AGI-2 and 44.4% and 40% respectively on Humanity’s Last Exam. The results point to the fact that with Flash, the ceiling of the complex reasoning and abstract problem solving is lower than with Pro, but that the use of tools and execution speed is better optimized for the former.

Both models, however, fail to get any significant lift with the full 1M-token context, with scores around 26% on MRCR v2. Whether you use a specific Gemini model or not, for workloads that have very large collections of documents, you still need a dedicated retrieval system.

Comparison Of Pricing: Google Gemini 3.1 Pro Vs Flash 3.5

The pricing of Gemini 3.5 Flash is completely different from that of Gemini 3.1 Pro. All in all:

Cost Factor

Gemini 3.5 Flash Gemini 3.1 Pro
Input (per 1,000,000 tokens) $1.50

$2.00

Output (per 1M tokens)

$9.00

$12.00

Cached input

$0.15/1Million

$0.50/1M

Speed

284 tokens/sec

109 tokens/sec

Latency (TTFT)

18.55 sec

29.71sec

Max output tokens

65,536

32,768

Context window

1M tokens

1M tokens

The disruptive aspect of Gemini 3.5 Flash pricing comes in when the input price is cached. Agent pipelines with stable system prompts are much lesser at $0.15 per million tokens than Pro’s RAG (Retrieval Augmented Generation) systems at $0.50 per million tokens. For a high frequency, 10,000/month knowledge base application, the price tag would amount to about $1,800/month on Pro and about $200/month with caching turned on for Flash. That’s a difference of six figures for a product of a reasonable size after one year.

A surprising feature: Most people are surprised to find that the maximum output tokens for Flash is 65,536, which is twice that of Pro’s 32,768. For making long reports, detailed code files or multiple document summaries, Flash is more than just the budget friendlier choice. 

5 Real-World Workloads: Which Model Wins Each

Instead of vague rules, here’s how the data equates to five specific decisions.

MCP-based agent workflows Use Gemini 3.5 Flash. This is a clear call as it tops the 5.4-point chart on MCP Atlas and 7.1-point chart on Toolathlon and features a 4× speed advantage in multi-step loops. There is no situation where Pro can go to this endgame and win just based on performance.

For long document retrieval (100k+ tokens), use Gemini 3.1 Pro. The gap of 7.6 points at MRCR v2 is consistent and relevant for legal review, compliance and contract extraction. If you need to upload the full document and then ask questions, Pro is more reliable with chunked retrieval with pre-processing is not an effective method.

Terminal coding agents Use Gemini 3.5 Flash. It leads Terminal-Bench 2.1 by 5.9 points and Blueprint-Bench 2 by 7.1 points. SWE-Bench is effectively tied (55.1% vs 54.2%). If there is no difference in the number of code fixes, then the lower latency and cost of Flash is its deciding factor.

Stable corpus: use Gemini 3.5 Flash with aggressive caching for high volume RAG. This is the lowest price to cache a single system prompt ($0.15/1M) so it is the most cost efficient for any application where the system prompt is consistent throughout the requests.

Abstract reasoning and research tasks Use Gemini 3.1 Pro. If the workload is one that requires novel deduction, pattern recognition in an unseen domain, or expert scientific reasoning, the leads on ARC-AGI-2 and Humanity’s Last Exam are consistent and truly represent a ceiling difference.

Facing all this information, the comparison between Gemini 3.1 Pro and 3.5 Flash just comes back to one question: What failure mode do you really have the resources to tolerate?

Flash is faster, cheaper and stronger when most applications are used most of the time. It has more output token limit. It has a significantly lower price, which it caches. It is not exactly a step forward over Pro with its GDPval-AA lead or the MCP Atlas win, and its Finance Agent result isn’t really an improvement. It’s another level of performance for that type of work.

Pro is still worth his price, however. This loss of 7.6 points at 128k tokens is real. Logically, the reasonings in the two arguments ARC-AGI-2 and Humanity’s Last Exam are consistent. In jobs that do matter and for serious pursuits such as law, medicine or research, they really do, and that’s a worthwhile investment of money.

The best production environment is when both are used: Flash for the steps that are clearly definable and in high volume, Pro for the decision points that involve reasoning or memory retrieval.

My Honest Take

Facing all this information, the comparison between Gemini 3.1 Pro and 3.5 Flash just comes back to one question: What failure mode do you really have the resources to tolerate?

Flash is faster, cheaper and stronger when most applications are used most of the time. It has more output token limit. It has a significantly lower price, which it caches. It is not exactly a step forward over Pro with its GDPval-AA lead or the MCP Atlas win, and its Finance Agent result isn’t really an improvement. It’s another level of performance for that type of work.

Pro is still worth his price, however. This loss of 7.6 points at 128k tokens is real. Logically, the reasonings in the two arguments ARC-AGI-2 and Humanity’s Last Exam are consistent. In jobs that do matter and for serious pursuits such as law, medicine or research, they really do, and that’s a worthwhile investment of money.

The best production environment is when both are used: Flash for the steps that are clearly definable and in high volume, Pro for the decision points that involve reasoning or memory retrieval.

Final Verdict

For building agent workflows, coding assistant, multimodal pipelines or high volume RAG systems, go for Gemini 3.5 Flash. It outperforms Pro in 8 of 12 benchmark categories, is 2.6 times faster, is much cheaper, and is able to output for a longer time. Flash is the default for most of the 2026 production builds.

Use Gemini 3.1 Pro for tasks that involve more than 100k+ tokens of a document, abstract reasoning, or tasks that require high accuracy like in legal or financial fields, or scientific research.

Always use both where possible. Throughput = Flash; Judgement = Pro. That hybrid method offers the best unit economics, and the best coverage, as per the data, these two models are meant to go hand-in-hand.

Frequently Asked Questions

Yes, Flash wins 8 of 12 benchmark categories including all agentic, tool-use, and multimodal tests. It runs 2.6× faster and costs significantly less per token, making it the stronger default choice for most modern AI applications and pipelines in 2026.
Gemini 3.5 Flash was released on May 19–20, 2026, at Google I/O. The Gemini 3.5 Flash API is available immediately through Google AI Studio, Vertex AI, the Gemini App, and Google Antigravity using the model string gemini-3.5-flash-preview.
Use Gemini 3.1 Pro for long-document retrieval over 100k tokens, abstract reasoning tasks like ARC-AGI-2 puzzles, and high-stakes professional workflows in law, medicine, or finance where Pro's measurable accuracy advantage directly reduces the cost of model errors.
Most Related