Gemini 3.7 Flash Review — Google's Cheap Agent Workhorse
On August 13, 2026, Google shipped Gemini 3.7 Flash, the successor to Gemini 3.6 Flash and the clearest statement yet of a strategy Google has been circling all year: make the cheap model the one you actually build agents on. The pitch is not a new frontier ceiling — the Pro line still owns that — but a workhorse that lands the agentic-coding tier of models costing two to three times as much, at an introductory $0.75 per million input tokens. This review covers what Gemini 3.7 Flash actually scores, what the pricing giveth and the calendar taketh away, and where it fits against the GPT-5.6 Terra and Luna tiers it is aimed squarely at.
TL;DR verdict
| Gemini 3.7 Flash | |
|---|---|
| Type | Cheap, high-throughput multimodal model |
| Released | August 13, 2026 (Google DeepMind) |
| Category peers | GPT-5.6 Terra, GPT-5.6 Luna, Gemini 3.6 Flash |
| Context window | 1,048,576 tokens (~65K max output) |
| Modalities | Text, image, audio, video in; text out |
| Introductory pricing | $0.75 in / $3.75 out per 1M (through 2026-12-31) |
| Scheduled pricing | $1.50 in / $7.50 out per 1M (from 2027-01-01) |
| Headline numbers | GPQA Diamond 90.4, Terminal-Bench 2.1 85.8, MMMU-Pro 81.2 (Google-reported) |
| Best for | High-volume agentic and coding workloads on a budget |
| Caveat | Introductory price doubles in January; launch numbers are Google-run |
If you read no further: Gemini 3.7 Flash is the strongest price-to-capability move of the month, and for high-volume agentic workflows it is an easy model to default to today. The rating stops short of the top because the number that makes it exciting — $0.75 input — is temporary, and the benchmarks that make it exciting are Google-reported and not yet on independent boards.
What Google actually reported
Gemini 3.7 Flash keeps the Flash line’s defining trait — a 1,048,576-token context window with full multimodal input (text, image, audio, and video) — and pushes the numbers up across the board. The figures Google led with at launch:
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash | GPT-5.6 Terra | Notes |
|---|---|---|---|---|
| GPQA Diamond | 90.4 | — | 87.0 | Google-reported; science reasoning |
| Terminal-Bench 2.1 | 85.8 | 78.0 | 87.4 | Google-reported; agentic terminal use |
| MMMU-Pro | 81.2 | — | — | Google-reported; multimodal reasoning |
| SWE-bench Pro | ~62 (est.) | 58.7 | 63.4 | Not published by Google; leaderboard estimate |
Two readings fall out of that table. First, on the benchmarks Google did publish, 3.7 Flash is genuinely in the agentic-coding tier: its 85.8 on Terminal-Bench 2.1 is a large jump from 3.6 Flash’s 78.0 and sits just behind GPT-5.6 Terra’s 87.4 — a model priced at $2.50 input, more than three times the introductory Flash rate. Google’s own coding-focused deltas tell the same story: FrontierCode 1.1 rose from 34.4% to 43.6% and DeepSWE v1.1 from 48.6% to 65.3% between 3.6 and 3.7 Flash.
Second, note what is missing. Google did not publish a SWE-bench Pro number, which is the memorisation-resistant coding benchmark we lean on most. On our leaderboard we therefore store SWE-bench Pro as an explicit conservative estimate (~62, anchored to 3.6 Flash’s 58.7 and the reported coding gains), greyed out and excluded from every ranking — rather than laundering a guess into a measurement. The GPQA, Terminal-Bench, and MMMU-Pro cells are stored as vendor-reported, with the caveat that they were run on Google’s own harness. For how much weight self-reported coding numbers deserve, the LLM Benchmark Comparison 2026 lays out the hierarchy.
One more calibration point worth keeping in view: on the newer, much harder Terminal-Bench 3.0, Gemini 3.7 Flash scores in the mid-teens — roughly level with Claude Sonnet 5 and below GPT-5.6 Terra. The saturation on 2.1 is real, and 3.0 is where the daylight between models still lives. A Flash-tier model clearing 85 on 2.1 is a statement about how fast the floor is rising, not a claim that the hardest agentic work is solved.
The pricing — and the January catch
This is the part to read slowly, because the headline and the reality diverge on a fixed date.
Introductory pricing (through December 31, 2026):
- Input: $0.75 per 1M tokens
- Output: $3.75 per 1M tokens
Scheduled pricing (from January 1, 2027):
- Input: $1.50 per 1M tokens
- Output: $7.50 per 1M tokens
At the introductory rate, Gemini 3.7 Flash is half the input price of Gemini 3.6 Flash ($1.50 in / $7.50 out) while scoring higher on every benchmark Google published — a rare strict improvement on both axes. It also undercuts GPT-5.6 Luna ($1.00 in / $6.00 out) and sits far below GPT-5.6 Terra ($2.50 in / $15.00 out), both of which it trades blows with on agentic evals. For a high-volume agent pipeline, that is a compelling default right now.
The catch is the calendar. On January 1, 2027 the price doubles back to exactly where Gemini 3.6 Flash sits today. That does not make it a bad deal — at $1.50 in / $7.50 out it is still competitive with its own predecessor at better scores — but it does mean the “cheaper than Luna” framing has a five-month expiry. If you are sizing a budget for a workload that ships in Q1 2027, model the $1.50/$7.50 numbers, not the launch-week ones. Prompt caching narrows the gap further for repeated-context workloads; the leaderboard’s cost calculator will run either price against your own token mix.
How it stacks up
Against its own predecessor the verdict is simple: Gemini 3.7 Flash strictly dominates Gemini 3.6 Flash — higher on every published benchmark, half the introductory input price, same 1M multimodal context. If you are on 3.6 Flash, this is an upgrade with no obvious downside beyond re-validating your prompts.
Against the OpenAI mid-tier the picture is a genuine contest. GPT-5.6 Terra still edges it on Terminal-Bench 2.1 (87.4 vs 85.8) and has a published SWE-bench Pro number (63.4) where Flash has only an estimate — but it costs more than three times as much per input token at the introductory rate. GPT-5.6 Luna is the closer price comparison, and here Flash’s published science and multimodal numbers plus the cheaper introductory input make it the more interesting pick for vision-heavy or long-context agent work. As with the Muse Spark 1.2 launch, the honest move is to weight the self-reported coding figure as directional until a neutral harness confirms it, then run your own bake-off.
Who should care
- Teams running high-volume agent pipelines: This is the headline audience. At $0.75 input, a 1M-context model landing the agentic tier is hard to argue with — just budget for the January price change if your workload outlives 2026.
- Anyone on Gemini 3.6 Flash: Upgrade. It is cheaper and better on every number Google published; the only cost is re-testing your prompts.
- Vision- and video-heavy applications: The 81.2 MMMU-Pro and full multimodal input make it a strong cheap option where GPT-5.6 Luna’s text-first profile is a worse fit.
- Benchmark-driven buyers: Add it to your suite, but treat the launch numbers as Google-run until Artificial Analysis and the public boards weigh in, and compare against the leaderboard rather than the launch blog.
FAQ
When was Gemini 3.7 Flash released? Google shipped it on August 13, 2026, as the successor to Gemini 3.6 Flash.
How much does it cost? Introductory pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026. Google has scheduled rates of $1.50 input and $7.50 output from January 1, 2027 — double the introductory rate.
How good is it at coding? Google reports Terminal-Bench 2.1 of 85.8%, up from 3.6 Flash’s 78.0% and just behind GPT-5.6 Terra’s 87.4%. Google did not publish a SWE-bench Pro figure, so our leaderboard carries an estimate for that cell.
Are the benchmarks verified? No — GPQA Diamond 90.4, Terminal-Bench 2.1 85.8, and MMMU-Pro 81.2 are Google-reported on Google’s own harness. Validate on your own tasks and watch the independent boards.
Is it multimodal? Yes. It accepts text, image, audio, and video input and returns text, with a 1,048,576-token context window.
Continue reading
- AI Models Leaderboard — Gemini 3.7 Flash versus 75+ models on benchmarks, pricing, context window, and throughput, with every figure sourced.
- GPT-5.6 Sol, Terra & Luna Review — the OpenAI mid-tier that Gemini 3.7 Flash is priced and benchmarked against.
- Meta Muse Code & Muse Spark 1.2 Review — another workhorse-tier launch, and the same lesson on reading self-reported coding numbers.
- LLM Benchmark Comparison 2026 — how to weigh Terminal-Bench, SWE-bench, and vendor-run figures.
- All Reviews — index of every head-to-head review on the site.