
Gemini 3.8 Flash launched Sept 2, 2026 at $0.75/$3.75 per 1M tokens, with a 1M context window. Here's what's new, the hidden token costs, and who should actually switch.
Vamsi Tallapudi
Manager, Architect Technology at Cognizant
Google shipped Gemini 3.8 Flash on September 2, 2026, its third Flash-tier release in six weeks. It's not a new model built from scratch. Google's own docs say flatly that "3.8 Flash is based on 3.7 Flash" — this is a tuning pass that makes the existing model work harder on coding, multi-step agent tasks, and specialized reasoning, while keeping the same speed and the same price tag as its predecessor.
That last part is the headline Google wants you to remember. Whether it's the whole story is a different question, and I'll get to that.
Google's framing is that 3.8 Flash "works harder." On complex problems it runs more reasoning steps and calls tools more iteratively before answering, instead of just predicting the next token faster. That's a real behavior change, not a speed upgrade.
The results show up in benchmarks Google is happy to publish:
What Google didn't put in a keynote slide: on Terminal-Bench 4.0, a benchmark for agentic terminal use, 3.8 Flash scores 19.1% against Opus 5's 51.8%. That's not a close race. So "near-Pro coding" is true in some benchmarks and badly wrong in others, depending on the task.
The feature set is wide, and it's genuinely useful if you're building anything agent-shaped:
Notice what's missing. There's no image generation, no audio output, no Live API here. If you want Gemini to talk back to you in real time, that's a separate product — Gemini 3.8 Live and Live Extended Thinking, which Google shipped two weeks later and which I wrote about here. Same version number, completely different model, which is a confusing naming choice on Google's part.
Here's where the pricing story gets more complicated than "same price as 3.7 Flash."
| Input | Output | |
|---|---|---|
| Gemini 3.8 Flash (intro, through Dec 31, 2026) | $0.75 / 1M tokens | $3.75 / 1M tokens |
| Gemini 3.8 Flash (2027 standard rate) | $1.50 / 1M tokens | $7.50 / 1M tokens |
| Cache reads | $0.075 / 1M tokens | — |
The per-token price matches Gemini 3.7 Flash exactly. But output pricing includes the model's thinking tokens, which you never see in the response. Independent benchmarking from Artificial Analysis found 3.8 Flash burns about 70% more output tokens than the median model to finish the same task. Same price per token, more tokens per task — your actual bill doesn't necessarily go down just because the rate card looks unchanged.
Google is upfront about this trade-off, to its credit. The official guidance for anyone who cares about cost or speed over raw capability: stick with Gemini 3.7 Flash, or drop the thinking level, because 3.8 Flash will cost more tokens to do the same job.
This is the part that bugs me about the "speedy" framing. Flash models exist for low latency, and on raw output speed 3.8 Flash still does well — Artificial Analysis clocked it at 302 tokens/sec, third-fastest among the models it tracks.
But time to first token is a different story: 13.3 seconds, about 4.5x slower than the median model. That's the gap between when you hit send and when the first word shows up. For a chatbot, a support widget, or anything a human is staring at waiting for a reply, 13 seconds feels broken. For a background coding agent that's going to run for two minutes anyway, nobody notices.
So "Flash" here describes the price tier and the token-generation speed, not the experience of waiting for a reply. If you're building something interactive, test this number yourself before you commit to it.
Alongside the main model, Google released Gemini 3.8 Flash Cyber, a variant tuned specifically for vulnerability discovery and patching, with "deliberately looser" safety mitigations than the public model. It's not something you can just sign up for.
Access runs through Google DeepMind's Fairwind Program, limited to vetted defenders — government agencies, critical infrastructure operators, and software maintainers. The results Google is showcasing are genuinely strong: the Chrome Security team reported 2.6x more correct patches than the best commercial models they tested, and the model hits 47.2% pass@1 on CWE-Bench for automated patching, roughly matching larger frontier models at a fraction of the cost.
If you're not already a vetted defender, you won't be running this one. It exists to find bugs before attackers do, not to ship inside your product.
Gemini 3.8 Flash rolled out across most of Google's surfaces on day one:
gemini-3.8-flashThat's about as wide a launch as Google does for a Flash-tier model, and it means most people reading this already have access without changing anything, assuming you're already paying for Gemini AI Pro or Ultra.
My honest take, after reading through the benchmarks and the model card: switch for agentic coding and document-heavy, multi-step work. Stay on 3.7 Flash for anything latency-sensitive or high-volume, like a support chatbot or a live search feature, where a 13-second pause kills the experience and the token verbosity actually raises your bill.
This is also worth saying plainly: Gemini 3.8 Flash is a tuning pass, not a new architecture. If you've already built a workflow around 3.7 Flash and it's working, I wouldn't rush the migration just because a bigger number showed up in the model picker. Test it on your actual workload first.
A few things got buried under the launch excitement:
None of these are dealbreakers on their own. Together, they're a good reminder that "upgraded workhorse model" doesn't mean "upgrade on every axis." Read the model card, not just the announcement post, before you move production traffic.
If you're also watching what Anthropic shipped around the same window, our Claude Opus 5 pricing breakdown covers the other side of this race. And if you want the wider context on why Google is iterating on Flash this fast, our recap of every AI model that launched in August 2026 has the full timeline, including Gemini 3.7 Flash itself.
Google's cadence here is the real story. Three Flash releases in six weeks (3.6 in July, 3.7 in August, 3.8 in September) tells you Google is iterating faster than it's innovating, shipping tuning passes between the actual architectural jumps. That's not a bad strategy. It just means the next "big" Gemini release is probably further out than the version number suggests.
New AI tools, automation workflows, and course drops — straight to your inbox. Join 2,400+ builders.

Google launched Nano Banana 2.1 on October 6, 2026, cutting image API prices in half to $0.0336 per 1K image while improving text rendering, editing, and character consistency. Here's what actually changed.

A solo dev built PaperRoute, a finished Paperboy-style browser game, with GPT-6 Astra, Blender and Meshy in 39 tracked hours and 1.56B tokens. Here's the 6-step process he used.

Anthropic is merging Claude Cowork and Claude chat into one app, rolling out to Pro and Max subscribers over the next few weeks.