Skip to content
ProductModel ReleaseAI & LLMs

Gemini 3.6 Flash Underwhelming

Gargeya SharmaFounder & Architect
July 21, 2026
5 min read
Gemini 3.6 Flash Underwhelming

Key Takeaways

  • Gemini 3.6 Flash improves token efficiency and lowers output cost but stays pricier than rivals.

  • Grok 4.5 and GPT-5.6 Luna outperform Gemini on SWE‑Bench, DeepSWE, and Terminal‑bench.

  • Fast inference times don’t offset higher per‑task cost for developers.

  • Only Google‑ecosystem users get a free upgrade; others see limited value.

Gemini 3.6 Flash dropped today. I’m still underwhelmed.

I spent a good chunk of the day reading the announcement, the charts, the pricing sheet, and the usual wave of “this is the new workhorse” posts. Google’s new Flash model is cleaner and more efficient than 3.5 Flash. That part is true. But the gap between the marketing energy and the actual competitive picture is wide enough that I keep coming back to the same feeling: this doesn’t move the needle the way it should.

What Google is selling

Gemini 3.6 Flash is the updated workhorse. Same $1.50 input price as 3.5 Flash, output cut from $9.00 to $7.50. The big claim (backed by Artificial Analysis) is 17% fewer output tokens on multi-step work. Less verbosity, faster completion, lower bills on agentic loops.

Here’s the token efficiency side-by-side they published:

On DeepSWE the drop is sharp. On the broader Artificial Analysis Index tasks it is more modest but still real. They also showed clear gains on agentic benchmarks versus their own previous generation:

So yes — better than 3.5 Flash. More efficient, slightly stronger on coding and computer-use tasks, strong long-context numbers. For anyone already locked into the Google ecosystem this is a free upgrade. No complaints there.

The comparison that actually matters

The problem is that “better than our last Flash” is not the bar anymore. The models that sit in the same practical price band for daily coding and agent work are Grok 4.5 and GPT-5.6 Luna. Those are the ones that should be getting the attention in any honest comparison.

Here’s the table restructured so the real competitors stand out:

Here is the data formatted as a Markdown table:

MetricGemini 3.6 FlashGrok 4.5GPT-5.6 LunaGemini 3.5 Flash
Input / Output ($/1M)$1.50 / $7.50$2.00 / $6.00$1.00 / $6.00$1.50 / $9.00
SWE-Bench Pro58.7%64.7%62.7%55.1%
DeepSWE v1.149%54%67%37%
Terminal-bench 2.178.0%83.3%84.7%76.2%
OSWorld-Verified83.0%72.6%78.4%
GDPVal-AA v2 (Elo)1421153515841349
CharXiv (no tools)85.2%81.6%82.7%84.2%

Look at the output price column. Grok 4.5 and Luna are both at $6.00. Gemini is still at $7.50 even after the cut. On the heavy coding and agentic benches that actually matter for building software (SWE-Bench Pro, DeepSWE, Terminal-bench), both competitors are ahead. Gemini only pulls ahead cleanly on OSWorld and some long-context/multimodal work.

That is the core of the underwhelmed feeling. Google is celebrating a 17% token reduction and a $1.50 output price drop while the models people are actually choosing for the same workloads are already cheaper on output and stronger on the exact tasks most of us care about.

Artificial Analysis view

Artificial Analysis already has 3.6 Flash on the board. Intelligence Index sits around the same 50 zone as 3.5 Flash. Grok 4.5 is higher (around 54). Luna is right there with it in the low 50s, while the heavier GPT-5.6 Sol variants sit higher still.

The interesting part on AA is speed and time-per-task. Gemini Flash models continue to be extremely fast:

3.6 Flash finishes Intelligence Index tasks in roughly 1.3 minutes — one of the quickest measured. That matches the token-efficiency story. Less chatter means less wall-clock time.

But when you look at the broader intelligence-versus-cost picture, the most attractive quadrant is still occupied by models that deliver solid intelligence at meaningfully lower cost per task. Gemini 3.6 Flash is competent and fast, not exceptional on value.

Intelligence Index vs  Cost per Intelligence Index Task (21 Jul '26)

Cost per Intelligence Index Task (21 Jul '26)

Output Tokens per Intelligence Index Task (21 Jul '26)

Time per Intelligence Index Task (21 Jul '26)

Why the emotion is there

I wanted to feel more excited. A new Flash model with better efficiency should be easy to celebrate. Instead I felt the same mild frustration I had earlier in the day: the announcement energy is high, the relative gains over 3.5 Flash are real, and yet the competitive picture barely moves.

For a solo builder who actually pays the bills and cares about how many tokens get burned on multi-step coding and agent loops, the question is simple. Why switch to a model that is still more expensive on output and weaker on several of the core coding benches when Grok 4.5 and GPT-5.6 Luna are sitting right there?

The 17% token cut helps. The lower output price helps. Neither one is enough to make 3.6 Flash the obvious default for people who are not already inside Google’s stack. That is the gap between the charts Google wants you to look at and the comparison that actually decides daily usage.

I’ll still test it on my own workloads this week. Efficiency gains are always welcome. But right now the story feels incremental in a market that has already moved past incremental.

Finished reading?

Connect with Gargeya Sharma on digital strategy and autonomous pipelines.

All Broadcast Articles