Release brief · July 22, 2026
Gemini 3.6 Flash lands with a thud: twice as fast, no smarter than 3.5
Gemini 3.6 Flash is twice as fast as 3.5, but no smarter. With GPT-5.6 Luna scoring higher for less money, the community reaction turned hostile.

Gemini 3.6 Flash is a serious efficiency upgrade disguised as a disappointing model launch. It cuts task time in half and ranks first for output speed, but independent testing shows no intelligence gain over 3.5 Flash. For buyers comparing raw capability per dollar, GPT-5.6 Luna currently makes Google's new workhorse look awkward.
Google launched Gemini 3.6 Flash as its new workhorse for coding, multimodal analysis, computer use, and production AI agents. The first independent benchmark verdict was brutal: it is not any smarter than the model it replaces.
On the updated Artificial Analysis Intelligence Index, Gemini 3.6 Flash at high reasoning scored 50, the same as Gemini 3.5 Flash. It placed #21 out of 186 models in the comparison class, one point behind GPT-5.6 Luna at maximum reasoning.
The ranking alone would not be fatal. A score of 50 is still well above the 31-point average for comparable models. The problem is the value comparison. Gemini costs $1.50 per million input tokens and $7.50 per million output tokens. GPT-5.6 Luna costs $1.00 and $6.00 while scoring 51.
Google released a new Flash model that is slightly less intelligent and more expensive than OpenAI’s competing lightweight model. The community saw the chart and immediately asked the obvious question: what exactly is the upgrade?
The launch chart and the leaderboard tell different stories
Google’s launch announcement presents 3.6 Flash as a clear generational step. Its selected benchmarks show gains over 3.5 Flash in long-horizon software engineering, machine-learning engineering, knowledge work, and computer use.

Those gains are real within the reported evaluations:
- DeepSWE: 49% for 3.6 Flash, up from 37% for 3.5 Flash.
- MLE-Bench: 63.9%, up from 49.7%.
- GDPval-AA v2: 1421, up from 1349.
- OSWorld-Verified: 83.0%, up from 78.4%.
But a broad composite index asks a different question. Artificial Analysis combines nine evaluations spanning agentic work, terminal use, coding, scientific reasoning, knowledge reliability, and long-context reasoning. In that wider test, the improvements were balanced by flat or weaker results elsewhere. GDPval-AA improved, while Humanity’s Last Exam slipped three percentage points to 38%. The total score stayed at 50.
This is the source of the launch-day disconnect. Google can truthfully say 3.6 Flash is better at several important production workloads. Critics can truthfully say it delivered zero overall intelligence gain on a major independent index.
How Gemini 3.6 Flash actually ranks
Artificial Analysis’ July 22 snapshot gives the model a split personality:
| Metric | Gemini 3.6 Flash (high) | GPT-5.6 Luna (max) | What it means |
|---|---|---|---|
| Intelligence Index | 50, #21 | 51, #16 | Luna has the small capability lead |
| Output speed | 303.6 tokens/s, #1 | 189.7 tokens/s, #12 | Gemini is roughly 60% faster once generation begins |
| Input price | $1.50 / 1M | $1.00 / 1M | Gemini charges 50% more |
| Output price | $7.50 / 1M | $6.00 / 1M | Gemini charges 25% more |
| Context window | 1M tokens | 1M tokens | No advantage either way |
| Index output tokens | 59M | 130M | Gemini is dramatically less verbose |
The headline criticism is therefore directionally fair but incomplete. Luna is one point smarter and has lower token prices. Gemini, however, is much faster and used less than half as many output tokens across the index run. Depending on the workload, that concision can reduce total cost and recovery time even when the price per token is higher.
Artificial Analysis measured an average 1.3 minutes per task, down from 2.7 minutes for Gemini 3.5 Flash. Its estimated cost per benchmark task fell about 18%, from $0.59 to $0.50. Output speed reached roughly 304 tokens per second, the fastest result in its comparison class at publication.
That is the honest ranking: not an intelligence winner, but an efficiency monster.
The community reaction was immediate and mostly hostile
The most widely shared criticism reduced the launch to one painful chart. Developer Marcos Hernanz compared Gemini 3.6 Flash with GPT-5.6 Luna and wrote: “What are we doing Google?” His post framed Gemini as worse at 2.5 times the cost.

Another viral reaction focused on the absence of progress: Gemini 3.6 Flash scored the same 50 points as 3.5 Flash despite receiving a new version number. Reddit threads across Gemini, Antigravity, Cursor, and general AI communities repeated three complaints:
- The intelligence score did not improve. For users expecting a new model to be broadly smarter, halving task time feels like an optimization release wearing a generational name.
- Luna makes the pricing look weak. OpenAI’s model is marginally ahead on the same index while charging less per input and output token.
- Google still has not shipped Gemini 3.5 Pro broadly. Some users read another Flash release as a distraction from the higher-end model they were waiting for.
The tone ranged from disappointed to gleefully hostile. One Antigravity thread called the model less intelligent and more expensive than Luna. A Gemini community thread argued that speed would need to be exceptional to justify the position. Cursor users focused on coding scores and the higher token price.
This is not a scientific user study. Social platforms reward outrage, benchmark charts compress complex tradeoffs, and a one-point index gap is not proof that one model will win every real task. But the repetition of the same objection across communities matters: Google failed to make the value proposition obvious at launch.
The case for Gemini 3.6 Flash is better than the backlash suggests
The positive case starts with the number critics tend to crop out: #1 in output speed.
Developers building agents do not only pay for correct final answers. They pay in wall-clock time, tool calls, failed loops, generated tokens, and the delay before an agent can recover from a mistake. A model that finishes an average test task in 1.3 minutes instead of 2.7 can be more useful even if its composite intelligence score does not move.
Google also claims 3.6 Flash uses 17% fewer output tokens than 3.5 Flash and takes fewer reasoning steps, turns, and tool calls. Its API documentation says the model makes fewer unwanted code edits, reduces debugging loops, improves instruction following, and now supports Computer Use as a native tool. It accepts text, images, audio, and video, has a 1 million-token context window, and can return up to 64,000 output tokens.
There are early positive reports too. Some developers say Gemini Flash feels better for quick frontend iteration and small fixes because of its speed. Others point out that a fast model can retry and recover more quickly in agentic coding, which is not captured by a single intelligence-versus-cost chart.
Google even acknowledges a weakness that benchmark marketing normally hides: human evaluators preferred earlier models for visual layout and styling. The official migration guide recommends explicit design instructions when using 3.6 Flash for UI work. That is a useful warning for vibe coders expecting the model to supply taste automatically.
Why the launch still feels disappointing
The name “3.6” implies a smarter successor. The evidence instead describes a production optimization:
- Same broad intelligence score.
- Half the task time.
- Fewer output tokens.
- Lower output price than 3.5 Flash.
- Better results on selected agentic and computer-use benchmarks.
- Worse value than GPT-5.6 Luna if your priority is maximum index score per API dollar.
That is a meaningful engineering release. It is not the obvious capability jump people have learned to expect from a new frontier-model launch.
Google made the perception problem worse by launching 3.6 Flash while saying Gemini 3.5 Pro is still testing with partners. Users who were waiting for the next major Gemini intelligence leap received a faster version of an existing model instead. Meanwhile, OpenAI’s Luna offers a cleaner benchmark-and-price story, and fast open or low-cost models keep narrowing the gap.
The result is a model that may perform extremely well inside production agents but looks bad in the two-second social-media test: the dot is lower and farther to the right.
Who should actually use it
Gemini 3.6 Flash makes sense for:
- high-throughput agents where task latency matters more than the last point of benchmark intelligence;
- multimodal pipelines that need native image, audio, and video input;
- computer-use workloads built around Google’s agent stack;
- long-context document, chart, and data analysis;
- teams already using Gemini 3.5 Flash that want lower task cost and fewer runaway loops.
It is harder to recommend for:
- buyers choosing purely on intelligence per API dollar;
- coding workloads where GPT-5.6 Luna already performs reliably;
- visual frontend generation without strong design specifications;
- users expecting a broad new capability tier rather than an operational refinement.
The bottom line
Gemini 3.6 Flash may be one of Google’s most practical agent models.
But the community’s disappointment is understandable. Independent testing found no intelligence gain over 3.5 Flash, while a cheaper OpenAI model scored slightly higher. Google shipped world-class speed and efficiency, then gave the market a version number that invited everyone to judge it as a capability leap.
The launch verdict is simple: Gemini 3.6 Flash finishes tasks quickly. It does not lead on broad intelligence or price per token.
Sources
- Google: Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- Google AI for Developers: Gemini 3.6 Flash API details and migration guide
- Google DeepMind: Gemini 3.6 Flash model card
- Artificial Analysis: Gemini 3.6 Flash model results
- Artificial Analysis: Gemini 3.6 Flash halves time per task without improving intelligence
- Artificial Analysis: GPT-5.6 Luna model results
- Marcos Hernanz’s launch-day criticism on X
- Gemini community discussion: “Less intelligence for more money”
- Antigravity community benchmark discussion