← Back to API models

Gemini 3.6 Flash

by Google DeepMind · new workhorse Flash · released July 21, 2026

Google's new default Flash — 17% fewer output tokens than 3.5 Flash, stronger coding, built-in Computer Use, and an output price cut from $9 to $7.50.

Input $1.50 / 1M
Output $7.50 / 1M
Context 1M tokens
Tier Workhorse
Google AI Studio ↗ Updated July 23, 2026
§ API pricing

Per-token rates.

Input
$1.50/1M tokens
Prompt tokens
  • Unchanged from 3.5 Flash
  • Vision, audio, video billed as input
  • Cached input just $0.15 / 1M
Output
$7.50/1M tokens
Completion tokens
  • Cut from $9 on 3.5 Flash
  • Plus 17% fewer output tokens used
  • Includes thinking tokens
Context
1Mtokens
Window · 64K out
  • Full 1M input, up to 64K output
  • Native multimodal across the window
  • Thinking controls + Computer Use
Batch / Flex
$0.75/$3.75 per 1M
Async tier
  • Roughly half the standard rate
  • For latency-tolerant workloads
  • Same model, cheaper throughput

Gemini 3.6 Flash holds the $1.50 / 1M input rate while cutting output to $7.50 / 1M (from $9). Combined with ~17% fewer output tokens, real-world cost drops meaningfully versus 3.5 Flash. Cached input is $0.15 / 1M.

What's new in Gemini 3.6 Flash

Released July 21, 2026, Gemini 3.6 Flash is Google's new workhorse — the default Flash that replaces Gemini 3.5 Flash. It arrived alongside a new Gemini 3.5 Flash-Lite and a restricted Gemini 3.5 Flash Cyber model, and Google used the launch to tease a larger Gemini 4 release still to come. The headline is efficiency: 3.6 Flash burns about 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index — up to 65% fewer on the DeepSWE coding benchmark — while the output price itself drops from $9 to $7.50 per 1M.

Quality moved up too. DeepSWE coding jumps to 49% (from 37% on 3.5 Flash), computer use improves from 78.4% to 83.0% on OSWorld-Verified, and generation runs around 275 tokens/second — fast for a reasoning model in this price tier. Built-in Computer Use makes it a practical driver for agentic UI tasks.

Capabilities

Flash inherits the native multimodality of the Gemini family: text, image, audio, and video all go in, across the full 1M context, with up to 64K output tokens and thinking controls. The 3.6 update leans hardest into coding and computer use, making it a strong fit for agent loops and IDE assistants. Pro still wins the hardest single-shot reasoning, but for most production work 3.6 Flash is now the obvious default.

Typical use cases

  • Coding assistants and IDE integrations — the headline use case
  • Agentic UI automation via built-in Computer Use
  • Production chat where quality and cost both matter
  • Long-document and video/audio analysis over 1M context
  • High-volume workloads via the cheaper Batch/Flex tier

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
Gemini 3.6 Flash$1.50$7.501M
Gemini 3.5 Flash-Lite$0.30$2.501M
Gemini 3.1 Pro$2$121M
GPT-5.6 Luna$1$61M

3.6 Flash sits just below Gemini 3.1 Pro on price and now leads it on coding — pick Pro for the hardest reasoning, 3.6 Flash for most everything else. Flash-Lite is the genuine cheap-volume tier. Its closest external rival is GPT-5.6 Luna, slightly cheaper but without native video.

← See all Google / Gemini plans