← Back to API models

Gemini 3.7 Flash

by Google DeepMind · newest workhorse Flash · released August 13, 2026

Google's newest workhorse for coding and agents — $0.75/$3.75 per 1M tokens through the end of 2026, with in-app access via the new "Spark" feature on AI Pro/Ultra.

Input (intro) $0.75 / 1M
Output (intro) $3.75 / 1M
Context 1M tokens
Reverts $1.50/$7.50 Jan 1, 2027
Google AI Studio ↗ Updated August 17, 2026
§ API pricing

Per-token rates, at an introductory discount.

Input (intro)
$0.75/1M tokens
Prompt tokens
  • Introductory rate through Dec 31, 2026
  • Vision, audio, video billed as input
  • Cached input just $0.075 / 1M
Output (intro)
$3.75/1M tokens
Completion tokens
  • Introductory rate through Dec 31, 2026
  • Same price as sibling Gemini 3.6 Flash
  • Includes thinking tokens
Context
1Mtokens
Window
  • Native multimodal input
  • Tuned for coding and agentic tasks
  • Context caching: $0.075/hr read, $0.50/hr write per 1M
Reverts Jan 1, 2027
$1.50/$7.50 per 1M
Standard rate
  • Intro pricing holds through Dec 31, 2026
  • Standard rate matches other Flash-tier models
  • No action needed — automatic on Google's end

Gemini 3.7 Flash bills $0.75 / 1M input and $3.75 / 1M output through Dec 31, 2026 as an introductory rate. Cached input is $0.075 / 1M. The rate reverts to a standard $1.50/$7.50 on Jan 1, 2027.

What's new in Gemini 3.7 Flash

Google launched Gemini 3.7 Flash on August 13, 2026, calling it its most intelligent workhorse model yet for coding and agentic work. It ships at an introductory $0.75 input / $3.75 output per 1M tokens — the same rate Google simultaneously cut sibling Gemini 3.6 Flash to, so the two models are priced identically through the end of the year. In the consumer Gemini app, AI Pro and AI Ultra subscribers get access to 3.7 Flash through a new feature called "Spark," rolled out in over 160 countries at launch.

The intro pricing window is temporary: Google's docs list Jan 1, 2027 as the date rates revert to a standard $1.50/$7.50 per 1M — the same level 3.6 Flash launched at back in July. Build cost models accordingly if your workload extends past the new year.

Capabilities

3.7 Flash inherits the native multimodality of the Gemini family — text, image, audio, and video input across a 1M token context — with Google positioning it specifically for coding and agent workloads, ahead of general-purpose chat. Context caching is priced separately per hour ($0.075 read / $0.50 write per 1M tokens), which matters for agents that keep large system prompts warm across many calls.

Typical use cases

  • Coding agents and IDE integrations
  • Multi-step agentic workflows and tool-use loops
  • In-app access via "Spark" for AI Pro/Ultra subscribers
  • Production workloads that benefit from context caching
  • General chat and document work at the discounted intro rate

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
Gemini 3.7 Flash (intro)$0.75$3.751M
Gemini 3.6 Flash (intro)$0.75$3.751M
Gemini 3.5 Flash-Lite$0.30$2.501M
Gemini 3.1 Pro$2$121M
GPT-5.6 Luna$0.20$1.201M

3.7 Flash and 3.6 Flash cost the same through 2026 — pick 3.7 Flash if coding/agent benchmarks matter most to your workload, since that's what Google explicitly tuned it for. Flash-Lite is still the cheaper-volume option, and GPT-5.6 Luna undercuts both Flash models on raw price but lacks native video.

← See all Google / Gemini plans