← Back to API models

Gemini 3.5 Flash-Lite

by Google DeepMind · fastest Flash tier · released July 21, 2026

The cheap-volume, low-latency tier — around 350 tokens/second, succeeding 3.1 Flash-Lite with significantly better quality at the same 1M context.

Input $0.30 / 1M
Output $2.50 / 1M
Context 1M tokens
Tier Fast / cheap
Google AI Studio ↗ Updated July 23, 2026
§ API pricing

Per-token rates.

Input
$0.30/1M tokens
Prompt tokens
  • Fifth of 3.6 Flash's input rate
  • Native multimodal input
  • Succeeds 3.1 Flash-Lite ($0.25)
Output
$2.50/1M tokens
Completion tokens
  • Third of 3.6 Flash's output rate
  • Great for high-volume output
  • Includes thinking tokens
Speed
~350tokens / sec
Fastest tier
  • Lowest latency in the Flash lineup
  • Built for high-throughput scenarios
  • Significantly better than 3.1 Flash-Lite
Context
1Mtokens
Window
  • Full 1M window
  • Native multimodal across the window
  • Matches the larger Flash models on length

Flash-Lite is the cheap-volume tier at $0.30 / 1M input and $2.50 / 1M output — a fifth of 3.6 Flash's input cost, at the fastest speeds Google offers in the Flash lineup.

What Flash-Lite is for

Released July 21, 2026 alongside Gemini 3.6 Flash, Gemini 3.5 Flash-Lite is the fastest, cheapest tier in the Gemini Flash lineup. Google positions it as the successor to 3.1 Flash-Lite with "significantly better quality," generating around 350 tokens per second — the tier you reach for when latency and per-call cost dominate.

At $0.30 input and $2.50 output per 1M tokens, it sits well below the 3.6 Flash workhorse while keeping the full 1M context and native multimodal input. It's the natural home for classification, routing, re-ranking, and bulk summarization at scale.

Capabilities

Flash-Lite handles everyday generation, extraction, and lightweight reasoning with very low latency and native multimodal input. It won't match 3.6 Flash on hard coding or agentic computer-use tasks — that's what the workhorse tier is for — but for high-volume pipelines it's dramatically cheaper and faster.

Typical use cases

  • High-volume classification, routing, and tagging
  • Structured extraction at scale
  • Latency-sensitive interactive features
  • Bulk summarization and re-ranking
  • Cost-sensitive sub-steps inside larger agent pipelines

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
Gemini 3.5 Flash-Lite$0.30$2.501M
Gemini 3.6 Flash$1.50$7.501M
GPT-5.4 mini$0.25$2272K
Claude Haiku 4.5$1$5200K

Flash-Lite is the cheap-volume tier below 3.6 Flash. Its closest rival is GPT-5.4 mini — slightly cheaper but with a shorter context and no native video. For anything needing stronger reasoning or coding, step up to 3.6 Flash.

← See all Google / Gemini plans