The cheap-volume, low-latency tier — around 350 tokens/second, succeeding 3.1 Flash-Lite with significantly better quality at the same 1M context.
Flash-Lite is the cheap-volume tier at $0.30 / 1M input and $2.50 / 1M output — a fifth of 3.6 Flash's input cost, at the fastest speeds Google offers in the Flash lineup.
Released July 21, 2026 alongside Gemini 3.6 Flash, Gemini 3.5 Flash-Lite is the fastest, cheapest tier in the Gemini Flash lineup. Google positions it as the successor to 3.1 Flash-Lite with "significantly better quality," generating around 350 tokens per second — the tier you reach for when latency and per-call cost dominate.
At $0.30 input and $2.50 output per 1M tokens, it sits well below the 3.6 Flash workhorse while keeping the full 1M context and native multimodal input. It's the natural home for classification, routing, re-ranking, and bulk summarization at scale.
Flash-Lite handles everyday generation, extraction, and lightweight reasoning with very low latency and native multimodal input. It won't match 3.6 Flash on hard coding or agentic computer-use tasks — that's what the workhorse tier is for — but for high-volume pipelines it's dramatically cheaper and faster.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M |
| Gemini 3.6 Flash | $1.50 | $7.50 | 1M |
| GPT-5.4 mini | $0.25 | $2 | 272K |
| Claude Haiku 4.5 | $1 | $5 | 200K |
Flash-Lite is the cheap-volume tier below 3.6 Flash. Its closest rival is GPT-5.4 mini — slightly cheaper but with a shorter context and no native video. For anything needing stronger reasoning or coding, step up to 3.6 Flash.