← Back to API models

DeepSeek V4-Flash

by DeepSeek · budget tier

Still the price floor of the market: $0.22/$0.66 per 1M off-peak, doubling at peak hours, with a 1M window and real reasoning ability.

Input (off-peak) $0.22 / 1M
Output (off-peak) $0.66 / 1M
Context 1M tokens
Peak (01:00-04:00 & 06:00-10:00 UTC weekdays) $0.44 / $1.32
DeepSeek Platform ↗ Updated August 24, 2026
§ API pricing

Per-token rates, now split by time of day.

Input (off-peak)
$0.22/1M tokens
Prompt tokens
  • Up from $0.14 before Aug 16, 2026
  • A third of V4-Pro's off-peak $0.66
  • 1M window at this price is unmatched
Output (off-peak)
$0.66/1M tokens
Completion tokens
  • Up from $0.28 before Aug 16, 2026
  • A third of V4-Pro's off-peak $1.98
  • Still far cheaper than any US flagship
Peak (01:00-04:00 & 06:00-10:00 UTC weekdays)
$0.44/$1.32 per 1M
Peak-hour tier
  • Exactly 2x the off-peak input/output rate
  • New as of Aug 16, 2026
  • First time-of-day billing on the DeepSeek API
Hosting
Chinabased infrastructure
Consideration
  • Hosted version routes via China
  • Open-weights self-hosting available
  • Check your compliance needs

Why V4-Flash still leads on price

V4-Flash remains the cheapest model on our API table, even after DeepSeek introduced peak/off-peak billing on August 16, 2026 — its first time-of-day pricing scheme. Off-peak, it now runs $0.22 input / $0.66 output per 1M tokens (up from a flat $0.14/$0.28), doubling to $0.44/$1.32 during the 01:00-04:00 & 06:00-10:00 UTC weekday peak window. It keeps the full 1M context window either way, and still isn't a toy: it inherits the V4 family's reasoning lineage and beats every non-Chinese rival with a comparable window on price.

Like its big sibling V4-Pro, the considerations are non-technical: the hosted API routes through Chinese infrastructure, a hard blocker for some compliance regimes, and open-weights self-hosting is part of DeepSeek's pitch for teams that need to avoid it. If your workload is flexible on timing, scheduling batch jobs outside the peak window keeps costs at the old, lower rate.

Capabilities

For the price class, capability is still absurd: usable reasoning on math, code, and logic, a 1M window, and throughput suited to volume pipelines. Most tasks that teams route to mini-tier US models run fine here at a fraction of the cost, even at the new peak rate.

The honest weakness: polish and ecosystem. Tooling, SDK maturity, rate-limit headroom, and English prose quality all trail the US providers, and the hardest reasoning belongs to V4-Pro or a frontier model. Billing now also requires tracking time-of-day, adding a small layer of planning that flat-rate rivals don't need.

Typical use cases

  • Absolute-cost-floor volume pipelines, scheduled off-peak
  • Long-document processing on a budget (1M window)
  • Math/code/logic tasks too smart for nano-tier models
  • Self-hosted deployments via open weights
  • Cheap tier under V4-Pro in a routing stack

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
DeepSeek V4-Flash (off-peak)$0.22$0.661M
DeepSeek V4-Flash (peak)$0.44$1.321M
DeepSeek V4-Pro (off-peak)$0.66$1.981M
GPT-5.4 nano$0.20$1.25400K
Mistral Small 4$0.15$0.60256K

Off-peak, V4-Flash's input price now sits close to GPT-5.4 nano's — but GPT-5.4 nano's 400K window is still under half of V4-Flash's 1M, and V4-Flash's off-peak output price still undercuts it. If Chinese hosting is acceptable (or you self-host), V4-Flash off-peak is still the rational default for cheap volume; if not, GPT-5.4 nano and Mistral Small are the compliant runners-up.

← See all DeepSeek plans