← Back to API models

DeepSeek V4-Pro

by DeepSeek AI · Hangzhou, China · GA since Aug 13, 2026

Top-tier reasoning at a small fraction of frontier prices — now billed on a peak/off-peak schedule, DeepSeek's first time-of-day pricing.

Input (off-peak) $0.66 / 1M
Output (off-peak) $1.98 / 1M
Context 1M tokens
Peak (01:00-04:00 & 06:00-10:00 UTC weekdays) $1.32 / $3.96
DeepSeek Platform ↗ Updated August 24, 2026
§ API pricing

Per-token rates, now split by time of day.

Input (off-peak)
$0.66/1M tokens
Prompt tokens
  • Cached input $0.022 / 1M off-peak
  • Up from $0.435 before Aug 16, 2026
  • Chain-of-thought billed in output
Output (off-peak)
$1.98/1M tokens
Completion tokens
  • Up from $0.87 before Aug 16, 2026
  • Still well under GPT-5.5 and Opus 5
  • Includes reasoning trace tokens
Peak (01:00-04:00 & 06:00-10:00 UTC weekdays)
$1.32/$3.96 per 1M
Peak-hour tier
  • Exactly 2x the off-peak input/output rate
  • Cached input $0.044 / 1M at peak
  • New as of Aug 16, 2026
Free chat
$0chat.deepseek.com
Consumer
  • V4-Pro is free in DeepSeek chat
  • No ads, no paywall — peak/off-peak billing is API-only
  • Web, iOS, Android

The price story

DeepSeek V4-Pro reached general availability on August 13, 2026 (build V4-Pro-0813), adding selectable low/high/max thinking-effort levels and native Responses API support after months in preview. Three days later, on August 16, DeepSeek introduced its first time-of-day pricing: a peak window from 16:00 to 00:30 UTC now bills exactly double the off-peak rate. Off-peak, V4-Pro runs $0.66 input / $1.98 output per 1M tokens — up from the flat $0.435/$0.87 it charged before the change. At peak, that becomes $1.32/$3.96.

Even at the new peak rate, V4-Pro remains far cheaper than western frontier models, but the gap has narrowed and the pricing is no longer a single flat number — budget around the daily peak window if usage volume matters. The honest counterweights are unchanged: DeepSeek is a Chinese company hosting in China, and data residency is a real consideration for regulated western enterprises. Open weights mitigate this since V4-Pro can be self-hosted by anyone with the GPUs.

Capabilities

V4-Pro is strongest at math, logic, and chain-of-thought reasoning, which has been DeepSeek's calling card since R1 in early 2025. The Aug 13 GA release added selectable thinking-effort levels (low/high/max), letting callers trade latency and cost against reasoning depth. Coding is competitive with GPT-5.4 on most public benchmarks, especially on algorithmic and competition-style problems.

Where it trails the western frontier: nuanced writing voice, tool-use polish, and instruction-following on edge cases. UX around the API (rate limits, observability, SLAs) is also less mature than OpenAI or Anthropic — and now comes with a peak-hour billing schedule to track.

Typical use cases

  • Math, logic, and quantitative reasoning at scale
  • Code generation and competitive-programming-style tasks
  • Cost-sensitive RAG and document QA at 1M context
  • Batch/offload workloads scheduled outside the peak window
  • Open-weight deployments behind a corporate firewall

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
DeepSeek V4-Pro (off-peak)$0.66$1.981M
DeepSeek V4-Pro (peak)$1.32$3.961M
DeepSeek V4-Flash (off-peak)$0.22$0.661M
Gemini 3.6 Flash$0.75$3.751M
GPT-5.5$5$301M

Even at peak pricing, V4-Pro undercuts GPT-5.5 by a wide margin. Off-peak, it now lands close to Gemini 3.6 Flash's intro price rather than clearly under it — the first time a DeepSeek flagship hasn't been the unambiguous budget leader in this comparison. V4-Flash remains the cheapest reasoning option on the ledger at either time of day.

← See all DeepSeek plans