$ AI PriceCompare
← Back to plan comparison
API rates

Every AI API, priced per million tokens.

Input, output and cached-input rates for 37 text models from 12 providers, plus image, video and voice generation APIs. Every price comes from the provider's official pricing page.

§ 01 — Text & reasoning models

37 models, grouped by provider.

API pricing is per-million-tokens. Roughly: 1M tokens ≈ 750,000 words ≈ a stack of paperbacks. Output tokens cost 3–5× input tokens, and cached input can cost up to 98% less. Below: every model worth using right now, grouped by provider — flagships first.

View
Model
Input / 1M
Output / 1M
Cache read / 1M
Cache write / 1M
Context
OpenAI
6 models
Most capable, smartest
$30.00
$180.00
1M
Cut 20% Aug 21 (promo thru Nov 21); >272K tokens: $8/$30, cache $0.80
$4.00
$20.00
$0.40
1M
Balanced 5.6; cut 20% Jul 30
$2.00
$12.00
$0.20
1M
Cheaper, well-defined tasks
$0.75
$4.50
$0.075
400K
Fastest, cheapest
$0.20
$1.25
$0.02
400K
Fast & cheapest 5.6 tier; cut 80% Jul 30
$0.20
$1.20
$0.02
1M
Anthropic
4 models
Most powerful; restored Jul 1 after export-control pause
$10.00
$50.00
$1.00
$12.50
1M
New Opus flagship (Jul 24); near-Fable 5 at half the price
$5.00
$25.00
$0.50
$6.25
1M
$2/$10 is now permanent (Sep 1 increase to $3/$15 cancelled)
$2.00
$10.00
$0.20
$2.50
1M
Fast, near Sonnet 4 quality
$1.00
$5.00
$0.10
$1.25
200K
Google
7 models
Frontier multimodal (>200K: $4/$18)
$2.00
$12.00
$0.20
1M
Long context (1M); >200K: $2.50/$15
$1.25
$10.00
$0.125
1M
New workhorse (Aug 13, 2026) for coding & agents; intro price thru Dec 31, reverts to $1.50/$7.50 Jan 1
$0.75
$3.75
$0.075
1M
Cut to match 3.7 Flash intro price thru Dec 31 2026; reverts to $1.50/$7.50 Jan 1
$0.75
$3.75
$0.075
1M
Fastest Flash tier (~350 tok/s); succeeds 3.1 Flash-Lite
$0.30
$2.50
$0.03
1M
Best price / performance
$0.30
$2.50
$0.03
1M
Fast, affordable
$0.10
$0.40
$0.01
1M
Mistral
5 models
Mistral Medium 3.5
Frontier-class multimodal; agentic & coding focus
$1.50
$7.50
256K
Mistral flagship
$0.50
$1.50
256K
Code generation
$0.30
$0.90
128K
Mid-tier general; replaced Small 3.1
$0.15
$0.60
256K
Budget multimodal; replaced Pixtral 12B
$0.20
$0.20
256K
SpaceXAI
3 models
New flagship (Aug 12, 2026); replaces 4.5 at same price; >200K tokens: $4/$12
$2.00
$6.00
$0.50
500K
Reasoning workhorse; replaced Grok 4; >200K tokens: $2.50/$5
$1.25
$2.50
$0.20
1M
Code tasks; replaced Grok Code Fast 1; >200K tokens: $2/$4
$1.00
$2.00
$0.20
256K
DeepSeek
3 models
Off-peak rate; peak (01:00-04:00 & 06:00-10:00 UTC weekdays) is 2x: $1.32/$3.96
$0.66
$1.98
$0.022
1M
Off-peak rate; peak (01:00-04:00 & 06:00-10:00 UTC weekdays) is 2x: $0.44/$1.32
$0.22
$0.66
$0.007
1M
DeepSeek V4-Flash-Vision-Exp
Experimental multimodal/vision variant of V4-Flash (Aug 21, 2026); same rate card
$0.22
$0.66
$0.007
1M
Meta
Meta's first paid model, agentic; 1.2 succeeds 1.1 at same price
$1.25
$4.25
$0.15
1M
Alibaba
3 models
Qwen3.8-Max
New Alibaba flagship (Aug 3, 2026); 2.4T MoE
$2.00
$6.00
$0.17
1M
Qwen3.7-Max
Prior Qwen flagship; strong multilingual
$1.25
$3.75
$0.13
1M
Qwen3.7-Plus
Balanced Qwen workhorse
$0.32
$1.28
Moonshot AI
Kimi K3
Top open-weights agentic model
$3.00
$15.00
$0.30
1M
Z.ai
GLM-5.3
GLM flagship; replaced 5.2 at same price; strong coding value
$1.40
$4.40
$0.26
1M
Perplexity
2 models
Sonar Pro
Web search built in; per-request fees apply
$3.00
$15.00
200K
Sonar
Cheap web-grounded answers; per-request fees apply
$1.00
$1.00
128K
Cohere
Command A+
Succeeds Command A; open-weight (Apache 2.0); hosted access via Cohere/Azure is enterprise-only — no public metered price
256K

“—” means the provider doesn’t publish that rate. Cache read is the discounted price for input tokens the provider has recently seen. Only Anthropic charges a cache-write fee (+25% of input, 5-minute TTL); OpenAI, Google, SpaceXAI and DeepSeek cache automatically at no extra write cost — Google’s explicit caching bills storage per hour instead.

§ 02 — Beyond text

Image, video & voice APIs.

Generation APIs don't bill per token. Image models charge per picture, video models per second of output, voice models per minute or per million characters.

Image generation APIs

Model
Price
Most powerful
GPT-Image-2
OpenAI · Newest OpenAI flagship; photoreal edits
≈$0.032 / image
GPT-Image-1.5
OpenAI · Being retired Dec 1, 2026 — migrate to GPT-Image-2, now the top-rated model
≈$0.034 / image
Gemini 3.1 Flash Image
Google · Tiered by resolution: $0.045 @512px, $0.067 @1024px, $0.151 @4K
$0.045–0.151 / image
Gemini 3 Pro Image
Google · "Nano Banana Pro"; premium tier above Flash Image, up to 4K
$0.134–0.24 / image
Seedream 4.5
ByteDance · Strong realism & spatial understanding
$0.04 / image
FLUX 1.1 Pro
BFL · Open-ecosystem favorite
$0.04 / image
Recraft V3
Recraft · Design, vector styles & long text
$0.04 / image
Fast & budget
Ideogram V3 Turbo
Ideogram · Fast; best-in-class text rendering
$0.03 / image
GPT-Image-1-mini
OpenAI · Cheap OpenAI tier
≈$0.008 / image
FLUX Schnell
BFL · Cheapest solid quality; open weights
$0.003 / image

Video generation APIs

Model
Price
Most powerful
Sora 2 Pro
OpenAI · Top cinematic quality (720p–1080p)
$0.30–0.70 / s
Veo 3.1
Google · Frontier quality with audio (720p–4K)
$0.40–0.60 / s
Grok Imagine 1.5
SpaceXAI · Cheapest frontier tier (480p–1080p)
$0.08–0.25 / s
Fast & budget
Sora 2
OpenAI · Standard tier, 720p
$0.10 / s
Veo 3.1 Fast
Google · Faster, cheaper Veo (720p–4K)
$0.10–0.30 / s
Veo 3.1 Lite
Google · Cheapest Veo tier (720p–1080p)
$0.05–0.08 / s
Grok Imagine Video
SpaceXAI · Cheapest video API (480p–720p)
$0.05–0.07 / s

Voice & audio APIs

Model
Price
Most powerful
GPT-Realtime-2.1
OpenAI · Speech-to-speech agents with tool use
$32 in / $64 out per 1M audio tok
Grok Voice Think Fast 2.0
SpaceXAI · Real-time voice with reasoning
$0.08 / min
Gemini 3.1 Flash TTS
Google · High-quality controllable TTS
$20 / 1M audio tok out
Fast & budget
TTS-1 HD
OpenAI · HD text-to-speech
$30 / 1M chars
TTS-1
OpenAI · Standard text-to-speech
$15 / 1M chars
Grok Voice Think Fast 1.0
SpaceXAI · Budget real-time voice
$0.05 / min
Whisper
OpenAI · Cheap transcription (speech-to-text)
$0.006 / min

Image prices are per generated image at default 1024×1024 quality — ≈ marks estimates converted from token-based billing. Video prices are per second of output at the listed resolutions. Sources: official OpenAI, Google and SpaceXAI pricing pages and Replicate list prices, verified Aug 4, 2026.

Worked example

Translating a 2,500-word essay (≈ 6,666 tokens, half in / half out):

  • Gemini 2.5 Flash-Lite$0.002
  • GPT-5.6 Luna$0.005
  • Claude Haiku 4.5$0.020
  • Gemini 3.1 Pro$0.047
  • Claude Fable 5$0.200
  • GPT-5.5 pro$0.700

TipStart cheap, upgrade only if the output isn't good enough. The cheapest model is usually 95% as good for simple tasks.

§ 03 — FAQ

API pricing, explained.

How is AI API pricing calculated?

Almost every provider bills per million tokens, with separate rates for input (what you send) and output (what the model writes). A token is roughly ¾ of an English word, so 1M tokens ≈ 750,000 words. Output tokens usually cost 3–5× more than input tokens, and input the provider has recently seen (cache reads) can cost up to 90% less.

Which AI API is the cheapest in 2026?

Ministral 3 14B is the cheapest text model on this page at $0.20 / $0.20 per 1M tokens (input / output). Among flagship-tier models, Grok 4.6 is the cheapest at $2.00 / $6.00. The most expensive is GPT-5.5 pro at $30.00 / $180.00.

How much do GPT, Claude and Gemini cost via the API?

Flagship rates per 1M input / output tokens: GPT-5.5 pro $30.00 / $180.00; GPT-5.6 Sol $4.00 / $20.00; Claude Fable 5 $10.00 / $50.00; Claude Opus 5 $5.00 / $25.00; Gemini 3.1 Pro $2.00 / $12.00. Each provider also sells cheaper mid-tier and budget models that handle most everyday tasks.

What is prompt caching, and how much does it save?

Prompt caching lets a provider reuse input it has already processed — a long system prompt, a document, a conversation history. Cached input is billed at the cache-read rate, typically 50–90% below the normal input price. Only Anthropic charges a cache-write fee (+25% of the input price); OpenAI, Google, SpaceXAI and DeepSeek cache automatically at no extra cost.

Is the API cheaper than a ChatGPT or Claude subscription?

For light use, yes: a few hundred short exchanges a month on a mid-tier model cost cents, not $20. Heavy daily use of a flagship model can exceed a subscription. Use the cost calculator to estimate your monthly bill from messages per day and message length.