Input, output and cached-input rates for 37 text models from 12 providers, plus image, video and voice generation APIs. Every price comes from the provider's official pricing page.
API pricing is per-million-tokens. Roughly: 1M tokens ≈ 750,000 words ≈ a stack of paperbacks. Output tokens cost 3–5× input tokens, and cached input can cost up to 98% less. Below: every model worth using right now, grouped by provider — flagships first.
“—” means the provider doesn’t publish that rate. Cache read is the discounted price for input tokens the provider has recently seen. Only Anthropic charges a cache-write fee (+25% of input, 5-minute TTL); OpenAI, Google, SpaceXAI and DeepSeek cache automatically at no extra write cost — Google’s explicit caching bills storage per hour instead.
Generation APIs don't bill per token. Image models charge per picture, video models per second of output, voice models per minute or per million characters.
Image prices are per generated image at default 1024×1024 quality — ≈ marks estimates converted from token-based billing. Video prices are per second of output at the listed resolutions. Sources: official OpenAI, Google and SpaceXAI pricing pages and Replicate list prices, verified Aug 4, 2026.
Translating a 2,500-word essay (≈ 6,666 tokens, half in / half out):
TipStart cheap, upgrade only if the output isn't good enough. The cheapest model is usually 95% as good for simple tasks.
Almost every provider bills per million tokens, with separate rates for input (what you send) and output (what the model writes). A token is roughly ¾ of an English word, so 1M tokens ≈ 750,000 words. Output tokens usually cost 3–5× more than input tokens, and input the provider has recently seen (cache reads) can cost up to 90% less.
Ministral 3 14B is the cheapest text model on this page at $0.20 / $0.20 per 1M tokens (input / output). Among flagship-tier models, Grok 4.6 is the cheapest at $2.00 / $6.00. The most expensive is GPT-5.5 pro at $30.00 / $180.00.
Flagship rates per 1M input / output tokens: GPT-5.5 pro $30.00 / $180.00; GPT-5.6 Sol $4.00 / $20.00; Claude Fable 5 $10.00 / $50.00; Claude Opus 5 $5.00 / $25.00; Gemini 3.1 Pro $2.00 / $12.00. Each provider also sells cheaper mid-tier and budget models that handle most everyday tasks.
Prompt caching lets a provider reuse input it has already processed — a long system prompt, a document, a conversation history. Cached input is billed at the cache-read rate, typically 50–90% below the normal input price. Only Anthropic charges a cache-write fee (+25% of the input price); OpenAI, Google, SpaceXAI and DeepSeek cache automatically at no extra cost.
For light use, yes: a few hundred short exchanges a month on a mid-tier model cost cents, not $20. Heavy daily use of a flagship model can exceed a subscription. Use the cost calculator to estimate your monthly bill from messages per day and message length.