← Back to API models

GPT-5.6 Luna

by OpenAI · fast & affordable · released July 9, 2026

The fastest, cheapest member of the GPT-5.6 family — strong capability at OpenAI's lowest GPT-5.6 price, built for high-volume and latency-sensitive work.

Input $0.20 / 1M
Output $1.20 / 1M
Context 1M
Tier Fast / cheap
OpenAI Platform ↗ Updated August 3, 2026
§ API pricing

Per-token rates.

Input
$0.20/1M tokens
Prompt tokens
  • Cut from $1 on Jul 30, 2026
  • Vision inputs billed as tokens
  • Cache reads get ~90% discount
Output
$1.20/1M tokens
Completion tokens
  • Cut from $6 on Jul 30, 2026
  • Includes reasoning tokens
  • Great for high-volume output
Context
1Mtokens
Window
  • Large window despite low price
  • Official figure not published
  • Shares the GPT-5.6 architecture
Caching
1.25×on cache writes
Prompt caching
  • Explicit cache breakpoints
  • 30-minute minimum cache life
  • Cache reads keep the 90% discount

Luna is the entry point to the GPT-5.6 family at $0.20 / 1M input and $1.20 / 1M output — cut from $1/$6 on July 30, 2026, and now a fraction of Sol's output cost for latency-sensitive, high-throughput work.

What's new in GPT-5.6 Luna

Luna is the fast-and-affordable tier of the GPT-5.6 family, released publicly on July 9, 2026 alongside Sol and Terra at $1 input and $6 output per 1M, bringing GPT-5.6-generation capability to workloads that were previously served by mini-class models but need more headroom. OpenAI cut the price further to $0.20 / $1.20 per 1M tokens on July 30, 2026.

Luna is the model to reach for when throughput and cost dominate: classification, routing, extraction, and high-volume chat where each request is relatively simple but there are a lot of them. It doesn't carry Sol's max/ultra reasoning modes, and it trails Terra on the hardest tasks, but it's dramatically cheaper.

Capabilities

Luna handles everyday generation, summarization, extraction, and lightweight tool use with low latency. It's a natural fit for pipelines that fan out many small calls, and for interactive apps where responsiveness matters more than frontier reasoning. For harder problems, step up to Terra or Sol.

Typical use cases

  • High-volume classification, routing, and tagging
  • Structured extraction at scale
  • Latency-sensitive interactive chat
  • Draft generation and summarization pipelines
  • Cost-sensitive agent sub-steps

Sibling and rival comparison

ModelInput / 1MOutput / 1MContext
GPT-5.6 Luna$0.20$1.201M
GPT-5.6 Terra$2$121M
GPT-5.4 nano$0.20$1.25400K
GPT-5.4 mini$0.75$4.50400K
Claude Haiku 4.5$1$5200K
Gemini 3.5 Flash$1.50$91M

Luna is now priced right alongside GPT-5.4 nano — matching its input rate and slightly undercutting it on output — while carrying full GPT-5.6-generation capability and a 1M context window. Its closest full-size rivals are Claude Haiku 4.5 and Gemini 3.5 Flash.

← See all ChatGPT / OpenAI plans