OpenAI cuts GPT-5.6 Luna 80% and Terra 20% on the API rate card

Serving-efficiency gains, OpenAI says — though the timing tracks a cost war. Sol holds, Priority becomes Fast mode, and your production bill just dropped.

Nowline AUG 3 12:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Luna and Terra get cheaper on the same rate card

    GPT-5.6 Luna drops 80% to $0.20 per million input tokens and $1.20 output; Terra falls 20% to $2.00 input and $12 output. The cut is already live — no opt-in and no new model IDs, so anything already calling these tiers just got cheaper.

  • Sol holds — and Priority Processing becomes Fast mode

    The flagship GPT-5.6 Sol stays at $5/$30 per million tokens. OpenAI also swapped Priority Processing for a new Fast mode (~2x standard rates) for latency-sensitive calls, while batch stays at half price and Luna's cached input drops to about $0.02 per million.

  • Why now: the token-price floor keeps dropping

    OpenAI credits serving-efficiency gains, but Luna at $0.20/$1.20 now lands in cheap-open-model territory that Qwen 3.7 Flash ($0.03/$0.13) and DeepSeek's V4-Flash have been carving out. For high-volume or agentic loops, re-checking whether you still need a pricier tier is worth an afternoon.