OpenAI cuts GPT-5.6 Luna 80%, Terra 20%; Sol gets Fast Mode

Luna falls to $0.20/$1.20 per million tokens; the cuts flow into Codex and ChatGPT Work credits, and Fast Mode gives Sol 2.5x speed at 2x cost.

Nowline Jul 31 12:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Luna craters 80% to $0.20/$1.20 per M

    GPT-5.6 Luna's input drops to $0.20 and output to $1.20 per million tokens — an 80% cut that lands it below most open-weights hosting. High-volume extraction, classification, and routing just got roughly 5x cheaper on OpenAI's fastest model.

  • Terra trimmed 20% to $2/$12; Sol unchanged

    The balanced everyday model, GPT-5.6 Terra, now runs $2 input / $12 output per million. Frontier Sol pricing holds, so the economics reward routing routine work down to Luna and Terra and saving Sol for the hard calls.

  • Fast Mode replaces Priority Processing

    For Sol, the new Fast mode delivers up to 2.5x faster responses at 2x the standard price, with no change in intelligence. It supersedes Priority Processing and is backward compatible — requests already tagged `priority` auto-upgrade, so no code change is needed.

  • Your Codex and ChatGPT Work credits stretch further

    OpenAI says the lower Luna and Terra prices are reflected in how usage is metered inside Codex and ChatGPT Work, so the same subscription budget buys more calls. Plan prices and quota budgets are unchanged.

  • Build this weekend

    Re-price your agent loops: move retrieval, tool-routing, and summarization to Luna and you can roughly 5x call volume on the same spend. Wire a confidence-based router — cheap Luna draft, escalate to Sol only when confidence is low — and measure the quality delta.