OpenAI cuts GPT-5.6 Luna 80% to $0.20/$1.20 per million tokens

Terra drops 20% too, Fast mode trades 2.5x speed for 2x price, GPT Transcribe undercuts Whisper, and Moonshot's Kimi K3 opens 2.8T weights.

Nowline AUG 1 9:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Luna falls 80%, Terra 20%

    GPT-5.6 Luna now runs $0.20/$1.20 per million input/output tokens — an 80% cut — while the balanced Terra drops 20% to $2/$12. Luna is suddenly priced like a small model while keeping frontier-family reasoning, which resets the cost math for high-volume agent loops and batch jobs.

  • Fast mode replaces Priority Processing

    A new Fast mode delivers up to 2.5x standard speed at twice the price with no change in intelligence, superseding Priority Processing on GPT-5.6 Sol. Reach for it when latency — not token cost — is the bottleneck in interactive agents and live tools.

  • GPT Transcribe undercuts Whisper

    OpenAI's new GPT Transcribe and GPT Live Transcribe hit an 8.98% word error rate versus Whisper-1's 15.21%, with keyword hints and low-latency streaming at roughly $0.27/$1.02 per hour. A near drop-in upgrade for voice-in apps, meeting notes, and live captioning.

  • Kimi K3 opens 2.8T weights

    Moonshot's Kimi K3 is out under open weights — 2.8T total / 104B active params and a 1M-token context — shipping with MXFP4 quantization and ranking near the top of the intelligence index. It's the largest open-weight model yet if you can feed the VRAM to self-host it.