OpenAI cuts GPT-5.6 Luna 80% to $0.20/$1.20 per million tokens
Terra drops 20% too, Fast mode trades 2.5x speed for 2x price, GPT Transcribe undercuts Whisper, and Moonshot's Kimi K3 opens 2.8T weights.

Copy markdown
Luna falls 80%, Terra 20%
GPT-5.6 Luna now runs $0.20/$1.20 per million input/output tokens — an 80% cut — while the balanced Terra drops 20% to $2/$12. Luna is suddenly priced like a small model while keeping frontier-family reasoning, which resets the cost math for high-volume agent loops and batch jobs.
Fast mode replaces Priority Processing
A new Fast mode delivers up to 2.5x standard speed at twice the price with no change in intelligence, superseding Priority Processing on GPT-5.6 Sol. Reach for it when latency — not token cost — is the bottleneck in interactive agents and live tools.
GPT Transcribe undercuts Whisper
OpenAI's new GPT Transcribe and GPT Live Transcribe hit an 8.98% word error rate versus Whisper-1's 15.21%, with keyword hints and low-latency streaming at roughly $0.27/$1.02 per hour. A near drop-in upgrade for voice-in apps, meeting notes, and live captioning.
Kimi K3 opens 2.8T weights
Moonshot's Kimi K3 is out under open weights — 2.8T total / 104B active params and a 1M-token context — shipping with MXFP4 quantization and ranking near the top of the intelligence index. It's the largest open-weight model yet if you can feed the VRAM to self-host it.