Report: Claude Code appears to be A/B testing lower effort on Opus

Paying Opus prices for slower, over-thinking runs? Builders are gauge-testing sessions and rerouting to Codex — plus Tencent's open 33-language translator.

Nowline AUG 23 1:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • What builders are seeing

    Front-page reports say Opus 4.9/5 sessions now burn effort far past the ask — one user clocked a 2-minute Opus 4.6 job at 43 minutes on Opus 5 for identical output, with thinking roughly 50x over the request. Shifting reasoning budgets and unprompted sandboxes and test suites read like an undisclosed effort A/B test.

  • How to tell if you're bucketed

    Give each session a fixed gauge problem and watch the thinking-token count and wall-clock against a pinned older model. A sudden 10-50x jump in reasoning for the same task is the tell that you've landed in a low-effort variant, not that the work got harder.

  • The escape hatches today

    Until Anthropic comments, builders are pinning Opus 4.6/4.8, switching on the 'Concise' output style added in Claude Code 2.1.237, or moving jobs to Codex — which several say 'feels like using Claude Code for the first time again.' No official statement yet.

  • Elsewhere: an open 33-language translator that runs local

    Tencent open-sourced Hy-MT2-30B-A3B, a fast-thinking MoE with just 3B active params covering 33 languages that fits on a single consumer GPU (1.8B and 7B siblings too). It's on Hugging Face and OpenRouter, so you can drop an offline translation layer into an app with no per-token meter.