Report: builders measure a quality regression across Claude 5 models
Reproducible benchmarks show weaker nonsense-detection and ~2x verbosity vs 4.8; users also report Fable 5 silently rerouting to Opus 4.8. No official word.

Copy markdown
The nonsense-detection numbers
On a public 'BullshitBench' run in the report, gen-5 models caught fabricated premises about 52% of the time versus 87% for Opus 4.6/4.8 — confidence intervals that don't overlap — and refused zero nonsense prompts. If you rely on Claude to push back on bad specs, test it before you trust it.
You pay for the extra tokens
At matched reasoning effort the same tasks got wordier: Opus 5 reportedly ran ~107% more tokens and Sonnet 5 ~84% more than their 4.x predecessors, with Sonnet 5 at max effort taking ~59s on a single prompt. That's more cost and latency for the same answer.
The silent reroute to 4.8
Users report Fable 5 sessions silently rerouting to Opus 4.8 mid-conversation with no disclosure — so the model label may not match what's actually serving you. Pin your model explicitly and log responses if that matters to your pipeline.
What to do today
Opus 4.8 and Sonnet 4.6 are still selectable, and several developers report reverting to them because they 'feel better.' Until this settles, pin a known-good model rather than trusting the gen-5 defaults.
Where it stands
The complaints reached the Hacker News front page today and span multiple open GitHub issues plus third-party 'downgrade' write-ups. Anthropic hasn't publicly responded — treat all of this as reports, not confirmed regressions.
Elsewhere: GitHub Copilot retires six models Sept 1
Copilot drops Opus 4.5/4.6, Sonnet 4.5/4.6, Gemini 3.1 Pro and Raptor Mini on September 1; Sonnet 4.6 survives only for annual individual plans. Switch your workflows to a supported model before the date.