Report: builders measure a quality regression across Claude 5 models

Reproducible benchmarks show weaker nonsense-detection and ~2x verbosity vs 4.8; users also report Fable 5 silently rerouting to Opus 4.8. No official word.

Nowline AUG 15 12:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The nonsense-detection numbers

    On a public 'BullshitBench' run in the report, gen-5 models caught fabricated premises about 52% of the time versus 87% for Opus 4.6/4.8 — confidence intervals that don't overlap — and refused zero nonsense prompts. If you rely on Claude to push back on bad specs, test it before you trust it.

  • You pay for the extra tokens

    At matched reasoning effort the same tasks got wordier: Opus 5 reportedly ran ~107% more tokens and Sonnet 5 ~84% more than their 4.x predecessors, with Sonnet 5 at max effort taking ~59s on a single prompt. That's more cost and latency for the same answer.

  • The silent reroute to 4.8

    Users report Fable 5 sessions silently rerouting to Opus 4.8 mid-conversation with no disclosure — so the model label may not match what's actually serving you. Pin your model explicitly and log responses if that matters to your pipeline.

  • What to do today

    Opus 4.8 and Sonnet 4.6 are still selectable, and several developers report reverting to them because they 'feel better.' Until this settles, pin a known-good model rather than trusting the gen-5 defaults.

  • Where it stands

    The complaints reached the Hacker News front page today and span multiple open GitHub issues plus third-party 'downgrade' write-ups. Anthropic hasn't publicly responded — treat all of this as reports, not confirmed regressions.

  • Elsewhere: GitHub Copilot retires six models Sept 1

    Copilot drops Opus 4.5/4.6, Sonnet 4.5/4.6, Gemini 3.1 Pro and Raptor Mini on September 1; Sonnet 4.6 survives only for annual individual plans. Switch your workflows to a supported model before the date.