Claude Sonnet 5 vs Opus 4.8: Is the Cheaper Model Good Enough Now?
Sonnet 5 costs a fraction of Opus 4.8 and lands close to it on agentic benchmarks. Where the gap actually closes, and where it doesn't.

We covered Sonnet 5’s launch as a footnote to a bigger story about export controls. It deserves its own look, because the actual question teams are asking is narrower and more practical than “what’s Anthropic’s new model”: should you keep defaulting to Opus 4.8, or is Sonnet 5 now good enough to be the model you reach for first?
The price gap is the whole reason this question exists
Anthropic’s own pricing page lays the three current models side by side. Opus 4.8 runs $5 per million input tokens and $25 per million output tokens. Sonnet 5 is $2 and $10 through an introductory window that runs until the end of August, rising to $3 and $15 afterward, still a fraction of Opus 4.8’s rate on either side of that date. Fable 5, the frontier tier, sits above both at $10 and $50.
That gap is the entire reason this comparison matters. A model that costs a fifth to a third of another one isn’t interesting because it’s cheaper in the abstract; it’s interesting because Anthropic is explicitly positioning it as close enough on capability that the price difference stops being a rounding error and starts being the deciding factor. The question isn’t which model scores higher. It’s whether Sonnet 5 closes the gap enough that paying triple for Opus 4.8 needs its own justification now, instead of being the default nobody questions.
Effort is the lever most people aren’t using yet
Both models ship with an effort setting (low, medium, high, extra-high) that trades cost and latency for how hard the model works on a given answer. Most people leave it on medium and never touch it, which means most comparisons between Sonnet 5 and Opus 4.8 are actually comparisons of Sonnet 5’s default against Opus 4.8’s default, not of what either model can do at its ceiling.
That matters here specifically because Sonnet 5 at high or extra-high effort is still, on a per-task basis, cheaper than Opus 4.8 at medium, and on Anthropic’s own agentic benchmarks, that’s where the pass-rate curves for the two models sit closest together. The honest move before reaching for Opus 4.8 on a task Sonnet 5 struggled with is to push the effort slider up first, not the model. Only once Sonnet 5 at its highest effort setting still comes up short does switching models actually buy you something the effort setting couldn’t.
A real build, side by side
Numbers on a pricing page are one thing; a real task run twice is another. Duncan Rogoff, who runs the YouTube channel Learn Claude Code and sells a paid Claude Code skill for building product sites, ran exactly that test: the same 3D product website, built from the same product photo, once with Sonnet 5 and once with Opus 4.8, both left to run unattended.
Sonnet 5 finished in 2 minutes 28 seconds for $1.77. Opus 4.8 finished in 4 minutes 19 seconds for $34, almost two minutes slower and roughly nineteen times the cost for the same brief. Rogoff’s own read on the output was that Opus 4.8’s site looked “a little bit cleaner and a little bit more consistent,” which is a real difference, but a marginal one next to a twentyfold price gap. Treat this as one creator’s own test on one task, not an independent benchmark: the channel sells a Claude Code skill, so the numbers are demonstrating a product, not auditing one. But the shape of the result, a large cost gap next to a small quality gap on a genuinely multi-step agentic task, lines up with what Anthropic’s own benchmark charts show.
Where Opus 4.8 still earns the extra cost
None of this means Opus 4.8 is now pointless. Its edge shows up specifically in long, autonomous, multi-step work: large refactors, agents that hold context across many turns, anything where a mistake early in a long unattended run compounds instead of getting caught. That’s a narrower slice of most teams’ actual workload than the marketing around “agentic AI” suggests, but it’s a real one, and it’s the slice where the highest Sonnet 5 effort setting can still fall short.
There’s also a quieter case for Opus 4.8: consistency you’ve already measured in production is worth something a cheaper model hasn’t earned yet. If Opus 4.8 is already tuned into a workflow that works, the savings from switching have to clear the cost of re-testing that workflow on a different model, not just the sticker price.
How to actually decide
Start every new task on Sonnet 5. If it’s a short, well-scoped job, leave the effort on medium and move on. If it’s a harder multi-step task, push the effort slider to high or extra-high before you touch the model selector: that’s the cheaper lever, and on Anthropic’s own numbers it’s the one that closes most of the gap with Opus 4.8. Reach for Opus 4.8 only once Sonnet 5 at its highest effort setting has actually been tried and actually came up short, not out of habit because it used to be the only model capable of the job.
Frequently asked questions
Is Sonnet 5 as good as Opus 4.8?
On short, well-scoped tasks, close enough that reviewers often can’t reliably tell the outputs apart. On long, autonomous, multi-step work, Opus 4.8 still has a real edge, though pushing Sonnet 5’s effort setting to high or extra-high narrows that gap significantly before you need to switch models at all.
How much cheaper is Sonnet 5 than Opus 4.8?
Through the end of August, Sonnet 5 is $2 per million input tokens and $10 per million output tokens against Opus 4.8’s $5 and $25, roughly a fifth of the cost on both sides. After the introductory window ends, Sonnet 5 rises to $3 and $15, still well under half of Opus 4.8’s rate.
What does the effort setting actually change?
It controls how much the model reasons before answering: low, medium, high, or extra-high. Higher effort costs more and answers more slowly, but closes most of the capability gap with a pricier model on the same task. It’s worth trying before switching models on a task that didn’t go well.
Should I switch my whole workflow from Opus 4.8 to Sonnet 5?
Test it on a slice of real work first, with the effort setting turned up on the harder tasks in that slice. If your work is mostly short and well-scoped, the switch likely pays off immediately. If it leans on long autonomous runs or large multi-step jobs, keep Opus 4.8 there and move only the rest.