Claude Sonnet 5 Is Here: Anthropic’s Most Agentic Sonnet Yet, at Near-Opus Performance

Claude Sonnet 5 lands with near-Opus benchmarks at Sonnet pricing. See the full SWE-bench, Terminal-Bench, and OSWorld scores vs Sonnet 4.6 and Opus 4.8.
Anthropic just shipped Claude Sonnet 5, and the benchmark gap to its own flagship Opus model has nearly closed. Here is what changed, what the numbers actually mean, and whether the switch is worth it today.
Claude Sonnet 5 launches with near-Opus performance at Sonnet pricing
Claude Sonnet 5 is Anthropic’s newest mid-tier model, built to plan tasks, operate browsers and terminals, and run autonomously at a level that used to require Opus-class compute. It posts 63.2% on SWE-bench Pro and 81.2% on OSWorld-Verified for computer use, both within single digits of Opus 4.8. It is now the default model on Free and Pro plans and live across the Claude apps and Claude Platform today.
What’s actually new in Claude Sonnet 5
Anthropic is positioning this release around one word: agentic. Not faster autocomplete, not a longer context window as the headline feature, but a model built to hold a plan together across many steps without a human nudging it forward at each one.
According to Anthropic, early access partners reported three consistent behaviors that separate Sonnet 5 from prior Sonnet releases:
- It finishes complex, multi-stage tasks where earlier Sonnet versions stalled out or asked for clarification midway through.
- It checks its own output against the original instructions without being prompted to self-review.
- It does this agentic work at a price point Anthropic is calling attractive relative to the jump in capability.
That third point matters more than it sounds. Every prior generation of frontier model has made the same capability-versus-cost tradeoff: better reasoning meant a bigger, pricier model. Sonnet 5 is the latest entry in Anthropic’s pattern of pushing flagship-adjacent intelligence into the mid-tier price band instead.
Claude Sonnet 5 benchmark scores compared to Sonnet 4.6 and Opus 4.8
This is the table that matters if you’re deciding whether to upgrade. Anthropic’s own release data shows Sonnet 5 closing most of the gap to Opus 4.8 while staying meaningfully ahead of the model it replaces.
| Benchmark | What it measures | Sonnet 5 | Sonnet 4.6 | Opus 4.8 (reference) |
|---|---|---|---|---|
| SWE-bench Pro | Agentic coding | 63.2% | 58.1% | 69.2% |
| Terminal-Bench 2.1 | Agentic coding (CLI tasks) | 80.4% | 67.0% | 82.7% |
| Humanity’s Last Exam, no tools | Multidisciplinary reasoning | 43.2% | 34.6% | 49.8% |
| Humanity’s Last Exam, with tools | Multidisciplinary reasoning | 57.4% | 46.8% | 57.9% |
| OSWorld-Verified | Computer use | 81.2% | 78.5% | 83.4% |
| GDPval-AA v2 | Knowledge work | 1618 | 1395 | 1615 |
A few things stand out reading this table straight. Sonnet 5 beats Sonnet 4.6 by a wide margin on every single metric, with the Terminal-Bench jump (67.0% to 80.4%) being the largest single gain. And on GDPval-AA v2, the knowledge work benchmark, Sonnet 5 actually edges past Opus 4.8 outright, scoring 1618 against Opus’s 1615. That is not a typo. The mid-tier model is outscoring the flagship on one of six listed evaluations.
On Humanity’s Last Exam with tools enabled, Sonnet 5’s 57.4% sits less than half a point behind Opus 4.8’s 57.9%, which is about as close as a Sonnet generation has ever gotten to its Opus sibling on a reasoning benchmark.
Is Claude Sonnet 5 worth switching to from Sonnet 4.6?
Yes, for nearly every workflow that involves multi-step agentic tasks, coding, or computer use. The benchmark gains over Sonnet 4.6 are not marginal: Terminal-Bench 2.1 jumped over 13 points and SWE-bench Pro gained 5 points, both signs of a model that holds a plan together longer before needing correction.
The harder question is whether to pay for Opus 4.8 at all anymore. With Sonnet 5 landing within roughly six points of Opus on coding benchmarks and actually surpassing it on knowledge work, the case for defaulting to Opus for routine agentic tasks gets weaker. Anthropic itself seems to be steering users this direction by making Sonnet 5 the new default on Free and Pro tiers rather than leaving Opus as the recommended starting point.
Pricing and availability
Sonnet 5 is rolling out with introductory pricing through August 31, available across the Claude apps and the Claude Platform starting today. It is the default model for Free and Pro users, and accessible to Max, Team, and Enterprise plans as well. Anthropic has not published the post-introductory rate yet, so anyone budgeting API usage past August should watch for that update directly from Anthropic rather than assuming current pricing holds.
How Claude Sonnet 5 compares to Claude Opus 4.8
The honest read of this release is that Anthropic has narrowed, not closed, the gap between its mid-tier and flagship models. Opus 4.8 still leads on SWE-bench Pro (69.2% vs 63.2%), Terminal-Bench 2.1 (82.7% vs 80.4%), and computer use (83.4% vs 81.2%). For the hardest coding and agentic workloads, Opus remains the stronger pick.
But the margin has shrunk to a point where many teams will find Sonnet 5 good enough, and at the lower price tier, that changes the default choice for day-to-day work. The pattern matches what Anthropic did with Sonnet 4.5 closing in on Opus 4.1 last year: the previous generation’s top-tier reasoning becomes the new generation’s mid-tier baseline, one release cycle later.
What this means if you’re building with Claude
If your workflow leans on agents that need to run for a while without supervision (research pipelines, coding agents, browser automation, multi-document analysis) Sonnet 5’s jump in Terminal-Bench and OSWorld scores is the number to pay attention to. Those are the two benchmarks that most directly track “can this model keep working without me stepping in.”
If your workload is closer to single-shot answers or short conversational tasks, the gains will be less visible day to day, though the Humanity’s Last Exam improvement suggests sharper reasoning even on one-off questions.
FAQ
When did Claude Sonnet 5 launch? Claude Sonnet 5 launched today across all Claude apps and the Claude Platform, with introductory pricing in effect through August 31.
Is Claude Sonnet 5 better than Opus 4.8? Not across the board. Opus 4.8 still leads on agentic coding and computer-use benchmarks, but Sonnet 5 actually scores higher on the GDPval-AA v2 knowledge work benchmark, and trails by less than a point on Humanity’s Last Exam with tools.
Is Claude Sonnet 5 the default model now? Yes, it is now the default on Free and Pro plans, and available to Max, Team, and Enterprise users.
What is the biggest improvement over Sonnet 4.6? Terminal-Bench 2.1 shows the largest jump, climbing from 67.0% to 80.4%, indicating a substantial gain in agentic command-line and tool-use tasks.
Source: Anthropic’s official Claude Sonnet 5 launch announcement.
[…] Read about Sonnet 5 here […]
[…] Claude Code, create the skill directory and pull the […]