Same $5 / $25 per million tokens as Opus 4.8. That was the first thing I checked when Anthropic put Claude Opus 5 into general availability on July 24, model id `claude-opus-5`, new default on Claude Max, top option on Pro. The pitch is near-Fable-5 intelligence at half Fable's cost. Anthropic's own write-up claims state of the art on Frontier-Bench and GDPval-AA for coding and knowledge work, still trailing Mythos 5 on cyber. Fast mode carries over: about 2.5× default speed at 2× price on the Claude Platform and via Claude Code usage credits.
Anthropic's benches, Anthropic's harness
Skip the vanity single number. The chart they push is effort versus cost. Frontier-Bench v0.1: Opus 5 ahead of the field, more than double Opus 4.8 at lower cost per task, per Anthropic. CursorBench 3.2 at max effort: within 0.5% of Fable 5's peak at half the cost per task. They also say it wins cost-performance at high, xhigh, and max effort. Footnotes: internal mini-SWE-agent harness on GKE, mean reward over five attempts, Opus 4.8 fallback when classifiers refuse Opus 5 or Fable 5. Not an independent bake-off you can reproduce from the blog alone.
Same post gets aggressive elsewhere. ARC-AGI 3: about 3× the next-best score. Zapier AutomationBench: ~1.5× next-best pass rate at the same cost per task. Even lowest effort beats everyone else in their comparison. OSWorld 2.0: best at a given cost, tops Fable 5's best at just over a third of the cost. Internal life-sciences: +10.2 points on an organic-chemistry spectroscopy task, +7.7 on protein-variant functional prediction versus Opus 4.8, relevant if you're already in that workflow, noise if you're not. GDPval-AA is in the SOTA claim set too. I'm not going to re-litigate every axis here.
Customer quotes on the announcement page are vendor-hosted, not field reporting, but at least they're named. Cursor: near-Fable at Opus speed/cost, CursorBench just under Fable, similar behaviors. Zapier: topped their leaderboard without spending more tokens than prior Claude, finished an end-to-end churn-prevention path from a raw account-health workbook that older models failed (100% on that path). Cognition/Devin: hard debugging, FrontierCode 1.1 approaching Fable at half cost. Directional. Coding agents, automation, long debug sessions, that's where Anthropic wants this to live. Anonymous "practitioners we spoke with" would have been worse.
Demos, safety scores, classifier policy
Early-access vignettes, unreproduced by outsiders: model gets a drawing of a machine part with no direct viewer, writes a CV pipeline, rebuilds the part in FreeCAD while competitors fail five times. Finds a root cause in a popular package manager and fixes an edge case a community patch missed. Trading-firm engineer story about a market-data feed for a new exchange plus a homemade test harness when no live feed existed. Anthropic says Opus 5 is better at checking its own work. Handpicked demos of thoroughness. Your monorepo may disagree. System Card beats the blog for hard claims. I skimmed the FreeCAD story twice and still don't know how much scaffolding the ambient environment provided, which is usually the tell in these write-ups.
Automated behavioral audit: 2.3 overall misaligned behavior, lowest among recent Anthropic models, better Constitution adherence than Opus 4.8, Sonnet 5, or Fable 5, lower deception, less sucker for misuse tricks. Dual-use: they say it does not advance the frontier. Still behind Mythos 5 on biology research and offensive cyber. They avoided training Opus 5 on cyber tasks. General gains still lifted vuln-finding toward Mythos, while exploit development stays "substantially behind," including on their OSS-Fuzz identify-vs-exploit gaps.
Cyber classifiers intervene ~85% less often than on Fable 5: source-code vuln hunting allowed. Binary scanning, pentesting, exploit gen blocked. Flagged asks on Claude.ai / Claude Code / Cowork fall back to Opus 4.8. API can auto-fallback instead of hard-blocking. Cyber Verification Program keeps a looser Opus 5 path for approved orgs. Biology requests that Fable blocked now route to Opus 5 rather than 4.8, Anthropic's way of making Opus 5 the strongest generally available research model under an Opus-4.8-like safeguard suite.
Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it's just under Fable 5 and has many of the same behaviors.
Same price card, different default
Already on Opus 4.8 rates? Quality bump, no simultaneous hike, cleaner than most frontier upgrades. Max defaults to it. Pro gets it as top option. API points at `claude-opus-5`. Fast mode doubles the bill for ~2.5× speed when latency beats token thrift. Two betas alongside: mid-conversation tool changes that don't bust the prompt cache, and API safety fallbacks so a refused Opus 5 / Fable 5 call can reroute. No data-retention requirements for general access, same as prior Opus.
The cache-friendly tool-change beta is the sort of thing agent builders will feel before they feel another half-point on CursorBench. If you've been rebuilding prompts every time the tool list shifted mid-session, you know the tax. Automatic fallbacks are less glamorous and more important when classifiers get twitchy, a hard block that kills a long agent run is the kind of outage users blame on "the model" when it's really policy plumbing.
Frame Anthropic wants: Fable when you need the ceiling and will pay. Opus 5 when you want most of that ceiling as a daily driver. Mythos when cyber/biology risk posture is the product. Whether independent benches keep Frontier-Bench and CursorBench this tight is next month's problem. I've watched vendor gaps evaporate and widen within a single release cycle. I'm not marrying a cost-per-task chart from the launch post.
"Half the price of Fable" only bites if your traffic was actually headed to Fable. A lot of teams were never going to pay Fable rates for everyday coding agents. For them Opus 5 is a same-price upgrade from 4.8 with louder evals, useful, less revolutionary than the headline. For teams that were about to put Fable on the default path, the math gets more interesting. I'd A/B a real workload for a week before rewriting the model router.
Where I'd actually point traffic first
Coding agents and long debug loops first, that's where the customer quotes and CursorBench story point. Knowledge-work automation second if Zapier's path looks like yours. Cyber-heavy workflows last, and maybe never, unless you're in the Cyber Verification Program and know why. Mythos still owns that corner of Anthropic's lineup on purpose.
Same price as Opus 4.8 means the switching cost is mostly eval time. Spend the eval time. The launch blog is not a substitute for your golden set. Neither is a screenshot of someone else's CursorBench chart, especially one captured on launch day.
If Fast mode is how you plan to "get Fable speed on an Opus bill," remember it's still 2× price for ~2.5× speed, Priority Processing with better branding. Budget for that explicitly or you'll discover it on the invoice next month.
- LLMs




