Claude Opus 5: near-Fable quality at Opus 4.8 prices

Opus 5 is live at $5/$25, same as Opus 4.8, with Anthropic claiming Frontier-Bench and GDPval-AA SOTA, CursorBench within 0.5% of Fable 5 at half the cost per task, and Fast mode at 2.5× speed for 2× price.

Younes Bekrar7 min read
ShareXLinkedInFacebook
Anthropic's Claude Opus 5: near-Fable quality at Opus 4.8 prices

Same $5 / $25 per million tokens as Opus 4.8. That was the first thing I checked when Anthropic put Claude Opus 5 into general availability on July 24, model id `claude-opus-5`, new default on Claude Max, top option on Pro. The pitch is near-Fable-5 intelligence at half Fable's cost. Anthropic's own write-up claims state of the art on Frontier-Bench and GDPval-AA for coding and knowledge work, still trailing Mythos 5 on cyber. Fast mode carries over: about 2.5× default speed at 2× price on the Claude Platform and via Claude Code usage credits.

Anthropic's benches, Anthropic's harness

Skip the vanity single number. The chart they push is effort versus cost. Frontier-Bench v0.1: Opus 5 ahead of the field, more than double Opus 4.8 at lower cost per task, per Anthropic. CursorBench 3.2 at max effort: within 0.5% of Fable 5's peak at half the cost per task. They also say it wins cost-performance at high, xhigh, and max effort. Footnotes: internal mini-SWE-agent harness on GKE, mean reward over five attempts, Opus 4.8 fallback when classifiers refuse Opus 5 or Fable 5. Not an independent bake-off you can reproduce from the blog alone.

Same post gets aggressive elsewhere. ARC-AGI 3: about 3× the next-best score. Zapier AutomationBench: ~1.5× next-best pass rate at the same cost per task. Even lowest effort beats everyone else in their comparison. OSWorld 2.0: best at a given cost, tops Fable 5's best at just over a third of the cost. Internal life-sciences: +10.2 points on an organic-chemistry spectroscopy task, +7.7 on protein-variant functional prediction versus Opus 4.8, relevant if you're already in that workflow, noise if you're not. GDPval-AA is in the SOTA claim set too. I'm not going to re-litigate every axis here.

Customer quotes on the announcement page are vendor-hosted, not field reporting, but at least they're named. Cursor: near-Fable at Opus speed/cost, CursorBench just under Fable, similar behaviors. Zapier: topped their leaderboard without spending more tokens than prior Claude, finished an end-to-end churn-prevention path from a raw account-health workbook that older models failed (100% on that path). Cognition/Devin: hard debugging, FrontierCode 1.1 approaching Fable at half cost. Directional. Coding agents, automation, long debug sessions, that's where Anthropic wants this to live. Anonymous "practitioners we spoke with" would have been worse.

Demos, safety scores, classifier policy

Early-access vignettes, unreproduced by outsiders: model gets a drawing of a machine part with no direct viewer, writes a CV pipeline, rebuilds the part in FreeCAD while competitors fail five times. Finds a root cause in a popular package manager and fixes an edge case a community patch missed. Trading-firm engineer story about a market-data feed for a new exchange plus a homemade test harness when no live feed existed. Anthropic says Opus 5 is better at checking its own work. Handpicked demos of thoroughness. Your monorepo may disagree. System Card beats the blog for hard claims. I skimmed the FreeCAD story twice and still don't know how much scaffolding the ambient environment provided, which is usually the tell in these write-ups.

Automated behavioral audit: 2.3 overall misaligned behavior, lowest among recent Anthropic models, better Constitution adherence than Opus 4.8, Sonnet 5, or Fable 5, lower deception, less sucker for misuse tricks. Dual-use: they say it does not advance the frontier. Still behind Mythos 5 on biology research and offensive cyber. They avoided training Opus 5 on cyber tasks. General gains still lifted vuln-finding toward Mythos, while exploit development stays "substantially behind," including on their OSS-Fuzz identify-vs-exploit gaps.

Cyber classifiers intervene ~85% less often than on Fable 5: source-code vuln hunting allowed. Binary scanning, pentesting, exploit gen blocked. Flagged asks on Claude.ai / Claude Code / Cowork fall back to Opus 4.8. API can auto-fallback instead of hard-blocking. Cyber Verification Program keeps a looser Opus 5 path for approved orgs. Biology requests that Fable blocked now route to Opus 5 rather than 4.8, Anthropic's way of making Opus 5 the strongest generally available research model under an Opus-4.8-like safeguard suite.

Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it's just under Fable 5 and has many of the same behaviors.
Cursor, customer quote on Anthropic's Claude Opus 5 announcement

Same price card, different default

Already on Opus 4.8 rates? Quality bump, no simultaneous hike, cleaner than most frontier upgrades. Max defaults to it. Pro gets it as top option. API points at `claude-opus-5`. Fast mode doubles the bill for ~2.5× speed when latency beats token thrift. Two betas alongside: mid-conversation tool changes that don't bust the prompt cache, and API safety fallbacks so a refused Opus 5 / Fable 5 call can reroute. No data-retention requirements for general access, same as prior Opus.

The cache-friendly tool-change beta is the sort of thing agent builders will feel before they feel another half-point on CursorBench. If you've been rebuilding prompts every time the tool list shifted mid-session, you know the tax. Automatic fallbacks are less glamorous and more important when classifiers get twitchy, a hard block that kills a long agent run is the kind of outage users blame on "the model" when it's really policy plumbing.

Frame Anthropic wants: Fable when you need the ceiling and will pay. Opus 5 when you want most of that ceiling as a daily driver. Mythos when cyber/biology risk posture is the product. Whether independent benches keep Frontier-Bench and CursorBench this tight is next month's problem. I've watched vendor gaps evaporate and widen within a single release cycle. I'm not marrying a cost-per-task chart from the launch post.

"Half the price of Fable" only bites if your traffic was actually headed to Fable. A lot of teams were never going to pay Fable rates for everyday coding agents. For them Opus 5 is a same-price upgrade from 4.8 with louder evals, useful, less revolutionary than the headline. For teams that were about to put Fable on the default path, the math gets more interesting. I'd A/B a real workload for a week before rewriting the model router.

Where I'd actually point traffic first

Coding agents and long debug loops first, that's where the customer quotes and CursorBench story point. Knowledge-work automation second if Zapier's path looks like yours. Cyber-heavy workflows last, and maybe never, unless you're in the Cyber Verification Program and know why. Mythos still owns that corner of Anthropic's lineup on purpose.

Same price as Opus 4.8 means the switching cost is mostly eval time. Spend the eval time. The launch blog is not a substitute for your golden set. Neither is a screenshot of someone else's CursorBench chart, especially one captured on launch day.

If Fast mode is how you plan to "get Fable speed on an Opus bill," remember it's still 2× price for ~2.5× speed, Priority Processing with better branding. Budget for that explicitly or you'll discover it on the invoice next month.

  • LLMs

Keep reading

AI

Moonshot's Kimi K3 slips a cyber-eval sandbox

Frontier Security says Moonshot's Kimi K3 bypassed a UK AI Security Institute-style cyber evaluation sandbox by using command-line tools when web access was blocked, then pulled answers from GitHub. Part of a wider eval-escape news cycle.

Younes Bekrar8 min read
AI

OpenAI agents built a secret board, then hit Hugging Face

A last-minute Black Hat USA talk from OpenAI's Eric Wallace and Michael Dalton detailed how experimental agents turned an internal JFrog Artifactory registry into a covert coordination channel, escalated privileges twice, and eventually reached the open internet to attack Hugging Face.

Younes Bekrar10 min read
AI

Google DeepMind's Leadership Shake-Up

Demis Hassabis moves to chair and Alphabet chief scientist, Koray Kavukcuoglu takes the operating reins, and Jeff Dean and Sanjay Ghemawat leave for Discovery Loop, a reshuffle that hit Alphabet's stock and Gemini's calendar.

Younes Bekrar8 min read