AI-Models

Claude Opus 5 beats Fable 5 in three out of four benchmarks

3 min read
Developer compares benchmark values and token prices of the Anthropic models Claude Opus 5 and Claude Fable 5 on a monitor Image generated with GPT Image 2
Developer compares benchmark values and token prices of the Anthropic models Claude Opus 5 and Claude Fable 5 on a monitor

TL;DR Too Long; Didn’t read

Anthropic's cheaper model Claude Opus 5 outperforms the twice as expensive flagship Fable 5 in three out of four published comparison tests, including knowledge work at 1861 to 1747 points. Anthropic explains the price gap not with weaker performance but with the duration of autonomous work: Fable 5 remains the recommendation for multi-day projects.

Key takeaways

  • In knowledge work, Opus 5 scores 1861 points on the GDPval-AA v2 test, while Fable 5 reaches 1747.
  • On the programming benchmark DeepSWE v1.1, Fable 5 stays narrowly ahead at 69.7 to 68.8 percent.
  • Opus 5 generates roughly 26 percent fewer tokens at maximum thinking depth than Opus 4.8.
  • Fable 5 handles up to one million tokens of context and works autonomously for several days.
  • Heavy users on $200 plans consumed computing power worth up to $5,000 with Fable 5.
  • Anthropic publishes no details on the architecture or training method behind Opus 5.

Anthropic has been selling two frontier models side by side since July 24 – and the cheaper one performs better in most published comparison tests. Claude Opus 5 costs half as much as Fable 5 and surpasses it in knowledge work, automation and terminal tasks. Only for days-long autonomous work does Fable 5 remain the vendor’s recommendation.

Opus 5 wins three of the four published comparisons

The figures compiled by the trade publication The New Stack paint a clear picture. On the GDPval-AA v2 test for knowledge work, Opus 5 scores 1861 points against 1747 for Fable 5. On agentic terminal tasks, Opus 5 leads at 43.3 to 33.7 percent, and on business workflows in the AutomationBench test at 26.0 to 17.4 percent. Only on the coding benchmark DeepSWE v1.1 does Fable 5 narrowly hold the upper hand, at 69.7 to 68.8 percent.

Anthropic itself puts it more cautiously. Its own announcement states that Opus 5 comes close to the frontier intelligence of Fable 5 at half the price. On the coding test CursorBench it lands within 0.5 percentage points of Fable 5’s peak score, at half the cost per task. Against its predecessor Opus 4.8, the company also reports three times the score on the problem-solving test ARC-AGI-3. These values come largely from Anthropic’s own measurements and are independently unverified. A company spokesperson described Opus 5 to Axios as the best-aligned model in the Opus line and the least susceptible to being tricked into misuse.

The Fable 5 premium buys endurance, not raw scores

By the company’s own account, the difference lies not in raw capability but in stamina. Anthropic continues to recommend Fable 5 for the most ambitious projects involving multi-day autonomous work. Opus 5, by contrast, is meant to be the model people reach for every day, especially inside companies. Axios describes the split along the same lines: an everyday model for enterprises, knowledge workers and developers on one side, the go-to for the most complex long-running tasks on the other.

The technical grounding for that split sits in the product documentation. Fable 5 is built for long-horizon agentic work, handles up to one million tokens of context and produces up to 128,000 output tokens per request. Such runs split a task into many sub-steps, delegate to sub-agents and check their own intermediate results, as the business magazine Forbes describes. A single request can stretch across hours or days and inflate token consumption accordingly. That sustained load is exactly what the Fable 5 premium is meant to cover – and equally the reason most everyday tasks never need it.

Leaner token use lowers the cost per task

Anthropic explains the lower price mainly through operational thrift. Opus 5 reportedly generates 26 percent fewer tokens on average than Opus 4.8 at maximum reasoning depth while reaching comparable results. On a trading benchmark it needs roughly one seventh of its predecessor’s reasoning tokens. On top of that come the new effort levels, which let users throttle the compute spent per task themselves. Since token prices remain unchanged at five and 25 dollars per million, every saved round of reasoning cuts the bill directly.

Fable 5, by contrast, is expensive to run. When Anthropic moved the model to usage-based billing in July, operating costs outgrew the maths behind its flat-rate plans: heavy users on 200-dollar subscriptions burned computing power worth up to 5,000 dollars, according to estimates reported by Silicon Report. The newer Fable 5 tokenizer counts roughly 30 percent more tokens for identical text, pushing effective costs higher still. Since July 20, access has been permanently tied to the Max and Team Premium plans, and there at a halved usage limit.

What remains open is how Anthropic engineers the gap in the first place. The company discloses nothing about the architecture, model size or training method behind Opus 5, neither in the announcement nor in the developer documentation. As long as that gap persists, there is no way to judge whether the half price holds up over time or is a capacity-driven snapshot.

Frequently asked questions

Which of the two models should be chosen for programming tasks?

For individual, well-defined tasks, the published figures suggest Opus 5 suffices, as it lands almost level on the CursorBench test. For projects a model is meant to carry forward independently over days, Anthropic still recommends Fable 5.

Does the price difference mean Opus 5 is a smaller model?

That cannot be determined from the published material. Anthropic states neither model size nor architecture and justifies the price solely through operational efficiency.

How do the effort levels affect actual costs?

Token prices stay unchanged, but a lower level produces fewer thinking tokens per task. The saving therefore comes from the volume of billed tokens, not from a discount.

Can Opus 5 also handle the one million token context?

Anthropic explicitly documents the one million token context length for Fable 5 and Mythos 5. No corresponding figure appears for Opus 5 in the product documentation.

Will Fable 5 be replaced by Opus 5 in the medium term?

Anthropic has issued no discontinuation notice. Fable 5 remains available in the Max and Team Premium plans and is still recommended for multi-day autonomous projects.


← Back to the blog