Skip to content
research

Claude's Civil War: Opus 5 Dethrones Fable

Anthropic's new Opus 5 model isn't just an upgrade; it's a strategic coup that redefines the AI cost-performance curve. But its most shocking feature is what they deliberately removed to make it even smarter.

Aki Tanaka
Claude's Civil War: Opus 5 Dethrones Fable

The Benchmark Bloodbath

Anthropic's new Opus 5 has dramatically reshaped the AI landscape, unexpectedly outperforming its flagship model, Fable 5, across nearly every critical benchmark. This successor to a previous mid-tier offering now sets a new standard, challenging the established hierarchy within Claude's model family.

The performance gains are stark. Opus 5 achieved a 10-point jump on Frontier Bench for agentic terminal coding, scoring 43 compared to Fable 5's 33. On OpenAI's GDP Val, a benchmark for real-world practical tasks, Opus 5 delivered a stunning 100-point improvement. It also registered a massive leap on Arc AGI 3, reaching 30% from a previous high of around 8%.

Further reinforcing its prowess, Opus 5 showed a 4-point improvement on OS World for computer use and a 9-point increase on Automation Bench. It also edged out Fable 5 on Browse Comp (90% versus 87%), demonstrating broad capabilities.

However, Opus 5 exhibits a specialized capability profile. It saw a slight drop on the Deep Swee benchmark and a decrease on the legal benchmark (from 13.3 to 11.7). Health Bench Professional also showed a decline, suggesting that while Opus 5 excels in many areas, its strengths are not uniformly distributed across all domains.

The New Bottom Line: Cost-Per-Task

Price-per-token, long a simplistic metric for AI utility, offers a misleading glimpse into true operational cost. The real determinant of value is cost-per-task, a holistic measure accounting for both token price and model efficiency. A model half the price per token might consume twice as many, making the actual cost a wash, as seen with Kimmy K3's performance against competitors.

Analysis of performance-versus-cost charts, such as those for OS World and Automation Bench, clearly illustrates Opus 5’s dominance. On OS World, where the Y-axis tracks total score and the X-axis measures cost-per-task, Opus 5 consistently plots higher and further to the left. This indicates both superior performance and lower cost compared to Fable 5 and Opus 4.8.

Even against GPT 5.6 Sol, Opus 5 demonstrates remarkable economic efficiency; matching Opus 5’s performance can cost GPT 5.6 Sol users over twice as much. This translates into tangible benefits: users gain a more capable model that completes tasks more reliably for the same or less money. Opus 5 represents a significant leap in AI economic viability, moving beyond mere benchmark scores to deliver demonstrable value.

A Quantum Leap in Reasoning

Opus 5’s most striking advance lies in its raw reasoning ability, evidenced by an unprecedented 30% score on the Arc AGI 3 benchmark. This represents a monumental leap from the previous best of approximately 8%. The Arc AGI 3 test measures pure, zero-shot problem-solving: models receive only a task environment, without instructions or context, challenging them to infer rules and achieve goals autonomously, much like a human encountering a completely novel puzzle.

This fundamental leap in reasoning directly translates to enhanced capabilities for complex enterprise applications. Sponsor Box’s internal findings confirm Opus 5’s significant gains in demanding, multi-step knowledge work, outperforming prior models. The model demonstrates superior performance in critical analytical tasks such as:

  • Due diligence
  • Data analysis
  • Report drafting from data

These exhaustive analyses are crucial for driving robust, real-world business decisions across diverse industries.

A model capable of solving novel problems is a foundational step toward more autonomous AI agents. Opus 5’s improved scores on benchmarks like OS World, which assesses an AI's ability to control a computer and interact with its interface—identifying elements, clicking buttons, and executing tasks—signal its readiness for sophisticated agentic workflows. This capability enables AI to operate tools and computers effectively, moving beyond simple task execution to genuine problem-solving in dynamic environments. For more details on its capabilities, see Introducing Claude Opus 5 - Anthropic.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Smarter Through Subtraction

Opus 5’s most counter-intuitive design decision involved a surgical nerf: its overall intelligence increased despite a deliberate reduction in cybersecurity exploit capabilities. While Opus 5 remains "substantially behind Mythos 5 at developing exploits," it paradoxically outperforms its predecessor, Opus 4.8, on general cybersecurity tasks. This suggests a strategic pruning, where limiting potentially harmful or tangential capacities allowed for a deeper, more refined focus on core reasoning.

This selective subtraction challenges the long-held belief that adding safety guardrails inevitably degrades a model's general performance. Instead, it hints at a more sophisticated approach to AI alignment, where targeted constraints can, much like removing distractions, actually foster a net gain in beneficial intelligence by allowing resources to be reallocated to more productive cognitive pathways.

Anthropic further augmented this safety-first philosophy with new developer-focused features designed for robust production use. These include automatic API fallbacks for safety-flagged requests, ensuring seamless operation even when sensitive content is detected. Furthermore, the ability to implement mid-conversation tool changes enhances the model's adaptability and reliability, providing developers with greater control and stability in their applications.

Frequently Asked Questions

What is Claude Opus 5?

Claude Opus 5 is the latest top-tier model from Anthropic. Surprisingly, it outperforms their previous flagship model, Fable 5, on most benchmarks while maintaining the same price as its predecessor, Opus 4.8.

How is Opus 5 better than Fable 5?

Opus 5 shows significant gains over Fable 5 in areas like agentic coding (Frontier Bench), real-world task completion (GDP Val), and novel problem-solving (Arc AGI 3). It also achieves this at a lower cost-per-task, making it more efficient.

What does 'cost per task' mean for AI models?

Cost per task is the true price of getting a job done, factoring in both the token price and the number of tokens required. A model can have cheap tokens but be expensive if it's inefficient. Opus 5 excels by being highly effective, thus reducing the overall cost per successful task.

Is Opus 5 good at everything?

No. While broadly more capable, Opus 5 shows slight performance decreases on certain specialized benchmarks, including legal (Legal Bench) and professional health (Health Bench Professional). It was also intentionally designed to be less capable at cybersecurity exploit development.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only