Claude Opus 5 Arrives: Half the Price, Near-Fable 5 Performance
Lead
Anthropic’s Claude Opus 5 has officially launched, and the immediate story is easy to grasp: it is much cheaper than Fable 5 while matching or beating it in several hard evaluations. But the more interesting point is not simply price-performance. Across community tests and Anthropic’s own examples, Opus 5 appears to be better at organizing work, building verification loops and checking whether its outputs actually meet the target.
Key points
- Same Opus price, stronger results. Opus 5 remains priced at $5 per million input tokens and $25 per million output tokens, about half of Fable 5’s price. On Frontier-Bench v0.1, it scored 43.3%, ahead of Fable 5’s 33.7% and more than double Opus 4.8’s 21.1%.
- Developer demos show stronger code generation. Users tested it on a Rocket League-style clone, skiing scenes, Minecraft-like worlds, a single-file HTML street evolving from 1945 to 2055, Apollo launch visuals and physics-heavy destruction tasks. The common theme was not just visual polish, but more coherent interaction and simulation.
- Self-verification is the standout behavior. In one Boeing 747 browser-based 3D reconstruction test, Fable 5 delivered faster, but Opus 5 created more modules and added measurement scripts. It extracted orthographic outlines from rendered images and compared wingspan, wheelbase, sweep angle and other metrics with public Boeing 747-400 data.
- Claude Code now needs less scaffolding. Anthropic says it removed more than 80% of Claude Code’s system prompt for the Opus 5 and Fable 5 generation without lowering coding benchmark scores. A rigid rule such as “do not write comments by default” has been replaced by a more flexible instruction to match the surrounding code style.
Why it matters
The release suggests a broader transition from prompt-following models to models that can manage a process. If a model can decide when to call tools, when to store memory and when to build a test harness, developers do not need to encode every behavior as a long system prompt.
For coding AI, that is a meaningful step. The next competition will not be only about who can generate the flashiest web demo. It will be about who can identify missing validation, build reliable checks and prove that an output matches real requirements. Opus 5 still needs more independent testing, but the examples in the source point to a model race increasingly centered on engineering autonomy.
Source: QbitAI
Comments
Checking sign-in status...
Loading comments...