Claude Opus 5 - Anthropic's latest flagship model
The Claude Opus 5 is Anthropic's latest flagship model, positioned as "a cutting-edge smart device close to the Fable 5, priced at about half the latter." Compared to the Opus 4.8, it offers significantly improved performance without changing costs...
What is Claude Opus 5?
The Claude Opus 5 is Anthropic's latest flagship model, positioned as "cutting-edge intelligence close to the Fable 5, but at about half the price." Compared to the Opus 4.8, it offers significantly improved performance without increasing cost, making it the default model for Claude Max and the strongest choice for Claude Pro. It introduces adjustable effort levels, allowing users to balance stronger inference with lower latency and less token consumption. It tops the charts in numerous programming, knowledge work, business automation, and computer operation benchmarks, including Frontier-Bench, GDPval-AA, ARC-AGI 3, AutomationBench, and OSWorld 2.0, only lagging behind the Mythos 5 in a few tasks such as cybersecurity.
Main functions of Claude Opus 5
-
Adjustable effort thought fileBit: Supports low, medium, high, xhigh, max and other levels. Increasing the level will result in stronger inference, while decreasing it will result in faster and more token-efficient inference, making it easier to balance quality, speed and cost according to the task.
-
Agentic programming skillsFor terminal and software engineering tasks, it scored 43.3% in Frontier-Bench v0.1, higher than Fable 5, Opus 4.8 and GPT-5.6 Sol; the official claim is that it has more than double the performance of Opus 4.8 and lower cost per task.
-
Code collaboration optimizationAt the highest setting in CursorBench 3.2, it is only 0.5% slower than Fable 5 at the peak performance, and costs about half the price; at the high, xhigh, and max settings, it outperforms other models within the same budget.
-
Knowledge Work and WritingIt covers analysis, research, document, and office tasks, with a GDPval-AA v2 score of 1861, outperforming Fable 5, GPT-5.6 Sol, and Opus 4.8.
-
Intelligent search and browsingIt possesses agentic search capabilities, scoring 90.8% on BrowseComp, slightly higher than GPT-5.6 Sol's 90.4%, making it suitable for complex data retrieval and synthesis.
-
Computer skillsIt scored 70.6% in OSWorld 2.0, higher than Fable 5 and GPT-5.6 Sol. The official claim is that it can achieve the best score of Fable 5 at about one-third the cost.
-
Business Process AutomationIt can complete commercial tasks end-to-end, with an AutomationBench pass rate of 26.0%, which is about 1.5 times that of the second-best model at the same cost; the number of tasks completed at the lowest effort level also exceeds all competitors.
-
Solving new problemsThe ARC-AGI 3 scored 30.2%, about three times that of the second-place GPT-5.6 Sol, highlighting its ability to handle tasks not seen in training.
-
Scientific research assistanceThe life sciences assessment comprehensively surpasses Opus 4.8, covering structural biology, organic chemistry, and bioinformatics; the spectral inference of molecular structure is improved by 10.2 percentage points, and the prediction of the functional impact of protein sequence variations is improved by 7.7 percentage points.
-
More proactive and more prudentThe official positioning is that it is a more proactive and thoughtful model that can reduce back-and-forth confirmation, self-checking, and recovery from errors in the task; at the same time, it emphasizes safety barriers.
-
High-frequency daily life and product integrationIt boasts higher operating efficiency and is suitable for daily use; it has become the default model for Claude Max and is also the highest capability option available on Claude Pro, with the interface providing access to settings such as "Opus 5 High".
How to use Claude Opus 5
-
Access PlatformOpen the Claude app or web version, and switch to Claude Opus 5 in the model selector.
-
Use by subscription tierClaude Max has set Opus 5 as the default model; Claude Pro users can choose it as the current highest-capability option.
-
Adjust the effort level according to the task.For simple Q&A, summarizing, and polishing, use the lower setting for faster and more economical results; for code debugging, complex analysis, and research tasks, switch to high, xhigh, or max for stronger reasoning.
-
Give clear goals and constraintsSpecify deliverables, format, word count, deadline, available tools, and acceptance criteria to reduce back-and-forth confirmations and leverage a more proactive approach.
-
Programming scenariosPaste the repository structure, error logs, API documentation, and expected behavior. The strongest approach with Opus 5 is to locate problems, modify code, run tests, and explain trade-offs.
-
Knowledge workProvide background information, target audience, and output templates, then conduct research, outline, draft, proofreading, and rewriting; GDPval-AA v2 demonstrates its leading position in knowledge work.
-
Search and ReviewThe requirement to "search first and then summarize, mark the source, and distinguish between facts and speculation" corresponds to its agentic search ability, which scores 90.8% on BrowseComp.
-
Automated processesBreaking down repetitive office processes into steps, inputs, outputs, and exception handling allows for end-to-end execution; AutomationBench shows that its pass rate at the same cost is approximately 1.5 times that of the second-best.
-
Computer operation tasksClearly define the scope of allowed software and files and the prohibited actions before allowing execution; Opus 5 scored 70.6% on OSWorld 2.0.
-
Scientific research scenarioIt is used in structural biology, organic chemistry, bioinformatics and other fields, such as spectroscopic inference of molecular structure and analysis of the functional impact of protein variation; however, it still requires expert review.
-
Controlling costsFirst, test with a low gear, then upgrade to a higher gear for difficult tasks; break long tasks down into smaller batches and prioritize reserving the maximum effort for critical and difficult tasks, because higher gears are smarter but consume more tokens.
Claude Opus 5's core advantages
-
Half price approaching flagshipThe intelligence level is close to that of Fable 5, but the price is about half that of the latter; the highest setting of CursorBench 3.2 is only 0.5% different from the peak of Fable 5, but the cost is about half.
-
More for the same priceCompared to Opus 4.8, it significantly improves performance without changing the cost. The official claim is that the performance of Frontier-Bench programming tasks is more than doubled and the cost per task is lower.
-
Most outstanding programming skillsFrontier-Bench v0.1 scored 43.3%, higher than Fable 5's 33.7%, GPT-5.6 Sol's 34.4% and Opus 4.8's 21.1%.
-
More capable with the same budgetAt the high, xhigh, and max levels, it outperforms other models at the same cost; the official narrative can be summarized as "spend half the money and accomplish nearly 99% of the work".
-
effort, elasticity scarceOne model covers multiple needs, from fast to economical to deep reasoning. The low-end model can be used frequently every day, while the high-end model tackles difficult problems, and the cost curve is more controllable.
-
Knowledge work leadingGDPval-AA v2 score of 1861, surpassing Fable 5, GPT-5.6 Sol and Opus 4.8, is suitable for research, analysis, writing and office decision-making.
-
New problem solving leapThe ARC-AGI 3 scored 30.2%, about three times that of the second-place GPT-5.6 Sol, indicating a stronger ability to generalize to unfamiliar tasks.
-
Automation and computer operation offer high cost-effectiveness: AutomationBench has a pass rate that is about 1.5 times higher than the second-best at the same cost; OSWorld 2.0 can surpass the best performance of Fable 5 at about one-third the cost.
-
Search and integration capabilities are among the best.BrowseComp has a success rate of 90.8%, slightly higher than GPT-5.6 Sol's 90.4%, making it suitable for complex searches, data comparisons, and review generation.
-
Clearly enhanced research capabilitiesThe scores in all life sciences assessments exceeded Opus 4.8, with the greatest improvement in organic chemistry; the percentage of molecular structure inference from spectroscopic inferences increased by 10.2 percentage points, and the percentage of protein variation effect prediction increased by 7.7 percentage points.
Claude Opus 5 project address
- Claude Opus 5:https://www.anthropic.com/news/claude-opus-5
Claude Opus 5 vs. Competitors
| Comparison items | Claude Opus 5 | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| position | High-performance flagship at a great price: Near-cutting-edge intelligence, suitable for high-frequency daily and professional work. | Higher-order frontier models are Opus 5's primary target for catching up with. | Strong external competitors, ranking among the top tier in multiple benchmark tests. |
| Price/Cost | It costs about half as much as the Fable 5; compared to the Opus 4.8, the cost remains the same but the performance is significantly improved. | The price is approximately twice that of the Opus 5. | The overall cost curve is weaker than Opus 5, and it lags behind in several aspects at the same cost. |
| Frontier-Bench v0.1 Terminal Programming | 43.3%, first | 33.7% | 34.4% |
| CursorBench 3.2 | The maximum setting is only 0.5% lower than the peak value of Fable 5, and costs about half. | Highest peak performance, but also highest cost. | Overall, it's lower than the Opus 5 within the same budget. |
| GDPval-AA v2 Knowledge Work | 1861, First | 1747 | 1736 |
| Solving new problems in ARC-AGI-3 | 30.2%, approximately three times that of the second-place finisher. | — | 7.8% |
| BrowseComp Smart Search | 90.8%, slightly higher than the previous year. | 87.4% | 90.4% |
| OSWorld 2.0 PC Operation | 70.6%, first place; about one-third of the cost can exceed Fable 5's best performance. | 66.1% | 62.6% |
| AutomationBench Business Processes | 26.0%, first place; at the same cost, it is about 1.5 times the second best. | 17.4% | 18.1% |
| DeepSWE v1.1 | 68.8% | 69.7% | 72.7%, ranking first in this category. |
| FrontierCode v1.1 | 53.4% | 53.5%, a slight increase; also the item where the official bolding error occurred. | 47.5% |
| HLE Multidisciplinary Reasoning | 56.3% without tools, 64.7% with tools. | Without tools: 56.5% narrow win; with tools: 63.9% win. | — |
| Legal / Health | Legal 11.7%;Health 59.8% | Legal (13.3%) is stronger; Fable is not listed in the Health table, Mythos 5 is 66.0%. | Legal 2.5%;Health 60.5% |
| More suitable | For everyday users in programming, knowledge work, automation, and scientific research who seek "near-flagship performance + cost control". | Highly demanding tasks with ample budgets and a pursuit of peak performance. | Teams already within the OpenAI ecosystem, or those that value specific strengths of DeepSWE/Health. |
Application scenarios of Claude Opus 5
-
Software Development and RefactoringSuitable for requirements clarification, code generation, bug location, test completion, refactoring explanation, and terminal command assistance; Frontier-Bench terminal programming ranks first with 43.3%, making it its strongest scenario.
-
AI programming assistants in daily life: Perform long task decomposition, cross-file modification, and regression checks in Cursor-type workflows; CursorBench's highest setting is close to Fable 5's peak but costs about half, making it suitable for high-frequency calls.
-
Enterprise Knowledge Work: Writing proposals, conducting research, compiling meeting minutes, generating reports, revising resumes/official documents/business documents; GDPval-AA v2 1861 is leading-edge and suitable for high-frequency office use.
-
In-depth search and reviewIt automatically retrieves and compares sources, summarizes conclusions, and marks uncertainties after a topic is given; BrowseComp has an accuracy of 90.8%, making it suitable for competitor research, industry research, and literature review.
-
Business Process AutomationIt enables end-to-end processes such as form processing, data transfer, approval drafts, customer follow-up, and report generation; AutomationBench achieves a pass rate approximately 1.5 times that of the second-best at the same cost.
-
Computer operation and lightweight RPAOperating software, organizing files, and batch processing spreadsheets and web pages within clearly defined permission boundaries; OSWorld 2.0 achieves 70.6% faster performance and is more cost-effective than Fable 5.
Frequently Asked Questions about Claude Opus 5
A: For maximum peak performance and a sufficient budget, choose the Fable 5; for better value and everyday high-frequency use, choose the Opus 5. At the highest setting on CursorBench, the peak performance of the Opus 5 differs from that of the Fable 5 by only 0.5%, and the cost is about half.
A: The official statement is that the cost remains the same while the performance is greatly improved; the score in the Frontier-Bench programming task increased from 21.1% to 43.3%, and it is claimed that the performance has more than doubled and the cost per task is lower.
A: Roughly speaking, Opus 5 is a flagship phone with high cost-performance, Fable 5 is a higher-end phone, and Mythos 5 is stronger in a few challenging tasks; Opus 5 is clearly inferior to Mythos 5 only in tasks such as network security.
A: Claude Max has it as the default model; Claude Pro users can choose it as the highest capability option. A gear selection option similar to "Opus 5 High" will be visible in the interface.
A: Effort is a level of thinking intensity. Higher settings make you smarter, while lower settings make you faster and save tokens. Use the low setting for simple tasks, and then upgrade to high/xhigh/max for difficult tasks such as coding, research, and automation.
A: Our strongest areas are software engineering and terminal programming; followed by knowledge work, intelligent search, business process automation, computer operation, and new problem solving. We lead in multiple benchmarks including Frontier-Bench, GDPval-AA, BrowseComp, AutomationBench, OSWorld, and ARC-AGI.