99MODELS

Claude Opus 5

New

Anthropic's flagship for demanding reasoning, coding and long-horizon agentic work.

Released Jul 24, 2026

2,640.00

per 10 lakh output tokens

Input: ₹528.00 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,00,000 tokens
Max output
1,28,000 tokens
Accepts
Text, Images, PDF files
Reasoning
On by default
Effort levels
low, medium, high, xhigh, max
Tool use
Yes
Structured output
Yes
Code execution
No
Intelligence rank
#3 of 54
Value rank
#22 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input528.00$5.50
Output2,640.00$27.50
Cached input52.80$0.55

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 93.2%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 54.9%

    HLE

    Humanity's Last Exam

  • 56.4%

    SciCode

    SciCode - scientific code generation

  • 79.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 89.1%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 53.4%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 90.4%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 68.3%

    OSWorld 2

    OSWorld 2 - agentic computer use (partial credit)

  • 1691

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About Claude Opus 5

Anthropic's model for complex agentic coding and enterprise work, described by the lab as a step change over Opus 4.8 rather than an incremental one, reaching close to Fable 5's frontier intelligence at half the price. Its largest gains are in deep reasoning, long-horizon agentic tasks and test-time compute scaling, with strong results in code review and bug finding, vision, long-context work and document tasks. Thinking is adaptive and on by default, with effort selectable from low up to max.

Anthropic released Opus 5 as the model to reach for every day rather than the one held back for the hardest hour: it became the default on Claude Max and the strongest model offered on Claude Pro on the day it shipped. The lab positions it as coming close to Claude Fable 5's frontier intelligence at half the price, and it charges exactly what Opus 4.8 charged before it, so the generational gain arrives without a rate change.

Anthropic presents almost all of its results as cost-per-task curves across the effort dial rather than as single peak scores, which is the honest way to read a model whose thinking budget the caller sets. On Frontier-Bench v0.1, its software engineering evaluation, the lab reports Opus 5 ahead of every other model and more than double Opus 4.8 at a lower cost per task. On CursorBench 3.2 at max effort it lands within 0.5% of Fable 5's peak score for half the cost per task, and beats every other model at a given cost on the high, xhigh and max rungs.

The same pattern holds away from code. Anthropic reports three times the next-best model's score on ARC-AGI 3, which scores models on problems they have not seen before; a pass rate around 1.5 times the next-best model on Zapier's AutomationBench at the same cost per task, with more tasks passed than any other model even at the lowest effort setting; and better results than any model at any cost on the OSWorld 2.0 computer-use evaluation, passing Fable 5's best score at just over a third of the cost. It is also the lab's best and most cost-efficient model on GDPval-AA v2, HLE AutomationBench and DeepSearchQA.

What early-access testers described was less a score than a working habit: the model checks itself. Anthropic's own examples include a Frontier-Bench task where the model was given a mechanical drawing and deliberately no way to view it, and answered by writing a computer vision pipeline to recover the geometry from raw pixels; a real bug in a widely used package manager where it found the root cause an accepted community patch had missed; and a market data feed built in one session, with its own test harness written because no live feed existed to validate against. On science the lab reports Opus 5 ahead of Opus 4.8 on every one of its internal life-sciences evaluations, by 10.2 points on inferring molecular structures from spectroscopy and 7.7 points on predicting how protein sequence variants behave.

Anthropic is also specific about where the model stops. It says Opus 5 does not advance the frontier in dual-use capability: on its OSS-Fuzz evaluation it comes close to Mythos 5 at finding software vulnerabilities but stays far behind at turning them into working exploits, and Mythos 5 remains the stronger model for long-running autonomous biology research. Its safety classifiers allow source-code vulnerability review while blocking binary-based scanning, penetration testing and exploit generation. On the lab's automated behavioural audit it scored 2.3 for overall misaligned behaviour, the lowest of any recent Anthropic model.

What Anthropic announced at launch

Software engineering
On Frontier-Bench v0.1 Anthropic reports Opus 5 ahead of every other model and more than double Opus 4.8's score, at a lower cost per task. On CursorBench 3.2 at max effort it lands within 0.5% of Claude Fable 5's peak for half the cost.
Problems it has not seen
On ARC-AGI 3, which tests reasoning on genuinely novel puzzles, the lab reports a score three times as high as the next-best model.
Business tasks end to end
On Zapier's AutomationBench Anthropic reports a pass rate around 1.5 times the next-best model at the same cost per task, and more tasks passed than any other model even at the lowest effort setting.
Computer use
On OSWorld 2.0 the lab reports better results than every other model at any given cost, passing Claude Fable 5's best score at just over a third of the cost.
Scientific research
Ahead of Opus 4.8 on every internal life-sciences evaluation Anthropic ran, by 10.2 points on inferring molecular structures from spectroscopy and 7.7 points on protein sequence-variant tasks.
What it is not for
Anthropic reports Opus 5 close to Mythos 5 at finding software vulnerabilities but far behind at developing exploits, and behind it on long-running autonomous biology research. Binary-based vulnerability scanning, penetration testing and exploit generation are blocked.

Indian languages

Claude Opus 5 answers in 15 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Claude Opus 5.

When was Claude Opus 5 released?

Anthropic released Claude Opus 5 on Jul 24, 2026.

Who built Claude Opus 5?

Claude Opus 5 is developed by Anthropic. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Claude Opus 5?

It is ranked 3 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Claude Opus 5 cost?

Usage costs ₹528.00 per million input tokens and ₹2,640.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Claude Opus 5 pricing in US dollars?

The provider charges $5.50 per million input tokens and $27.50 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Claude Opus 5 hold?

Its context window is 10 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Claude Opus 5 rank for value?

It ranks 22 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Claude Opus 5 support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Claude Opus 5 support?

It answers in 15 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Anthropic

Catalog updated Sep 9, 2026