99MODELS

کلوڈ اوپس 5

نیا

Anthropic's flagship for demanding reasoning, coding and long-horizon agentic work.

تاریخ اجراء: 24 جولائی، 2026

2,640.00

فی 10 لاکھ آؤٹ پٹ Tokens

ان پٹ: ₹528.00 فی 10 لاکھ Tokens

فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔

تکنیکی تفصیلات

Context ونڈو
10,00,000 Tokens
زیادہ سے زیادہ آؤٹ پٹ
1,28,000 Tokens
سپورٹ
ٹیکسٹ, تصاویر, PDF فائلیں
Reasoning
پہلے سے آن
ایفرٹ لیولز
low, medium, high, xhigh, max
ٹول کا استعمال
ہاں
سٹرکچرڈ آؤٹ پٹ
ہاں
کوڈ ایگزیکیوشن
نہیں
ذہانت کا درجہ
54 میں سے #3
ویلیو کا درجہ
54 میں سے #22

قیمتیں

قیمتیں
فی 10 lakh TokensINRUSD
ان پٹ528.00$5.50
آؤٹ پٹ2,640.00$27.50
کیشڈ ان پٹ52.80$0.55

Benchmarks

تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔

  • 93.2%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 54.9%

    HLE

    Humanity's Last Exam

  • 56.4%

    SciCode

    SciCode - scientific code generation

  • 79.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 89.1%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 53.4%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 90.4%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 68.3%

    OSWorld 2

    OSWorld 2 - agentic computer use (partial credit)

  • 1691

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

کلوڈ اوپس 5 کے بارے میں

Anthropic's model for complex agentic coding and enterprise work, described by the lab as a step change over Opus 4.8 rather than an incremental one, reaching close to Fable 5's frontier intelligence at half the price. Its largest gains are in deep reasoning, long-horizon agentic tasks and test-time compute scaling, with strong results in code review and bug finding, vision, long-context work and document tasks. Thinking is adaptive and on by default, with effort selectable from low up to max.

Anthropic released Opus 5 as the model to reach for every day rather than the one held back for the hardest hour: it became the default on Claude Max and the strongest model offered on Claude Pro on the day it shipped. The lab positions it as coming close to Claude Fable 5's frontier intelligence at half the price, and it charges exactly what Opus 4.8 charged before it, so the generational gain arrives without a rate change.

Anthropic presents almost all of its results as cost-per-task curves across the effort dial rather than as single peak scores, which is the honest way to read a model whose thinking budget the caller sets. On Frontier-Bench v0.1, its software engineering evaluation, the lab reports Opus 5 ahead of every other model and more than double Opus 4.8 at a lower cost per task. On CursorBench 3.2 at max effort it lands within 0.5% of Fable 5's peak score for half the cost per task, and beats every other model at a given cost on the high, xhigh and max rungs.

The same pattern holds away from code. Anthropic reports three times the next-best model's score on ARC-AGI 3, which scores models on problems they have not seen before; a pass rate around 1.5 times the next-best model on Zapier's AutomationBench at the same cost per task, with more tasks passed than any other model even at the lowest effort setting; and better results than any model at any cost on the OSWorld 2.0 computer-use evaluation, passing Fable 5's best score at just over a third of the cost. It is also the lab's best and most cost-efficient model on GDPval-AA v2, HLE AutomationBench and DeepSearchQA.

What early-access testers described was less a score than a working habit: the model checks itself. Anthropic's own examples include a Frontier-Bench task where the model was given a mechanical drawing and deliberately no way to view it, and answered by writing a computer vision pipeline to recover the geometry from raw pixels; a real bug in a widely used package manager where it found the root cause an accepted community patch had missed; and a market data feed built in one session, with its own test harness written because no live feed existed to validate against. On science the lab reports Opus 5 ahead of Opus 4.8 on every one of its internal life-sciences evaluations, by 10.2 points on inferring molecular structures from spectroscopy and 7.7 points on predicting how protein sequence variants behave.

Anthropic is also specific about where the model stops. It says Opus 5 does not advance the frontier in dual-use capability: on its OSS-Fuzz evaluation it comes close to Mythos 5 at finding software vulnerabilities but stays far behind at turning them into working exploits, and Mythos 5 remains the stronger model for long-running autonomous biology research. Its safety classifiers allow source-code vulnerability review while blocking binary-based scanning, penetration testing and exploit generation. On the lab's automated behavioural audit it scored 2.3 for overall misaligned behaviour, the lowest of any recent Anthropic model.

لانچ کے وقت Anthropic کا بیان

Software engineering
On Frontier-Bench v0.1 Anthropic reports Opus 5 ahead of every other model and more than double Opus 4.8's score, at a lower cost per task. On CursorBench 3.2 at max effort it lands within 0.5% of Claude Fable 5's peak for half the cost.
Problems it has not seen
On ARC-AGI 3, which tests reasoning on genuinely novel puzzles, the lab reports a score three times as high as the next-best model.
Business tasks end to end
On Zapier's AutomationBench Anthropic reports a pass rate around 1.5 times the next-best model at the same cost per task, and more tasks passed than any other model even at the lowest effort setting.
Computer use
On OSWorld 2.0 the lab reports better results than every other model at any given cost, passing Claude Fable 5's best score at just over a third of the cost.
Scientific research
Ahead of Opus 4.8 on every internal life-sciences evaluation Anthropic ran, by 10.2 points on inferring molecular structures from spectroscopy and 7.7 points on protein sequence-variant tasks.
What it is not for
Anthropic reports Opus 5 close to Mythos 5 at finding software vulnerabilities but far behind at developing exploits, and behind it on long-running autonomous biology research. Binary-based vulnerability scanning, penetration testing and exploit generation are blocked.

ہندوستانی زبانیں

کلوڈ اوپس 5 15 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے زبان کا انتخاب کریں۔

Frequently Asked Questions

کلوڈ اوپس 5 کے بارے میں اکثر پوچھے جانے والے سوالات۔

کلوڈ اوپس 5 کب جاری ہوا تھا؟

Anthropic نے کلوڈ اوپس 5 کو 24 جولائی، 2026 کو جاری کیا۔

کلوڈ اوپس 5 کس نے بنایا ہے؟

کلوڈ اوپس 5 کو Anthropic نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔

کلوڈ اوپس 5 کتنا ذہین ہے؟

یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 3 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔

کلوڈ اوپس 5 کا کتنا خرچ آتا ہے؟

استعمال کی لاگت ₹528.00 فی 10 لاکھ ان پٹ Tokens اور ₹2,640.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔

امریکی ڈالر میں کلوڈ اوپس 5 کی قیمت کیا ہے؟

فراہم کنندہ $5.50 فی 10 لاکھ ان پٹ Tokens اور $27.50 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔

کلوڈ اوپس 5 کتنی لمبی گفتگو یاد رکھ سکتا ہے؟

اس کی Context حد 10 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔

کیا کلوڈ اوپس 5 مناسب قیمت میں بہترین کارکردگی دیتا ہے؟

یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 22 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔

کیا کلوڈ اوپس 5 جواب دینے سے پہلے سوچتا ہے؟

پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔

کلوڈ اوپس 5 کن ہندوستانی زبانوں میں جواب دیتا ہے؟

یہ 15 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے اپنی زبان منتخب کریں۔

Anthropic کے مزید ماڈلز

کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026