99MODELS

Inkling 975B

نیا

Open-weight multimodal MoE that reasons natively over text, images and audio.

تاریخ اجراء: 17 جولائی، 2026

388.80

فی 10 لاکھ آؤٹ پٹ Tokens

ان پٹ: ₹96.00 فی 10 لاکھ Tokens

فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔

تکنیکی تفصیلات

Context ونڈو
10,48,576 Tokens
زیادہ سے زیادہ آؤٹ پٹ
2,62,144 Tokens
سپورٹ
ٹیکسٹ, تصاویر, آڈیو
Reasoning
پہلے سے آن
ایفرٹ لیولز
none, minimal, low, medium, high, max
ٹول کا استعمال
ہاں
سٹرکچرڈ آؤٹ پٹ
نہیں
کوڈ ایگزیکیوشن
نہیں
پیرامیٹرز
975B A41B MoE
ذہانت کا درجہ
54 میں سے #37
ویلیو کا درجہ
54 میں سے #38

قیمتیں

قیمتیں
فی 10 lakh TokensINRUSD
ان پٹ96.00$1.00
آؤٹ پٹ388.80$4.05
کیشڈ ان پٹ16.32$0.17

Benchmarks

تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔

  • 87.2%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 31.9%

    HLE

    Humanity's Last Exam

  • 47.0%

    SciCode

    SciCode - scientific code generation

  • 77.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 55.1%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 14.0%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 36.5%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1408

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

Inkling 975B کے بارے میں

Thinking Machines' first model: an open-weights Mixture-of-Experts transformer with 975B total and 41B active parameters, pretrained on 45 trillion tokens of text, images, audio and video. It reasons natively over all three input modalities through an encoder-free architecture that embeds image patches and audio spectrograms alongside text tokens, so it can transcribe speech, follow spoken instructions and reason over recordings without a separate pipeline. Thinking effort is continuously controllable rather than stepped, and the lab's case for it is token efficiency: reaching a given score with fewer tokens than comparable models. Thinking Machines is direct that this is a base to customise rather than a frontier model, saying it is not the strongest model available today, open or closed.

Inkling is the first model Thinking Machines has trained itself, and the post announcing it spends most of its length explaining a choice rather than a score. The lab did not set out to win a leaderboard. It set out to publish a base that other people can take apart: full weights, a broad and deliberately balanced capability profile, and a fine-tuning path that exists on day one rather than as a promise. The argument behind that is a claim about where the remaining value sits - that many real problems are not solved well by even the best generalist model, and that the gap is closed by fine-tuning on knowledge a lab does not have and cannot buy.

The architecture follows from it. Inkling is a Mixture-of-Experts transformer of 975B total parameters with 41B active, 256 routed experts per layer and six active per token, using sigmoid routing, relative positional embeddings rather than rotary ones, and alternating sliding-window and global attention. It was pretrained on 45 trillion tokens spanning text, images, audio and video, and it carries a context window of up to a million tokens. None of that is arranged around a single headline capability; it is arranged around being a reasonable starting point for a lot of different endings.

Multimodality here is architectural rather than adjacent. Thinking Machines describes an encoder-free design: images arrive as 40x40 pixel patches through a four-layer hMLP, audio as dMel spectrograms through a light embedding layer, and both are then processed jointly with text tokens in the same stack. The practical consequence is that audio is a first-class input, not a transcription step bolted to the front - the model can transcribe speech, follow spoken instructions, answer questions about a recording and reason over longer-form audio without a separate pipeline in between.

The efficiency claim is the one the lab makes hardest, and it is about tokens rather than latency. Thinking effort on Inkling is continuous rather than stepped, set as a value between 0.2 and 0.99, and the lab's charts are drawn as score against tokens spent instead of score alone. On that framing it reports Inkling reaching a given score at fewer tokens than the models it compares against, matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. Its reported figures include 77.6% on SWEBench Verified, 63.8% on Terminal Bench 2.1, 87.2% on GPQA Diamond, 97.1% on AIME 2026, and on the multimodal side 73.5% on MMMU Pro and 91.4% on VoiceBench.

Thinking Machines is unusually direct about the ceiling. Its own words are that Inkling "is not the strongest overall model available today, open or closed", and that this is a first release in a family it intends to keep building on - it expects the multimodal side in particular to improve. Read that as the shape of the offer: reach for Inkling when you want a capable, cheap, genuinely multimodal model you can inspect, host or specialise, and reach for something else when you want the highest possible score on a task you are not going to fine-tune for.

لانچ کے وقت Thinking Machines کا بیان

Built to be customised
Thinking Machines frames Inkling as a base rather than a finished product - multimodal, efficient at thinking, and available for fine-tuning on its own Tinker platform from launch day, on the argument that specialised knowledge is what closes the remaining gap.
Open weights, 975B total
A Mixture-of-Experts transformer with 975B total and 41B active parameters, 256 experts per layer with six active per token, pretrained on 45 trillion tokens of text, images, audio and video.
Encoder-free multimodality
Images enter as 40x40 patches through a four-layer hMLP and audio as dMel spectrograms, both embedded and processed jointly with text tokens rather than through a separate perception model.
Audio as a real input
It transcribes speech, follows spoken instructions, answers questions about recordings and reasons over longer-form audio - capabilities the lab reports directly rather than through a transcription step.
Effort as a dial, not a ladder
Thinking effort is continuous between 0.2 and 0.99, and the lab measures itself on score per token: it reports matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens.
What it is not for
Thinking Machines states plainly that Inkling is not the strongest overall model available today, open or closed, and calls it the first of a family. If you need peak capability on a task you will not fine-tune for, this is not that model.

Frequently Asked Questions

Inkling 975B کے بارے میں اکثر پوچھے جانے والے سوالات۔

Inkling 975B کب جاری ہوا تھا؟

Thinking Machines نے Inkling 975B کو 17 جولائی، 2026 کو جاری کیا۔

Inkling 975B کس نے بنایا ہے؟

Inkling 975B کو Thinking Machines نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔

Inkling 975B کتنا ذہین ہے؟

یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 37 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔

Inkling 975B کا کتنا خرچ آتا ہے؟

استعمال کی لاگت ₹96.00 فی 10 لاکھ ان پٹ Tokens اور ₹388.80 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔

امریکی ڈالر میں Inkling 975B کی قیمت کیا ہے؟

فراہم کنندہ $1.00 فی 10 لاکھ ان پٹ Tokens اور $4.05 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔

Inkling 975B کتنی لمبی گفتگو یاد رکھ سکتا ہے؟

اس کی Context حد 10.5 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔

کیا Inkling 975B مناسب قیمت میں بہترین کارکردگی دیتا ہے؟

یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 38 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔

کیا Inkling 975B جواب دینے سے پہلے سوچتا ہے؟

پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔

کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026