Inkling 975B
نیاOpen-weight multimodal MoE that reasons natively over text, images and audio.
تاریخ اجراء: 17 جولائی، 2026
₹388.80
فی 10 لاکھ آؤٹ پٹ Tokens
ان پٹ: ₹96.00 فی 10 لاکھ Tokens
فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔
تکنیکی تفصیلات
- Context ونڈو
- 10,48,576 Tokens
- زیادہ سے زیادہ آؤٹ پٹ
- 2,62,144 Tokens
- سپورٹ
- ٹیکسٹ, تصاویر, آڈیو
- Reasoning
- پہلے سے آن
- ایفرٹ لیولز
- none, minimal, low, medium, high, max
- ٹول کا استعمال
- ہاں
- سٹرکچرڈ آؤٹ پٹ
- نہیں
- کوڈ ایگزیکیوشن
- نہیں
- پیرامیٹرز
- 975B A41B MoE
- ذہانت کا درجہ
- 54 میں سے #37
- ویلیو کا درجہ
- 54 میں سے #38
قیمتیں
Benchmarks
تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔
87.2%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
31.9%
HLE
Humanity's Last Exam
47.0%
SciCode
SciCode - scientific code generation
77.3%
Long Context
Long Context Reasoning - reasoning over long inputs
55.1%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
14.0%
FrontierCode
FrontierCode - long-horizon production coding tasks
36.5%
ARC-AGI-2
ARC-AGI-2 - abstract reasoning on novel puzzles
1408
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
Inkling 975B کے بارے میں
Thinking Machines' first model: an open-weights Mixture-of-Experts transformer with 975B total and 41B active parameters, pretrained on 45 trillion tokens of text, images, audio and video. It reasons natively over all three input modalities through an encoder-free architecture that embeds image patches and audio spectrograms alongside text tokens, so it can transcribe speech, follow spoken instructions and reason over recordings without a separate pipeline. Thinking effort is continuously controllable rather than stepped, and the lab's case for it is token efficiency: reaching a given score with fewer tokens than comparable models. Thinking Machines is direct that this is a base to customise rather than a frontier model, saying it is not the strongest model available today, open or closed.
Inkling is the first model Thinking Machines has trained itself, and the post announcing it spends most of its length explaining a choice rather than a score. The lab did not set out to win a leaderboard. It set out to publish a base that other people can take apart: full weights, a broad and deliberately balanced capability profile, and a fine-tuning path that exists on day one rather than as a promise. The argument behind that is a claim about where the remaining value sits - that many real problems are not solved well by even the best generalist model, and that the gap is closed by fine-tuning on knowledge a lab does not have and cannot buy.
The architecture follows from it. Inkling is a Mixture-of-Experts transformer of 975B total parameters with 41B active, 256 routed experts per layer and six active per token, using sigmoid routing, relative positional embeddings rather than rotary ones, and alternating sliding-window and global attention. It was pretrained on 45 trillion tokens spanning text, images, audio and video, and it carries a context window of up to a million tokens. None of that is arranged around a single headline capability; it is arranged around being a reasonable starting point for a lot of different endings.
Multimodality here is architectural rather than adjacent. Thinking Machines describes an encoder-free design: images arrive as 40x40 pixel patches through a four-layer hMLP, audio as dMel spectrograms through a light embedding layer, and both are then processed jointly with text tokens in the same stack. The practical consequence is that audio is a first-class input, not a transcription step bolted to the front - the model can transcribe speech, follow spoken instructions, answer questions about a recording and reason over longer-form audio without a separate pipeline in between.
The efficiency claim is the one the lab makes hardest, and it is about tokens rather than latency. Thinking effort on Inkling is continuous rather than stepped, set as a value between 0.2 and 0.99, and the lab's charts are drawn as score against tokens spent instead of score alone. On that framing it reports Inkling reaching a given score at fewer tokens than the models it compares against, matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. Its reported figures include 77.6% on SWEBench Verified, 63.8% on Terminal Bench 2.1, 87.2% on GPQA Diamond, 97.1% on AIME 2026, and on the multimodal side 73.5% on MMMU Pro and 91.4% on VoiceBench.
Thinking Machines is unusually direct about the ceiling. Its own words are that Inkling "is not the strongest overall model available today, open or closed", and that this is a first release in a family it intends to keep building on - it expects the multimodal side in particular to improve. Read that as the shape of the offer: reach for Inkling when you want a capable, cheap, genuinely multimodal model you can inspect, host or specialise, and reach for something else when you want the highest possible score on a task you are not going to fine-tune for.
لانچ کے وقت Thinking Machines کا بیان
- Built to be customised
- Thinking Machines frames Inkling as a base rather than a finished product - multimodal, efficient at thinking, and available for fine-tuning on its own Tinker platform from launch day, on the argument that specialised knowledge is what closes the remaining gap.
- Open weights, 975B total
- A Mixture-of-Experts transformer with 975B total and 41B active parameters, 256 experts per layer with six active per token, pretrained on 45 trillion tokens of text, images, audio and video.
- Encoder-free multimodality
- Images enter as 40x40 patches through a four-layer hMLP and audio as dMel spectrograms, both embedded and processed jointly with text tokens rather than through a separate perception model.
- Audio as a real input
- It transcribes speech, follows spoken instructions, answers questions about recordings and reasons over longer-form audio - capabilities the lab reports directly rather than through a transcription step.
- Effort as a dial, not a ladder
- Thinking effort is continuous between 0.2 and 0.99, and the lab measures itself on score per token: it reports matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens.
- What it is not for
- Thinking Machines states plainly that Inkling is not the strongest overall model available today, open or closed, and calls it the first of a family. If you need peak capability on a task you will not fine-tune for, this is not that model.
Frequently Asked Questions
Inkling 975B کے بارے میں اکثر پوچھے جانے والے سوالات۔
Inkling 975B کب جاری ہوا تھا؟
Thinking Machines نے Inkling 975B کو 17 جولائی، 2026 کو جاری کیا۔
Inkling 975B کس نے بنایا ہے؟
Inkling 975B کو Thinking Machines نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔
Inkling 975B کتنا ذہین ہے؟
یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 37 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔
Inkling 975B کا کتنا خرچ آتا ہے؟
استعمال کی لاگت ₹96.00 فی 10 لاکھ ان پٹ Tokens اور ₹388.80 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔
امریکی ڈالر میں Inkling 975B کی قیمت کیا ہے؟
فراہم کنندہ $1.00 فی 10 لاکھ ان پٹ Tokens اور $4.05 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔
Inkling 975B کتنی لمبی گفتگو یاد رکھ سکتا ہے؟
اس کی Context حد 10.5 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔
کیا Inkling 975B مناسب قیمت میں بہترین کارکردگی دیتا ہے؟
یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 38 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔
کیا Inkling 975B جواب دینے سے پہلے سوچتا ہے؟
پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔
کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026