99MODELS

Inkling 975B

नयाँ

Open-weight multimodal MoE that reasons natively over text, images and audio.

रिलिज मिति: 2026 जुलाई 17

388.80

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹96.00 प्रति 10 लाख Tokens

प्रति अमेरिकी डलर ₹96 मा 0% मार्कअपका साथ प्रदायककै दरमा गणना गरिन्छ।

विवरण

Context विन्डो
10,48,576 Tokens
अधिकतम आउटपुट
2,62,144 Tokens
स्वीकार गर्छ
टेक्स्ट, तस्बिरहरू, अडियो
Reasoning
सुरुमै चालु
प्रयास स्तर
none, minimal, low, medium, high, max
टुल प्रयोग
संरचित आउटपुट
छैन
कोड कार्यान्वयन
छैन
प्यारामिटरहरू
975B A41B MoE
इन्टेलिजेन्स र्‍याङ्क
54 मध्ये #37
भ्याल्यू र्‍याङ्क
54 मध्ये #38

मूल्य

मूल्य
प्रति 10 lakh TokensINRUSD
इनपुट96.00$1.00
आउटपुट388.80$4.05
क्यास गरिएको इनपुट16.32$0.17

बेन्चमार्क

रेटिङ बाहेकका सबै स्कोरहरू प्रतिशतमा छन्। सबै बेन्चमार्कहरू स्वतन्त्र रूपमा मापन गरिएका हुन्।

  • 87.2%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 31.9%

    HLE

    Humanity's Last Exam

  • 47.0%

    SciCode

    SciCode - scientific code generation

  • 77.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 55.1%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 14.0%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 36.5%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1408

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

Inkling 975B को बारेमा

Thinking Machines' first model: an open-weights Mixture-of-Experts transformer with 975B total and 41B active parameters, pretrained on 45 trillion tokens of text, images, audio and video. It reasons natively over all three input modalities through an encoder-free architecture that embeds image patches and audio spectrograms alongside text tokens, so it can transcribe speech, follow spoken instructions and reason over recordings without a separate pipeline. Thinking effort is continuously controllable rather than stepped, and the lab's case for it is token efficiency: reaching a given score with fewer tokens than comparable models. Thinking Machines is direct that this is a base to customise rather than a frontier model, saying it is not the strongest model available today, open or closed.

Inkling is the first model Thinking Machines has trained itself, and the post announcing it spends most of its length explaining a choice rather than a score. The lab did not set out to win a leaderboard. It set out to publish a base that other people can take apart: full weights, a broad and deliberately balanced capability profile, and a fine-tuning path that exists on day one rather than as a promise. The argument behind that is a claim about where the remaining value sits - that many real problems are not solved well by even the best generalist model, and that the gap is closed by fine-tuning on knowledge a lab does not have and cannot buy.

The architecture follows from it. Inkling is a Mixture-of-Experts transformer of 975B total parameters with 41B active, 256 routed experts per layer and six active per token, using sigmoid routing, relative positional embeddings rather than rotary ones, and alternating sliding-window and global attention. It was pretrained on 45 trillion tokens spanning text, images, audio and video, and it carries a context window of up to a million tokens. None of that is arranged around a single headline capability; it is arranged around being a reasonable starting point for a lot of different endings.

Multimodality here is architectural rather than adjacent. Thinking Machines describes an encoder-free design: images arrive as 40x40 pixel patches through a four-layer hMLP, audio as dMel spectrograms through a light embedding layer, and both are then processed jointly with text tokens in the same stack. The practical consequence is that audio is a first-class input, not a transcription step bolted to the front - the model can transcribe speech, follow spoken instructions, answer questions about a recording and reason over longer-form audio without a separate pipeline in between.

The efficiency claim is the one the lab makes hardest, and it is about tokens rather than latency. Thinking effort on Inkling is continuous rather than stepped, set as a value between 0.2 and 0.99, and the lab's charts are drawn as score against tokens spent instead of score alone. On that framing it reports Inkling reaching a given score at fewer tokens than the models it compares against, matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. Its reported figures include 77.6% on SWEBench Verified, 63.8% on Terminal Bench 2.1, 87.2% on GPQA Diamond, 97.1% on AIME 2026, and on the multimodal side 73.5% on MMMU Pro and 91.4% on VoiceBench.

Thinking Machines is unusually direct about the ceiling. Its own words are that Inkling "is not the strongest overall model available today, open or closed", and that this is a first release in a family it intends to keep building on - it expects the multimodal side in particular to improve. Read that as the shape of the offer: reach for Inkling when you want a capable, cheap, genuinely multimodal model you can inspect, host or specialise, and reach for something else when you want the highest possible score on a task you are not going to fine-tune for.

लन्चको समयमा Thinking Machines ले के भन्यो

Built to be customised
Thinking Machines frames Inkling as a base rather than a finished product - multimodal, efficient at thinking, and available for fine-tuning on its own Tinker platform from launch day, on the argument that specialised knowledge is what closes the remaining gap.
Open weights, 975B total
A Mixture-of-Experts transformer with 975B total and 41B active parameters, 256 experts per layer with six active per token, pretrained on 45 trillion tokens of text, images, audio and video.
Encoder-free multimodality
Images enter as 40x40 patches through a four-layer hMLP and audio as dMel spectrograms, both embedded and processed jointly with text tokens rather than through a separate perception model.
Audio as a real input
It transcribes speech, follows spoken instructions, answers questions about recordings and reasons over longer-form audio - capabilities the lab reports directly rather than through a transcription step.
Effort as a dial, not a ladder
Thinking effort is continuous between 0.2 and 0.99, and the lab measures itself on score per token: it reports matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens.
What it is not for
Thinking Machines states plainly that Inkling is not the strongest overall model available today, open or closed, and calls it the first of a family. If you need peak capability on a task you will not fine-tune for, this is not that model.

Frequently Asked Questions

Inkling 975B सम्बन्धी प्रायः सोधिने प्रश्नहरू।

Inkling 975B कहिले रिलिज भएको हो?

Thinking Machines ले Inkling 975B लाई 2026 जुलाई 17 मा सार्वजनिक गरेको हो।

Inkling 975B कसले बनाएको हो?

Inkling 975B लाई Thinking Machines ले बनाएको हो। 99Models AI ले प्रदायककै दरमा सिधै जोड्दछ।

Inkling 975B कत्तिको सक्षम र बुद्धिमानी छ?

यो हाम्रो बौद्धिकता श्रेणीकरणमा 54 च्याट Models मध्ये 37 स्थानमा छ। यसको पूर्ण अङ्क माथिको Benchmarks तालिकामा हेर्न सकिन्छ।

Inkling 975B को लागत कति पर्छ?

यसमा 0% मार्कअपका साथ प्रति 10 लाख इनपुट Tokens को ₹96.00 र आउटपुटको ₹388.80 लाग्छ। कुनै सदस्यता छैन; तपाईंले प्रयोग गरेअनुसार मात्र भुक्तानी गर्नुहुन्छ।

डलरमा Inkling 975B को API मूल्य कति हो?

प्रदायकले प्रति 10 लाख इनपुट Tokens को $1.00 र आउटपुट Tokens को $4.05 शुल्क लिन्छ। यस पृष्ठका दरहरू प्रति अमेरिकी डलर ₹96 मा रूपान्तरण गरिएका हुन्।

Inkling 975B ले कति लामो कुराकानी सम्झन सक्छ?

यसको Context विन्डो 10.5 lakh Tokens हो। यसले एकल अनुरोधमा प्रक्रिया गर्न सक्ने कुराकानी र संलग्न फाइलहरूको कुल क्षमता यही हो।

के Inkling 975B लागत अनुसार उत्कृष्ट छ?

मूल्य र गुणस्तरको आधारमा यो 54 Models मध्ये 38 स्थानमा छ। यसले बौद्धिकता र Token लागतको तुलना गर्दछ।

के Inkling 975B ले जवाफ दिनुअघि विचार गर्छ?

सुरुमै चालु। Reasoning उपलब्ध भएको ठाउँमा तपाईंले सिधै सन्देश बक्समा सोच्ने क्षमता समायोजन गर्न सक्नुहुन्छ।

क्याटलग अद्यावधिक गरिएको मिति: 2026 सेप्टेम्बर 9