99MODELS

Inkling 975B

ନୂଆ

Open-weight multimodal MoE that reasons natively over text, images and audio.

ରିଲିଜ୍ ତାରିଖ ଜୁଲାଇ 17, 2026

388.80

ପ୍ରତି 10 ଲକ୍ଷ output Tokens

Input: ପ୍ରତି 10 ଲକ୍ଷ Tokens ପାଇଁ ₹96.00

0% ମାର୍କଅପ୍ ସହିତ ₹96 ପ୍ରତି US ଡଲାର ହିସାବରେ ପ୍ରୋଭାଇଡର୍ ରେଟ୍ ରେ ବିଲ୍ କରାଯାଇଛି।

ସ୍ପେସିଫିକେସନ୍

Context window
10,48,576 Tokens
ସର୍ବାଧିକ Output
2,62,144 Tokens
ଗ୍ରହଣ କରେ
ଟେକ୍ସଟ୍, ଛବି, Audio
Reasoning
ଡିଫଲ୍ଟ ଭାବରେ On
Effort ସ୍ତର
none, minimal, low, medium, high, max
Tool ବ୍ୟବହାର
ହଁ
Structured output
ନାହିଁ
Code execution
ନାହିଁ
Parameters
975B A41B MoE
Intelligence rank
54 ମଧ୍ୟରୁ #37

ମୂଲ୍ୟ

ମୂଲ୍ୟ
ପ୍ରତି 10 ଲକ୍ଷ TokensINRUSD
Input96.00$1.00
Output388.80$4.05
Cached input16.32$0.17

Benchmarks

ରେଟିଂ ଭାବେ ଚିହ୍ନିତ ନ ହେଲେ ସ୍କୋରଗୁଡ଼ିକ ପ୍ରତିଶତ ଅଟେ। ସମସ୍ତ benchmark ସ୍ୱତନ୍ତ୍ର ଭାବରେ ମପାଯାଇଛି।

  • 87.2%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 31.9%

    HLE

    Humanity's Last Exam

  • 47.0%

    SciCode

    SciCode - scientific code generation

  • 77.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 55.1%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 14.0%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 36.5%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1408

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

Inkling 975B ବିଷୟରେ

Thinking Machines' first model: an open-weights Mixture-of-Experts transformer with 975B total and 41B active parameters, pretrained on 45 trillion tokens of text, images, audio and video. It reasons natively over all three input modalities through an encoder-free architecture that embeds image patches and audio spectrograms alongside text tokens, so it can transcribe speech, follow spoken instructions and reason over recordings without a separate pipeline. Thinking effort is continuously controllable rather than stepped, and the lab's case for it is token efficiency: reaching a given score with fewer tokens than comparable models. Thinking Machines is direct that this is a base to customise rather than a frontier model, saying it is not the strongest model available today, open or closed.

Inkling is the first model Thinking Machines has trained itself, and the post announcing it spends most of its length explaining a choice rather than a score. The lab did not set out to win a leaderboard. It set out to publish a base that other people can take apart: full weights, a broad and deliberately balanced capability profile, and a fine-tuning path that exists on day one rather than as a promise. The argument behind that is a claim about where the remaining value sits - that many real problems are not solved well by even the best generalist model, and that the gap is closed by fine-tuning on knowledge a lab does not have and cannot buy.

The architecture follows from it. Inkling is a Mixture-of-Experts transformer of 975B total parameters with 41B active, 256 routed experts per layer and six active per token, using sigmoid routing, relative positional embeddings rather than rotary ones, and alternating sliding-window and global attention. It was pretrained on 45 trillion tokens spanning text, images, audio and video, and it carries a context window of up to a million tokens. None of that is arranged around a single headline capability; it is arranged around being a reasonable starting point for a lot of different endings.

Multimodality here is architectural rather than adjacent. Thinking Machines describes an encoder-free design: images arrive as 40x40 pixel patches through a four-layer hMLP, audio as dMel spectrograms through a light embedding layer, and both are then processed jointly with text tokens in the same stack. The practical consequence is that audio is a first-class input, not a transcription step bolted to the front - the model can transcribe speech, follow spoken instructions, answer questions about a recording and reason over longer-form audio without a separate pipeline in between.

The efficiency claim is the one the lab makes hardest, and it is about tokens rather than latency. Thinking effort on Inkling is continuous rather than stepped, set as a value between 0.2 and 0.99, and the lab's charts are drawn as score against tokens spent instead of score alone. On that framing it reports Inkling reaching a given score at fewer tokens than the models it compares against, matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. Its reported figures include 77.6% on SWEBench Verified, 63.8% on Terminal Bench 2.1, 87.2% on GPQA Diamond, 97.1% on AIME 2026, and on the multimodal side 73.5% on MMMU Pro and 91.4% on VoiceBench.

Thinking Machines is unusually direct about the ceiling. Its own words are that Inkling "is not the strongest overall model available today, open or closed", and that this is a first release in a family it intends to keep building on - it expects the multimodal side in particular to improve. Read that as the shape of the offer: reach for Inkling when you want a capable, cheap, genuinely multimodal model you can inspect, host or specialise, and reach for something else when you want the highest possible score on a task you are not going to fine-tune for.

ଲଞ୍ଚ ସମୟରେ Thinking Machines ଯାହା କହିଥିଲା

Built to be customised
Thinking Machines frames Inkling as a base rather than a finished product - multimodal, efficient at thinking, and available for fine-tuning on its own Tinker platform from launch day, on the argument that specialised knowledge is what closes the remaining gap.
Open weights, 975B total
A Mixture-of-Experts transformer with 975B total and 41B active parameters, 256 experts per layer with six active per token, pretrained on 45 trillion tokens of text, images, audio and video.
Encoder-free multimodality
Images enter as 40x40 patches through a four-layer hMLP and audio as dMel spectrograms, both embedded and processed jointly with text tokens rather than through a separate perception model.
Audio as a real input
It transcribes speech, follows spoken instructions, answers questions about recordings and reasons over longer-form audio - capabilities the lab reports directly rather than through a transcription step.
Effort as a dial, not a ladder
Thinking effort is continuous between 0.2 and 0.99, and the lab measures itself on score per token: it reports matching Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens.
What it is not for
Thinking Machines states plainly that Inkling is not the strongest overall model available today, open or closed, and calls it the first of a family. If you need peak capability on a task you will not fine-tune for, this is not that model.

Frequently Asked Questions

Inkling 975B ବିଷୟରେ ବାରମ୍ବାର ପଚରାଯାଉଥିବା ପ୍ରଶ୍ନ।

Inkling 975B କେବେ ଲଞ୍ଚ ହୋଇଥିଲା?

Thinking Machines ଜୁଲାଇ 17, 2026 ରେ Inkling 975B ଲଞ୍ଚ କରିଥିଲା।

Inkling 975B କିଏ ତିଆରି କରିଛି?

Inkling 975B କୁ Thinking Machines ତିଆରି କରିଛି। 99Models ଏହା ସହ ସିଧାସଳଖ ପ୍ରୋଭାଇଡରଙ୍କ ନିର୍ଦ୍ଧାରିତ ରେଟ୍ ରେ ସଂଯୋଗ କରେ।

Inkling 975B କେତେ ଶକ୍ତିଶାଳୀ?

ଇଣ୍ଟେଲିଜେନ୍ସ ତାଲିକାରେ 54 ଟି chat Model ମଧ୍ୟରୁ ଏହାର ରାଙ୍କ୍ 37। ସ୍ୱାଧୀନ ବେଞ୍ଚମାର୍କ ସ୍କୋର ଆଧାରରେ ଏହା ସ୍ଥିର କରାଯାଇଛି, ଯାହା ଉପରେ ଥିବା ଟେବୁଲରେ ଉପଲବ୍ଧ।

Inkling 975B ର ମୂଲ୍ୟ କେତେ?

ବ୍ୟବହାର ଖର୍ଚ୍ଚ ପ୍ରତି 10 ଲକ୍ଷ ଇନପୁଟ୍ Tokens ପାଇଁ ₹96.00 ଏବଂ ଆଉଟପୁଟ୍ Tokens ପାଇଁ ₹388.80, 0% ମାର୍କଅପ୍ ସହିତ। କୌଣସି ସବସ୍କ୍ରିପସନ୍ ନାହିଁ; ଆପଣ ଯେତିକି ବ୍ୟବହାର କରିବେ ସେତିକି ପେମେଣ୍ଟ କରିବେ।

US ଡଲାରରେ Inkling 975B ର ମୂଲ୍ୟ କେତେ?

ପ୍ରୋଭାଇଡର୍ ପ୍ରତି 10 ଲକ୍ଷ ଇନପୁଟ୍ Tokens ପାଇଁ $1.00 ଏବଂ ଆଉଟପୁଟ୍ Tokens ପାଇଁ $4.05 ଚାର୍ଜ କରେ। ଏହି ପୃଷ୍ଠାର ଟଙ୍କା ମୂଲ୍ୟ ₹96 ପ୍ରତି ଡଲାର ହିସାବରେ ରୂପାନ୍ତରିତ।

Inkling 975B କେତେ ଲମ୍ବା କଥାବାର୍ତ୍ତା ମନେ ରଖିପାରିବ?

ଏହାର Context window ହେଉଛି 10.5 lakh Tokens। ଗୋଟିଏ request ରେ ଏହା ସମୁଦାୟ କଥାବାର୍ତ୍ତା ଏବଂ ସଂଲଗ୍ନ ଫାଇଲ୍ ପଢ଼ିପାରିବ।

ମୂଲ୍ୟ ହିସାବରେ Inkling 975B କେତେ ଭଲ?

ଭ୍ୟାଲୁ ରାଙ୍କିଙ୍ଗରେ 54 ଟି Model ମଧ୍ୟରୁ ଏହାର ସ୍ଥାନ 38। ଏହି ରାଙ୍କିଙ୍ଗ ବେଞ୍ଚମାର୍କ କ୍ଷମତା ଏବଂ Token ଖର୍ଚ୍ଚକୁ ତୁଳନା କରି ସ୍ଥିର କରାଯାଏ।

Inkling 975B କ’ଣ Reasoning ସପୋର୍ଟ କରେ?

ଡିଫଲ୍ଟ ଭାବରେ On। ଯେଉଁଠାରେ Reasoning ଉପଲବ୍ଧ, ଆପଣ ମେସେଜ୍ ବକ୍ସରେ thinking effort ସ୍ତର ସେଟ୍ କରିପାରିବେ।

କାଟାଲଗ୍ ଅପଡେଟ୍ ହୋଇଛି: ସେପ୍ଟେମ୍ବର 9, 2026