99MODELS

ગ્રોક 4.6

નવું

xAI's smartest model; frontier coding, knowledge work and STEM.

રજૂઆત: 12 ઑગસ્ટ, 2026

633.60

10 લાખ આઉટપુટ Tokens દીઠ

ઇનપુટ: 10 લાખ Tokens દીઠ ₹211.20

પ્રોવાઇડરના ભાવે બિલિંગ, $1 = ₹96 મુજબ રૂપાંતરિત, 0% માર્કઅપ સાથે.

વિશિષ્ટતાઓ

Context વિન્ડો
5,00,000 Tokens
મહત્તમ આઉટપુટ
4,50,000 Tokens
સ્વીકારે છે
ટેક્સ્ટ, ઇમેજ, PDF ફાઇલો
Reasoning
ડિફૉલ્ટ રૂપે ચાલુ
એફર્ટ લેવલ
low, medium, high, xhigh
ટૂલ ઉપયોગ
હા
સ્ટ્રક્ચર્ડ આઉટપુટ
હા
કોડ એક્ઝિક્યુશન
ના
ઇન્ટેલિજન્સ રેન્ક
54 માંથી #8
Value રેન્ક
54 માંથી #10

કિંમત

કિંમત
પ્રતિ 10 lakh TokensINRUSD
ઇનપુટ211.20$2.20
આઉટપુટ633.60$6.60
કૅશ ઇનપુટ52.80$0.55

Benchmarks

સ્કોર ટકાવારીમાં છે સિવાય કે રેટિંગ તરીકે દર્શાવેલ હોય. તમામ બેન્ચમાર્ક સ્વતંત્ર રીતે માપવામાં આવ્યા છે.

  • 94.9%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 42.9%

    HLE

    Humanity's Last Exam

  • 56.5%

    SciCode

    SciCode - scientific code generation

  • 80.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 88.4%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 48.0%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 67.1%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1629

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

ગ્રોક 4.6 વિશે

SpaceXAI's frontier model, built for coding, agentic tasks and knowledge work. It builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work, and the lab calls it the most intelligent and fastest model it has built. It is strongest at sustaining complex multi-step projects and at turning broad product concepts into working prototypes; reasoning is always on.

xAI attributes the step up to more work after pre-training than 4.5 got: synthetic material the lab curated itself around reasoning and hard technical ground, engineering data it rates highly, and a reworked optimiser and recipe. It says that combination left a stronger base for the supervised and reinforcement stages that follow. Grok 4.5 was then put to work regenerating the supervised training traces itself - across every effort setting, several agent scaffolds, and subjects from science to software to office work - with a model-run filter discarding the traces it judged unsound. The reinforcement learning stage covered agentic tasks in knowledge work, general coding, and purpose-built environments for kernel optimisation, web development and computer-aided design.

What xAI says that produced is a model that holds a goal across many steps. Its own testing focused on projects chosen to stretch range and endurance, and the lab reports its strongest showing when a vague product brief has to become something that runs: reading up on a domain it does not know, deciding the shape of the application, building the interactions that matter, then tightening them over repeated passes. On longer runs xAI reports the model increasingly stopping to check itself, confirming what it has done before carrying on - which is the behaviour that decides whether an unattended agent finishes or quietly drifts.

The visual side is where the lab draws the sharpest line against the previous generation. xAI reports stronger first passes on visual and interactive projects than it typically saw from Grok 4.5, with the model able to establish an application's structure and visual language in a single pass from a concrete product idea. The argument is about workflow rather than polish: when the fastest route to a good result is to begin with something substantial and then iterate in the loop, a strong first pass is worth more than a clean sketch.

On its own evaluation table, run at high effort, xAI reports Grok 4.6 at 69.9% on CursorBench v3.2 against 66.7% for Grok 4.5, 65.9% on DeepSWE v1.1 against 54%, 61.3% on the extended FrontierCode v1.1 against 56.6%, 57.5% on APEX-Agents against 47.1%, and 56.4% on APEX-SWE. Every one of those is a clear generational gain over the model it replaces.

The same table shows where it does not lead. On Terminal-Bench v3.0 xAI reports 26% against 34.6% and 34.1% for the two frontier models it compares itself with, and its DeepSWE v1.1 result sits below both of them as well. Reasoning cannot be switched off here - the dial runs low, medium, high and xhigh, with high as the default - so there is no cheap non-thinking mode to fall back to on simple calls. The lab says safeguards were recalibrated to the model's wider capability surface, with its broadest pre-deployment test suite to date plus post-deployment and third-party testing, and that the stack is tuned to keep the model usable for vulnerability patching, engineering design work and AI research rather than refusing those outright.

xAI એ લોન્ચ વખતે શું જણાવ્યું હતું

Long-running agents
xAI built this generation to stay with a complex task across many steps - researching a topic, analysing information, working across a codebase, or carrying an idea through to a finished artifact.
A working first version
The lab reports the model is especially strong at turning a broad product idea into something that runs: it researches the domain, structures the application, implements the core interactions and keeps refining.
Checks its own work
On longer trajectories xAI says it saw more self-testing and verification, with the model validating a step before moving on to the next one.
Visual and interactive passes
Given a concrete product idea, xAI reports the model establishes an application's structure and visual language in one pass, which it did not typically see from Grok 4.5.
Agentic coding gains
On the lab's own table at high effort: 69.9% on CursorBench v3.2, 65.9% on DeepSWE v1.1 and 61.3% on the extended FrontierCode v1.1, against 66.7%, 54% and 56.6% for Grok 4.5.
What it is not for
On Terminal-Bench v3.0 xAI reports 26% against 34.6% and 34.1% for the frontier models it compares with, and reasoning cannot be disabled, so there is no cheap non-thinking path for trivial calls.

ભારતીય ભાષાઓ

ગ્રોક 4.6 કુલ 15 ભારતીય ભાષાઓમાં જવાબ આપી શકે છે. મેસેજ બોક્સની બાજુમાં આપેલા મેનૂમાંથી ભાષા પસંદ કરો.

Frequently Asked Questions

ગ્રોક 4.6 વિશે વારંવાર પૂછાતા પ્રશ્નો.

ગ્રોક 4.6 ક્યારે રજૂ કરવામાં આવ્યું હતું?

xAI એ 12 ઑગસ્ટ, 2026 ના રોજ ગ્રોક 4.6 રજૂ કર્યું હતું.

ગ્રોક 4.6 કોણે બનાવ્યું છે?

ગ્રોક 4.6 ને xAI લેબ દ્વારા બનાવવામાં આવ્યું છે. 99Models સીધા પ્રોવાઇડરના નિર્ધારિત દરે કનેક્ટ કરે છે.

ગ્રોક 4.6 કેટલું સક્ષમ અને બુદ્ધિશાળી છે?

ઇન્ટેલિજન્સ રેન્કિંગમાં તે 54 ચેટ Models માંથી 8 નંબર પર છે, જે સ્વતંત્ર બેન્ચમાર્ક સ્કોર પર આધારિત છે. તેના તમામ સ્કોર ઉપર બેન્ચમાર્ક ટેબલમાં જોઈ શકો છો.

ગ્રોક 4.6 નો ઉપયોગ કરવાનો ખર્ચ કેટલો છે?

વપરાશ ખર્ચ ઇનપુટ માટે ₹211.20 પ્રતિ 10 લાખ Tokens અને આઉટપુટ માટે ₹633.60 પ્રતિ 10 લાખ Tokens છે, જેમાં 0% માર્કઅપ છે. કોઈ સબ્સ્ક્રિપ્શન નથી; વપરાશ મુજબ જ પેમેન્ટ કરો.

ગ્રોક 4.6 નું અમેરિકન ડોલરમાં API પ્રાઇસિંગ શું છે?

પ્રોવાઇડર ઇનપુટ માટે $2.20 પ્રતિ 10 લાખ Tokens અને આઉટપુટ માટે $6.60 પ્રતિ 10 લાખ Tokens ચાર્જ કરે છે. રૂપિયાના દરો $1 = ₹96 ના આધારે ગણવામાં આવ્યા છે.

ગ્રોક 4.6 કેટલી લાંબી વાતચીત યાદ રાખી શકે છે?

તેની Context વિન્ડો 5 lakh Tokens ની છે. એટલે કે તે એક જ વિનંતીમાં અગાઉની વાતચીત અને જોડેલી ફાઇલો સહિત આટલું લખાણ વાંચી શકે છે.

શું ગ્રોક 4.6 વાપરવું પૈસા વસૂલ છે?

વેલ્યૂ રેન્કિંગમાં તે 54 માંથી 10 નંબર પર છે. આ રેન્કિંગ બેન્ચમાર્ક ક્ષમતા અને વાસ્તવિક Token ખર્ચની તુલના કરે છે.

શું ગ્રોક 4.6 જવાબ આપતાં પહેલાં વિચારે છે?

ડિફૉલ્ટ રૂપે ચાલુ. જ્યાં Reasoning સપોર્ટ ઉપલબ્ધ છે, ત્યાં તમે મેસેજ કમ્પોઝરમાં વિચારવાની ક્ષમતા (Thinking effort) જાતે સેટ કરી શકો છો.

ગ્રોક 4.6 કઈ ભારતીય ભાષાઓમાં જવાબ આપી શકે છે?

તે 15 ભારતીય ભાષાઓમાં જવાબ આપે છે. મેસેજ બોક્સ પાસેના મેનૂમાંથી ભાષા પસંદ કરો અને તે જ ભાષામાં જવાબ મેળવો.

xAI ના અન્ય Models

કેટલોગ અપડેટ: 9 સપ્ટે, 2026