Muse Spark 1.2
NewMeta's reasoning model for complex agentic tasks; accepts every modality.
Released Aug 5, 2026
₹408.00
per 10 lakh output tokens
Input: ₹120.00 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
90.4%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
45.5%
HLE
Humanity's Last Exam
57.4%
SciCode
SciCode - scientific code generation
79.0%
Long Context
Long Context Reasoning - reasoning over long inputs
80.1%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
1534
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
About Muse Spark 1.2
Meta Superintelligence Labs' multimodal reasoning model for agentic tasks, with major gains in tool and computer use, coding and multimodal understanding. Version 1.2 is the coding-optimised checkpoint, built for long-horizon multi-agentic workflows such as multi-file refactors and debugging sessions that run well past a single prompt, with a million-token window so an entire repository fits in one session. It was trained with asynchronous and parallel tool calls plus planning, goal conditioning and context compaction to hold focus across extended tasks.
Meta did not ship this checkpoint on its own. It arrived with Muse Code, a terminal coding agent, and the two were co-trained - the lab says it folded rejection-sampled harness trajectories into training along with recipe work for goals, context compaction and subagents, and integrated the agent's own toolset so the model and the harness would not disagree about how a task is run. That is the frame the release asks for: this is a model tuned to a specific way of working rather than a general upgrade.
The training emphasis is long-horizon work, and Meta names the material: whole-repository generation, large end-to-end projects and automated research. It credits three mechanisms for holding a run together over hours - planning to sequence the work, goal conditioning to keep direction, and context compaction to carry forward only what still matters. It also describes a self-improvement loop in which the previous Muse Spark generated hard coding environments and instruction-following templates, then graded candidate solutions against them, producing training data at a scale hand-authoring could not reach.
On its own charts Meta reports 82.9% on Terminal-Bench 2.1 run through Muse Code and 59.3% on DeepSWE 1.1, against 76.2% and 53.0% for the previous version measured in a lighter harness. The lab is not claiming the top of either chart: on both it places a frontier competitor above Muse Spark 1.2, and on the software-engineering set two competitors sit above it. The gain it is selling is against its own predecessor and against the cost of the tier.
The most telling result Meta published is not a benchmark at all. It set the model to optimise GPU kernels over more than a thousand tool calls and up to twenty-four hours, writing, compiling, profiling and improving Triton implementations against a reference, with third-party kernel libraries explicitly forbidden so the model had to implement the algorithms rather than wrap someone else's. It reports substantial and continuing improvement over the baseline across that window. The launch demo makes the multimodal half concrete in the same spirit: a walkthrough video handed to the terminal as a file, turned into a working booking page. Meta frames the release as a step rather than a destination, saying larger and much more capable models are on the way.
What Meta announced at launch
- Co-trained with its harness
- Meta trained the model together with its terminal coding agent, folding in rejection-sampled harness trajectories and the agent's own toolset so model and scaffold behave consistently.
- Trained on long-horizon work
- The lab names whole-repository generation, large end-to-end projects and automated research as training material, held together by planning, goal conditioning and context compaction.
- Coding benchmarks
- Meta reports 82.9% on Terminal-Bench 2.1 through its own agent and 59.3% on DeepSWE 1.1, against 76.2% and 53.0% for the previous Muse Spark release.
- Self-improvement loop
- The previous generation generated challenging coding environments and instruction-following templates, then graded candidate solutions, producing training data for this checkpoint at scale.
- A day-long kernel optimisation run
- Meta ran the model for over a thousand tool calls and up to twenty-four hours writing, compiling and profiling Triton GPU kernels, with third-party kernel libraries forbidden.
- What it is not for
- This is a coding-focused update, and on Meta's own two charts a competing frontier model scores above it on both. The lab describes it as a step, with larger models still to come.
Indian languages
Muse Spark 1.2 answers in 15 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about Muse Spark 1.2.
When was Muse Spark 1.2 released?
Meta released Muse Spark 1.2 on Aug 5, 2026.
Who built Muse Spark 1.2?
Muse Spark 1.2 is developed by Meta. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Muse Spark 1.2?
It is ranked 14 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Muse Spark 1.2 cost?
Usage costs ₹120.00 per million input tokens and ₹408.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Muse Spark 1.2 pricing in US dollars?
The provider charges $1.25 per million input tokens and $4.25 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Muse Spark 1.2 hold?
Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Muse Spark 1.2 rank for value?
It ranks 9 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Muse Spark 1.2 support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does Muse Spark 1.2 support?
It answers in 15 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Meta
Catalog updated Sep 9, 2026