Muse Glimmer 30B
NewOpen-weight dense 30B distilled from Muse Spark; runs fast and cheap.
Released Aug 9, 2026
₹144.00
per 10 lakh output tokens
Input: ₹33.60 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
83.5%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
22.0%
HLE
Humanity's Last Exam
44.9%
SciCode
SciCode - scientific code generation
83.3%
Long Context
Long Context Reasoning - reasoning over long inputs
51.7%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
About Muse Glimmer 30B
Meta's open-weight 30B model, optimised for local, always-on agent workflows. It is a dense transformer with a dedicated perception encoder, distilled from Muse Spark and purpose-built to run autonomous agents on consumer hardware. Meta aims it at end-to-end agentic task completion, reliable tool use, multi-step reasoning, failure recovery and use as an evaluation judge, with text and image input and selectable reasoning strength.
The argument Meta makes for this model is about where it runs. Most agent deployments assume a network and a datacentre; an agent that manages your calendar, drafts your messages and organises your files needs deep access to personal context, and Meta's answer is to put the model on the machine that already holds it. Everything about the design follows from that constraint - a compact architecture, a distillation recipe that pulls agentic reasoning out of a far larger teacher, and inference work aimed at latency rather than peak quality.
Meta breaks the recipe into three named stages. The first trains Glimmer on Muse Spark's own outputs rather than on raw text alone, so the small model inherits the large one's distribution. The second stretches the context and tilts the corpus towards agent transcripts that show their reasoning, kept diluted with ordinary text. The third layers three techniques over one another - supervised fine-tuning, on-policy distillation, then reinforcement learning - and applies them to general, reasoning, coding and agentic work alike. Meta says the result was assessed for open-weight release under its own Advanced AI Scaling Framework across every relevant risk category.
The capability list is written for agent builders rather than chat users, and one entry stands out: what happens after a tool call goes wrong. Meta says Glimmer was trained to read the failure, work out why it happened and try again, instead of halting on the error - which is the difference between an agent that finishes and one that needs a human. Alongside that sit precise function calling with real schemas across extended workflows, multi-step planning, a dedicated perception encoder that lets it read interleaved screenshots, charts and documents, selectable reasoning strength, and training data spanning more than a hundred languages.
Making it fit was its own engineering problem, and Meta shows the arithmetic. Thirty billion parameters at full precision needs over 55 GB; quantised to roughly four bits the language model drops under 20 GB, which leaves room for the working memory, the perception encoder and a speculative-decoding drafter inside a 24 GB or 32 GB budget - with, the lab says, minimal to no degradation on agentic tasks. A lightweight drafter proposes whole blocks of tokens that the main model verifies in parallel, which Meta reports as a large speed-up on a high-end consumer GPU and a meaningful one on laptop silicon.
Meta's own comparison table is unusually honest about the shape of the trade. Against the two open models near its size it leads on the agent-completion and tool-calling columns - the ones it was built for - while trailing on terminal-style coding, on computer use, and on general knowledge and hardest-exam questions. Read it as a model tuned for finishing tool-driven tasks locally, not as a small general-purpose frontier model.
What Meta announced at launch
- Distilled from the flagship
- Glimmer learns from Muse Spark's outputs rather than from raw text alone, then gains context length and agent behaviour in a second stage, before a third stage stacks fine-tuning, distillation and reinforcement learning.
- Recovers from failed tool calls
- Meta trained the model to diagnose an error and retry when a tool call fails or returns something unexpected, rather than halting the run.
- Reads screens and documents
- A dedicated perception encoder accepts interleaved text and images, so an agent can interpret screenshots, charts and documents alongside the conversation.
- Engineered to fit a consumer GPU
- At full precision the model needs over 55 GB; quantised to roughly four bits it drops under 20 GB, leaving room for working memory, the encoder and a drafter inside a 24 GB or 32 GB budget.
- Speculative decoding for responsiveness
- A lightweight drafter proposes blocks of tokens that the main model verifies in parallel, which Meta reports as a large generation speed-up on a high-end consumer GPU.
- What it is not for
- On Meta's own table it leads its size class on agent completion and tool calling but trails on terminal-style coding, computer use and general knowledge, so it is not a small general-purpose frontier model.
Indian languages
Muse Glimmer 30B answers in 9 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about Muse Glimmer 30B.
When was Muse Glimmer 30B released?
Meta released Muse Glimmer 30B on Aug 9, 2026.
Who built Muse Glimmer 30B?
Muse Glimmer 30B is developed by Meta. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Muse Glimmer 30B?
It is ranked 42 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Muse Glimmer 30B cost?
Usage costs ₹33.60 per million input tokens and ₹144.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Muse Glimmer 30B pricing in US dollars?
The provider charges $0.35 per million input tokens and $1.50 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Muse Glimmer 30B hold?
Its context window is 1.3 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Muse Glimmer 30B rank for value?
It ranks 37 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Muse Glimmer 30B support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does Muse Glimmer 30B support?
It answers in 9 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Meta
Catalog updated Sep 9, 2026