99MODELS

Muse Glimmer 30B

New

Open-weight dense 30B distilled from Muse Spark; runs fast and cheap.

Released Aug 9, 2026

144.00

per 10 lakh output tokens

Input: ₹33.60 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
1,31,072 tokens
Max output
1,17,964 tokens
Accepts
Text, Images
Reasoning
On by default
Effort levels
low, medium, high, xhigh
Tool use
Yes
Structured output
Yes
Code execution
No
Parameters
30B
Intelligence rank
#42 of 54
Value rank
#37 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input33.60$0.35
Output144.00$1.50
Cached input3.84$0.04

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 83.5%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 22.0%

    HLE

    Humanity's Last Exam

  • 44.9%

    SciCode

    SciCode - scientific code generation

  • 83.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 51.7%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

About Muse Glimmer 30B

Meta's open-weight 30B model, optimised for local, always-on agent workflows. It is a dense transformer with a dedicated perception encoder, distilled from Muse Spark and purpose-built to run autonomous agents on consumer hardware. Meta aims it at end-to-end agentic task completion, reliable tool use, multi-step reasoning, failure recovery and use as an evaluation judge, with text and image input and selectable reasoning strength.

The argument Meta makes for this model is about where it runs. Most agent deployments assume a network and a datacentre; an agent that manages your calendar, drafts your messages and organises your files needs deep access to personal context, and Meta's answer is to put the model on the machine that already holds it. Everything about the design follows from that constraint - a compact architecture, a distillation recipe that pulls agentic reasoning out of a far larger teacher, and inference work aimed at latency rather than peak quality.

Meta breaks the recipe into three named stages. The first trains Glimmer on Muse Spark's own outputs rather than on raw text alone, so the small model inherits the large one's distribution. The second stretches the context and tilts the corpus towards agent transcripts that show their reasoning, kept diluted with ordinary text. The third layers three techniques over one another - supervised fine-tuning, on-policy distillation, then reinforcement learning - and applies them to general, reasoning, coding and agentic work alike. Meta says the result was assessed for open-weight release under its own Advanced AI Scaling Framework across every relevant risk category.

The capability list is written for agent builders rather than chat users, and one entry stands out: what happens after a tool call goes wrong. Meta says Glimmer was trained to read the failure, work out why it happened and try again, instead of halting on the error - which is the difference between an agent that finishes and one that needs a human. Alongside that sit precise function calling with real schemas across extended workflows, multi-step planning, a dedicated perception encoder that lets it read interleaved screenshots, charts and documents, selectable reasoning strength, and training data spanning more than a hundred languages.

Making it fit was its own engineering problem, and Meta shows the arithmetic. Thirty billion parameters at full precision needs over 55 GB; quantised to roughly four bits the language model drops under 20 GB, which leaves room for the working memory, the perception encoder and a speculative-decoding drafter inside a 24 GB or 32 GB budget - with, the lab says, minimal to no degradation on agentic tasks. A lightweight drafter proposes whole blocks of tokens that the main model verifies in parallel, which Meta reports as a large speed-up on a high-end consumer GPU and a meaningful one on laptop silicon.

Meta's own comparison table is unusually honest about the shape of the trade. Against the two open models near its size it leads on the agent-completion and tool-calling columns - the ones it was built for - while trailing on terminal-style coding, on computer use, and on general knowledge and hardest-exam questions. Read it as a model tuned for finishing tool-driven tasks locally, not as a small general-purpose frontier model.

What Meta announced at launch

Distilled from the flagship
Glimmer learns from Muse Spark's outputs rather than from raw text alone, then gains context length and agent behaviour in a second stage, before a third stage stacks fine-tuning, distillation and reinforcement learning.
Recovers from failed tool calls
Meta trained the model to diagnose an error and retry when a tool call fails or returns something unexpected, rather than halting the run.
Reads screens and documents
A dedicated perception encoder accepts interleaved text and images, so an agent can interpret screenshots, charts and documents alongside the conversation.
Engineered to fit a consumer GPU
At full precision the model needs over 55 GB; quantised to roughly four bits it drops under 20 GB, leaving room for working memory, the encoder and a drafter inside a 24 GB or 32 GB budget.
Speculative decoding for responsiveness
A lightweight drafter proposes blocks of tokens that the main model verifies in parallel, which Meta reports as a large generation speed-up on a high-end consumer GPU.
What it is not for
On Meta's own table it leads its size class on agent completion and tool calling but trails on terminal-style coding, computer use and general knowledge, so it is not a small general-purpose frontier model.

Indian languages

Muse Glimmer 30B answers in 9 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Muse Glimmer 30B.

When was Muse Glimmer 30B released?

Meta released Muse Glimmer 30B on Aug 9, 2026.

Who built Muse Glimmer 30B?

Muse Glimmer 30B is developed by Meta. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Muse Glimmer 30B?

It is ranked 42 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Muse Glimmer 30B cost?

Usage costs ₹33.60 per million input tokens and ₹144.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Muse Glimmer 30B pricing in US dollars?

The provider charges $0.35 per million input tokens and $1.50 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Muse Glimmer 30B hold?

Its context window is 1.3 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Muse Glimmer 30B rank for value?

It ranks 37 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Muse Glimmer 30B support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Muse Glimmer 30B support?

It answers in 9 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Meta

Catalog updated Sep 9, 2026