Muse Glimmer 30B is a Meta open-weight catalog model. CompareLLM Coding Estimated rating is 1570. Active context window extends to 131k tokens. Commercial API token pricing is listed at $0.3/1M tok input and $1.1/1M tok output per 1M tokens. Meta's first open-weights release since Llama 4: a dense 30B multimodal model under Apache 2.0, distilled from Muse Spark for agents on local hardware. Numbers below are dated snapshots from empirical benchmark harnesses.
Independent evaluation answering: "Is Muse Glimmer 30B the right model for your workload & budget?"
Open weights still need local hardware; this card does not estimate VRAM.
Dated snapshot metrics aggregated from official evaluators and API providers with visual relative score bars.
Our category lists (Coding, Reasoning, Chat, Agents). Named suites such as SWE-bench Pro are recipe inputs, not this rating.
LMArena, SWE-bench Pro, LiveBench and similar suites appear with a source and as-of date. They are not CompareLLM's rating.
How much of the recipe is present. Missing inputs stay missing; they are never filled with 0 or 50.
Compact streaming class. Exact tok/s stays in the fact sheet, not in compact cells.
Compact class for time-to-first-token. Exact milliseconds stay in the fact sheet.
Commercial API price per 1 Million prompt/output tokens (~750k words).
| Benchmark Metric & Meaning | Reported Score & Capability Fill | CompareLLM Review |
|---|---|---|
Maximum output (tokens) Maximum Response Length | 118k tokens | Source not recorded · Sep 16, 2026Maximum completion tokens for one request on the same endpoint. |
Cached input token price Cost to Reuse Cached Input | $0.04/1M tok | Source not recorded · Sep 16, 2026Published cache-read price on the same endpoint. |
Context window (tokens) Context Window Capacity (tokens) | 131k tokens | Source not recorded · Sep 16, 2026Maximum input context on the same endpoint. |
Input token price Cost to Prompt (Input tokens) | $0.3/1M tok | Source not recorded · Sep 16, 2026Published input price on the endpoint this sheet is read from. |
Output token price Cost to Generate (Output tokens) | $1.1/1M tok | Source not recorded · Sep 16, 2026Published output price on the endpoint this sheet is read from. |
Generation speed (tok/s) Writing Speed (tok/s) | 50–100 tok/s | Source not recorded · Sep 16, 2026 |
Time to first token Waiting Time (before it replies) | 200–500ms | Source not recorded · Sep 16, 2026 |
Select any rival to launch a side-by-side empirical benchmark comparison with winner deltas.
Compare Muse Glimmer 30B against
Test your GPU VRAM, Apple Mac Mini / Studio, or Cloud Cluster across FP16, Q8, Q5, and Q4 quantization tiers with live KV cache calculations.
Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage. Coverage refers to the configured recipe, not confidence.
Release recent-models-2026-09-23-r1 · recipe reported-text-2026-09-11-r1 · method reported-with-estimates-v2 · research through 2026-09-23
Overall benchmark coverage includes missing applicable categories. Estimated categories contribute to the rating but add no benchmark coverage.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Muse Glimmer 30B (high) scores 35, and placed within the 48-66 band this catalog's measured results occupy for this category. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: repository-engineering, terminal-work, frontiercode.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Muse Glimmer 30B (high) scores 35, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: graduate-science, advanced-mathematics, broad-academic, interactive-abstraction.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Muse Glimmer 30B (high) scores 35, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: constrained-story-writing, instruction-compliance.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Muse Glimmer 30B (high) scores 35, and placed within the 59-66 band this catalog's measured results occupy for this category. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: professional-computer-tasks, desktop-use, workflow-automation, web-research.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Muse Glimmer 30B (high) scores 35, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: conversation-quality, support-task-completion, instruction-compliance.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Muse Glimmer 30B (high) scores 35, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: visual-grounding, spatial-reconstruction, visual-symbol-recognition.
What this model is available as, what it takes to run, and where its identity comes from.
The weights are published for download, so this model can be run on your own hardware.
Published weights — the provider's own download page.
muse-glimmer-30bmeta/muse-glimmer-30bConcise empirical overview formatted for citations and prompt context
Plain-English methodology and leaderboard answers
Elo is a crowd vote on which hidden answer people liked more — not a school test. What is preference Elo? · Methodology
Follow this model in your watchlist, set it as your global comparison baseline, or assign it to your custom production stack.
Track updates & rank changes
Compare all models against this
Assign to custom architecture
Compare side-by-side vs all
Plotted against all active catalog models (50th percentile = catalog median).
Ranks in the top tier (≥75th percentile) for Speed, Value.
Pre-computed production rankings across developer workloads based on empirical benchmark capability, throughput, and operational economics.
Multimodal models only, ranked on our rating with latency and price as tie-breakers.
Highest Reasoning rating among models you can self-host or buy as open weights.
CompareLLM Agents and Coding ratings for computer-use and multi-step tools.
Sonnet / Terra / Flash class. Measured quality without Opus or Sol prices.
CompareLLM Coding rating first, then price. For CI bots and repo agents that cannot burn Opus prices.
Output tokens per second first, among models with measured quality. For long completions, batch jobs, and agent loops.
Output price first among models with measured quality.