Meta has shipped an agentic model that runs on personal hardware without a cloud call. Released on August 10 by Meta Superintelligence Labs, Muse Glimmer is a 30-billion-parameter (30B) multimodal model whose weights went straight to Hugging Face under Apache 2.0. The story here is not the parameter count but the deployment target: where most frontier models depend on data centers and network access, Glimmer is built to run always-on and offline on one consumer GPU or a single laptop.
A 'small flagship' distilled from Muse Spark
Glimmer was not trained from scratch. It is a compressed version of Meta's closed top-tier model, Muse Spark, transferred through logit distillation. Meta built it in three phases: pre-training used logit distillation on Muse Spark's outputs; mid-training added longer-context, agent-heavy data with richer reasoning traces; and post-training combined supervised fine-tuning with on-policy distillation and reinforcement learning across general, reasoning, coding, and agentic domains.
The result puts part of the closed Spark's capability into a downloadable form. Spark stays closed-weight, while only the smaller Glimmer can be downloaded, fine-tuned, and self-hosted.
License Apache 2.0 (commercial use and modification allowed)
Context 131,072+ tokens · knowledge cutoff 2026-01-04
I/O text + image input, text output · 100+ languages
How Meta squeezed 30B into 24GB
At full precision a 30B model needs over 55GB of memory — more than any consumer GPU offers. Meta quantized the weights to roughly 4-bit, bringing the language model under 20GB. The remaining headroom holds the perception encoder for image understanding, the KV cache, and a drafter model for faster inference, all inside a 24GB or 32GB envelope. Two quantized builds ship:
| Build | Target VRAM | Avg. degradation |
|---|---|---|
| K-Quant-Dynamic | 32GB | 0.2% |
| K-Quant-17GB | 24GB | 1.0% |
Degradation is averaged over accuracy metrics across 15 common benchmarks; Meta says agentic tasks see minimal to no loss. Speed comes from DFlash, a block-level speculative decoding scheme: a drafter proposes a 16-token block that the main model verifies in parallel. That lifts throughput from 74.9 to 233.4 tokens/sec on an RTX 5090 (a 3.1x speedup), 26.6 to 50.2 on an M5 Max, and 23.7 to 37.8 on an M4 Max.
Benchmarks: strong on agents and reasoning, weaker on computer use
Meta compared Glimmer against same-class models Gemma4-31B and Qwen3.6-27B. Glimmer led MCP Atlas at 75.5 versus 54.2 and 62.5, and also led on DeepSearch QA (74.6), SWE-Bench Pro (51.2), and reasoning benchmarks including AIME 2026 (94.7) and IFBench (77.0). It trailed on tasks that directly drive a screen or terminal, where Qwen3.6-27B led OSWorld-Verified (75.6 vs. Glimmer's 65.9) and TerminalBench 2.1. In short: strong on agentic orchestration and reasoning, still weak on computer-use work.
Ecosystem and context
The weights are downloadable now from Hugging Face, with rolling support for local and edge frameworks (Ollama, LM Studio, llama.cpp, ExecuTorch, MLX), scaled serving (vLLM, SGLang), and hosting partners such as Together AI, Fireworks AI, and OpenRouter. Hardware optimization is underway with AMD, Arm, Dell, Intel, and NVIDIA.
The launch arrived as Zuckerberg urged the U.S. to remove barriers to open-source AI, feeding an American debate over whether increasingly capable models should be freely distributed or kept under tighter control. Meta effectively answered with a concrete artifact: a downloadable 30B agent model. On safety, Meta stated that Glimmer does not meet the "Frontier AI" definition in its Advanced AI Scaling Framework, rating chem/bio, cyber, and loss-of-control risk at moderate or lower.
- Meta released the 30B agentic model Muse Glimmer as Apache 2.0 open weights on August 10.
- Distilled from the closed Muse Spark; 4-bit quantization lets it run fully offline on a 24GB consumer GPU or a Mac.
- DFlash speculative decoding delivers a 3.1x speedup (233 tok/s) on an RTX 5090.
- Leads on MCP Atlas, DeepSearch QA, and SWE-Bench Pro; trails Qwen3.6-27B on OSWorld and terminal tasks.
- Lands amid Zuckerberg's push to loosen open-source AI rules, reigniting the U.S. openness debate.