Blog / Model guides
Muse Glimmer 30B hardware requirements: what one GPU needs
Meta published Muse Glimmer 30B on 10 August 2026 under the Apache 2.0 license (Meta AI Research, Hugging Face). It is Meta's first open-weight release since the Llama line (VentureBeat). The model card lists 29.6 billion parameters in a dense causal transformer, a 1.8-billion-parameter vision encoder, text and image input with text output, and a 131,072-token context. Meta built it for local agent workflows and validated it on 24 to 32 GB devices: an RTX 5090 and the M4 Max and M5 Max MacBooks.
The card reports MCP Atlas 75.5, SWE-Bench Pro 51.2 and AIME 2026 94.7. Meta's comparison set is Gemma 4 31B and Qwen3.6 27B.
Memory by precision
Weights only. The KV cache for your context length and concurrency comes on top.
| Precision | Weights in memory | Basis |
|---|---|---|
| bf16 | about 60 GB | Calculated: 29.6B parameters at 2 bytes each |
| 8-bit | about 30 GB | Calculated: 1 byte per parameter |
| 4-bit | under 20 GB | Meta's figure on the model card |
One 80 GB H100 holds bf16 with cache to spare. One 48 GB L40S, the GPU in an AWS g6e.xlarge, holds the 8-bit build. A 24 GB card holds 4-bit, the configuration Meta designed for. The card names vLLM, SGLang, llama.cpp, MLX and Ollama as supported engines and ships a DFlash drafter for speculative decoding; with it, Meta reports 233.4 tokens per second on an RTX 5090.
What the GPU costs
List prices checked on 22 August 2026, plus the rate we observe on DeepSeek V4 Flash deployments for comparison.
| Shape | Per hour | Source |
|---|---|---|
| AWS g6e.xlarge (1x L40S), Stockholm, on-demand | $1.974 | Spare Cores |
| AWS g6e.xlarge, Stockholm, spot | $0.604 | Spare Cores |
| RunPod L40S, community cloud | $0.79 | GetDeploying |
| DeepSeek V4 Flash, 2x H200 on RunPod | $7 to $14 | Observed on our deployments |
At the Stockholm on-demand rate, one g6e.xlarge is about $1,440 a month around the clock (730 hours) and about $350 a month for business hours (176 hours) on a wake/sleep schedule. The cost calculator does the same arithmetic for your own hours.
Questions people ask
Does Muse Glimmer 30B fit on a 24 GB GPU?
Yes at 4-bit. Meta states the 4-bit build is under 20 GB and validated it on an RTX 5090 (32 GB) and 24 GB devices. The 8-bit build needs about 30 GB; bf16 needs about 60 GB.
Is Muse Glimmer 30B Apache 2.0?
Yes. Meta released the weights on Hugging Face on 10 August 2026 under Apache 2.0, which permits commercial use, fine-tuning and redistribution.
Does Muse Glimmer 30B accept images?
Yes. The model has a 1.8-billion-parameter vision encoder and takes text and images as input; output is text only.