Blog / Model guides

Muse Glimmer 30B hardware requirements: what one GPU needs

Published ยท Requirements calculated from the published model card

Meta published Muse Glimmer 30B on 10 August 2026 under the Apache 2.0 license (Meta AI Research, Hugging Face). It is Meta's first open-weight release since the Llama line (VentureBeat). The model card lists 29.6 billion parameters in a dense causal transformer, a 1.8-billion-parameter vision encoder, text and image input with text output, and a 131,072-token context. Meta built it for local agent workflows and validated it on 24 to 32 GB devices: an RTX 5090 and the M4 Max and M5 Max MacBooks.

The card reports MCP Atlas 75.5, SWE-Bench Pro 51.2 and AIME 2026 94.7. Meta's comparison set is Gemma 4 31B and Qwen3.6 27B.

Memory by precision

Weights only. The KV cache for your context length and concurrency comes on top.

PrecisionWeights in memoryBasis
bf16about 60 GBCalculated: 29.6B parameters at 2 bytes each
8-bitabout 30 GBCalculated: 1 byte per parameter
4-bitunder 20 GBMeta's figure on the model card

One 80 GB H100 holds bf16 with cache to spare. One 48 GB L40S, the GPU in an AWS g6e.xlarge, holds the 8-bit build. A 24 GB card holds 4-bit, the configuration Meta designed for. The card names vLLM, SGLang, llama.cpp, MLX and Ollama as supported engines and ships a DFlash drafter for speculative decoding; with it, Meta reports 233.4 tokens per second on an RTX 5090.

What the GPU costs

List prices checked on 22 August 2026, plus the rate we observe on DeepSeek V4 Flash deployments for comparison.

ShapePer hourSource
AWS g6e.xlarge (1x L40S), Stockholm, on-demand$1.974Spare Cores
AWS g6e.xlarge, Stockholm, spot$0.604Spare Cores
RunPod L40S, community cloud$0.79GetDeploying
DeepSeek V4 Flash, 2x H200 on RunPod$7 to $14Observed on our deployments

At the Stockholm on-demand rate, one g6e.xlarge is about $1,440 a month around the clock (730 hours) and about $350 a month for business hours (176 hours) on a wake/sleep schedule. The cost calculator does the same arithmetic for your own hours.

Questions people ask

Does Muse Glimmer 30B fit on a 24 GB GPU?

Yes at 4-bit. Meta states the 4-bit build is under 20 GB and validated it on an RTX 5090 (32 GB) and 24 GB devices. The 8-bit build needs about 30 GB; bf16 needs about 60 GB.

Is Muse Glimmer 30B Apache 2.0?

Yes. Meta released the weights on Hugging Face on 10 August 2026 under Apache 2.0, which permits commercial use, fine-tuning and redistribution.

Does Muse Glimmer 30B accept images?

Yes. The model has a 1.8-billion-parameter vision encoder and takes text and images as input; output is text only.

Start a 24-hour trial