Meta has released Muse Glimmer, a 30 billion-parameter AI model that runs on laptops and consumer GPUs. The model is built for local agentic tasks like scheduling, coding, and tool use, giving users advanced AI without depending on cloud servers.
Muse Glimmer is open-weight and licensed under Apache 2.0, so anyone can use it for commercial projects. It supports multimodal input, multilingual tasks, and works with frameworks like llama.cpp and MLX.
Meta says the model uses quantization and speculative decoding to speed up performance, making it efficient enough for personal devices while still handling complex reasoning.
Meta built Muse Glimmer by shrinking down its bigger system, Muse Spark, so it’s powerful but still easy to use. Benchmarks show Glimmer is competitive with mid-sized models like Google’s Gemma and Alibaba’s Qwen, though it doesn’t match the scale of bigger cloud systems. Running locally gives developers more control over data and workflows.
On NVIDIA hardware, Glimmer can reach up to 20K tokens per second using Blackwell Ultra GPUs. It’s optimized for deployment through NIM containers, making it useful for edge cases like robotics and industrial automation.
Also Read: Meta withdraws Muse Image tool days after launch amid privacy concerns
With Muse Glimmer’s release under Apache 2.0, Meta is betting on community adoption and innovation around local AI agents.
Muse Glimmer is now available for download, with weights available on Hugging Face.






