Meta's Muse Glimmer: a 30B open agentic model for your laptop
Apache 2.0 and under 20GB quantized, it beats Gemma4-31B and Qwen3.6 on agentic benches — with day-one Ollama, Together, and Fireworks support.

Copy markdown
A 30B agent that fits on a MacBook
Muse Glimmer is a 30B model that quantizes to under 20GB at 4-bit — Meta tested it on a MacBook M4-Max, M5-Max, and an RTX 5090. That's a full tool-using agent running on hardware you already own, with no API bill.
It out-benches Gemma4-31B and Qwen3.6-27B
Meta pits Glimmer against Gemma4-31B and Qwen3.6-27B on DeepSearch QA, MCP-Atlas, τ-Bench, and SWE-Bench, leaning on reliable tool use, multi-step reasoning, and — notably — failure recovery when a tool call breaks mid-task.
Speculative decoding: 3.1x faster on a 5090
Built-in speculative decoding hits 3.1x decode speed on an RTX 5090, 1.8x on an M5 Max, and 1.5x on an M4 Max. Local agents that no longer feel like they're crawling.
Apache 2.0, live on Ollama, Together, and Fireworks
The weights are on Hugging Face under Apache 2.0, with day-one deployment through Ollama, Together AI, and Fireworks AI. Fine-tune it and ship commercially without a license fight.
Build this weekend: a private, offline agent
Multimodal input via a dedicated perception encoder, 100+ languages, and a controllable reasoning-effort dial mean you can wire an air-gapped agent over your own code, files, and tools that never phones home.
The play: Meta's first open superintelligence model
Glimmer is distilled from the closed Muse Spark, the Superintelligence Lab's flagship, making it MSL's first open release. Zuckerberg's 'The Future Is for Everyone' essay urges US policy to clear friction for open models, and an open Muse Spark is promised next.