Meta's Muse Glimmer: a 30B open agentic model for your laptop

Apache 2.0 and under 20GB quantized, it beats Gemma4-31B and Qwen3.6 on agentic benches — with day-one Ollama, Together, and Fireworks support.

Nowline AUG 10 6:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • A 30B agent that fits on a MacBook

    Muse Glimmer is a 30B model that quantizes to under 20GB at 4-bit — Meta tested it on a MacBook M4-Max, M5-Max, and an RTX 5090. That's a full tool-using agent running on hardware you already own, with no API bill.

  • It out-benches Gemma4-31B and Qwen3.6-27B

    Meta pits Glimmer against Gemma4-31B and Qwen3.6-27B on DeepSearch QA, MCP-Atlas, τ-Bench, and SWE-Bench, leaning on reliable tool use, multi-step reasoning, and — notably — failure recovery when a tool call breaks mid-task.

  • Speculative decoding: 3.1x faster on a 5090

    Built-in speculative decoding hits 3.1x decode speed on an RTX 5090, 1.8x on an M5 Max, and 1.5x on an M4 Max. Local agents that no longer feel like they're crawling.

  • Apache 2.0, live on Ollama, Together, and Fireworks

    The weights are on Hugging Face under Apache 2.0, with day-one deployment through Ollama, Together AI, and Fireworks AI. Fine-tune it and ship commercially without a license fight.

  • Build this weekend: a private, offline agent

    Multimodal input via a dedicated perception encoder, 100+ languages, and a controllable reasoning-effort dial mean you can wire an air-gapped agent over your own code, files, and tools that never phones home.

  • The play: Meta's first open superintelligence model

    Glimmer is distilled from the closed Muse Spark, the Superintelligence Lab's flagship, making it MSL's first open release. Zuckerberg's 'The Future Is for Everyone' essay urges US policy to clear friction for open models, and an open Muse Spark is promised next.