MAI-Code-1.1-Flash goes downloadable: run a 138B coder locally

A 138B sparse MoE (5B active, 72.6% on SWE-Bench) that runs on-device at zero inference cost — and Copilot CLI can now discover your local Ollama models.

Nowline OCT 8 9:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Download it, run it, pay nothing

    MAI-Code-1.1-Flash can now be pulled from Microsoft's repo and run on your own hardware, and local calls carry zero inference charges. It keeps the full 256K context on-device, and Microsoft points at machines with 120GB+ of RAM — so a 128GB workstation is the target. It's the first Copilot-grade coding model that never has to leave your laptop.

  • 138B params, only 5B lit per token

    It's a sparse mixture-of-experts: 138B total parameters with roughly 5B active per token, which is why it stays cheap to run. The model card lists 72.6% on SWE-Bench Verified and 62.9% on Terminal-Bench 2.1 — agentic-coding scores, not autocomplete. In Copilot it runs at about a quarter the cost of the 1.0 model from June's Build.

  • Copilot CLI now finds your Ollama models

    As of CLI 1.0.94-0, running /model lists models from a live local Ollama instance next to Copilot's cloud models, and you can add one mid-session without a restart. The model needs tool-calling and streaming support. It doesn't flip on offline mode by itself (COPILOT_OFFLINE=true does that), but it's the plumbing for driving a local coder from Copilot.