Xiaomi open-sources MiMo-V2.6: 1T omni-modal MoE, MIT-licensed
Three sizes — 1T, 311B, and a 9B that runs local — all read text, image, video and audio at 1M context; Xiaomi open-sourced the full RL stack too.

Copy markdown
Flash is the one you can actually self-host
MiMo-V2.6-Flash-RL is a 309B MoE with only 15B active per token, released under MIT — so it's commercial-safe to ship, not a look-but-don't-touch license. It reads text, image, video and audio at 1M context and posts Terminal Bench 2.1 87.6 and OSWorld-Verified 80.8, putting a frontier-class agentic coder on your own GPUs with no per-token bill.
Pro trades blows with GPT-5.6 Sol
The 1.02T / 42B-active Pro model edges GPT-5.6 Sol on Terminal Bench 2.1 (89.9 vs 88.8) and scores 94.0 on CyberGym, though it still trails Claude Opus 5 on the harder Terminal Bench 4.0 (34.9 vs 49.0). Open weights are now in the same bracket as closed frontier for terminal and agent work — and you can host it yourself.
A 9B VL distill for your laptop
MiMo-V2.6-Distill-Qwen-9B is a 9B image-text-to-text model small enough to run on a single consumer GPU. Build this weekend: a fully local screenshot-reader, UI agent, or multimodal pipeline with zero cloud dependency and the same MIT terms.
The RL machinery ships too, not just the weights
Xiaomi open-sourced the training dynamics, the RL environments (code, general, visual and cyber) and the framework behind the run — reportedly around $3.47M of RL compute. If you train agent models, you can fork the pipeline and RL-tune your own instead of rebuilding the harness from scratch.