Report: Qwen's 125B/6B open MoE, Flash-Next, previews Qwen 4

Weights aren't public yet and no benchmarks — but a 6B-active MoE you can self-host is a big deal. Plus: GPT-5.6 in Kiro, Copilot in Slack and Teams.

Nowline AUG 26 8:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 125B total, ~6B active, open weights

    Qwen 3.8-Flash-Next is reportedly a 125B-parameter mixture-of-experts that fires only ~6B parameters per token — the runtime footprint of a small model at a far larger capacity — with open weights expected on Hugging Face and ModelScope. The specs come from a ModelScope page that briefly went live and was pulled; there are no official benchmarks or license terms yet, so treat the "Sonnet-class coding" chatter as rumor until it actually ships.

  • GPT-5.6 lands in Kiro's spec-driven flow

    OpenAI added the GPT-5.6 family — Sol, Terra, and Luna — to Kiro, its spec-driven coding agent that handles planning, building, review, and testing. OpenAI cites roughly an 82% cost reduction on Terminal-Bench 2.1 versus the previous generation, so the same agentic runs get materially cheaper.

  • Copilot agents move into Slack and Teams

    GitHub Copilot is now in public preview inside Slack and Microsoft Teams — mention @GitHub in a channel, thread, or DM to start a collaborative agent session without leaving chat. It puts the Copilot CLI and app's agentic coding where your team already coordinates.

  • Elsewhere: a robot that learns from one video

    Skild AI released S1, a robot foundation model that picks up a new multi-step task from a single human video with no fine-tuning — reportedly handling routines up to 10 minutes long and hitting 66% on tasks it has never seen. There's no download for builders yet, but it's a real step toward in-context learning for robotics.