Step 5 Preview hits OpenRouter: 1M-context agent model at $1/M

StepFun's 600B sparse MoE joins a busy week: Claude Code 2.1.296, JetBrains' open Mellum2.1, GPT-6.1 Sol Ultrafast, and a 16.9MB on-device speech model.

Nowline OCT 10 1:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Step 5 Preview lands on OpenRouter

    StepFun's new agentic MoE — 600B total params, 27B active, 1M-token context — is live on OpenRouter at $1.00/M input and $2.70/M output (cache reads $0.05/M), with up to 64K output tokens. It's tuned for long-horizon coding and document work, so it's a cheap drop-in for repo- and doc-scale agents you can wire up today.

  • Claude Code 2.1.296 ships

    The Oct 9 release adds `autoCompactWindow` so subagents auto-compact earlier than the main conversation, a `CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL` env var to run every workflow agent on one model, and raises the MCP tool-description limit from 2,048 to 4,096 characters. Small knobs, but they directly change how multi-agent and MCP-heavy setups behave.

  • JetBrains open-sources Mellum2.1

    Mellum2.1 is a 12B mixture-of-experts coding model, RL-trained in real coding environments to explore repos, edit files, and verify its own changes — released on Hugging Face under Apache 2.0. A genuinely open, self-hostable agent model you can fine-tune or run locally instead of renting a closed API.

  • GPT-6.1 Sol gets an Ultrafast tier

    OpenAI added an `ultrafast` service_tier to the Responses API for GPT-6.1 Sol, shortening the gap between output tokens for latency-sensitive apps. It's open to all API users (subject to rate limits) with US and EU data residency; Ultrafast pricing sits on the pricing page.

  • Whistle: speech-to-text in 16.9MB

    Cactus Compute's Whistle is a 16.9MB, Apache-2.0 STT model that runs on CPU with no GPU and no dependencies — 4.31% WER on LibriSpeech test-clean (vs Whisper base's 4.9%) and ~11ms to first token on an M4 Pro. It covers 7 languages with word timestamps and keyword biasing, so you can ship offline voice features for phones, wearables, or robots this weekend.