Step 5 Preview hits OpenRouter: 1M-context agent model at $1/M
StepFun's 600B sparse MoE joins a busy week: Claude Code 2.1.296, JetBrains' open Mellum2.1, GPT-6.1 Sol Ultrafast, and a 16.9MB on-device speech model.

Copy markdown
Step 5 Preview lands on OpenRouter
StepFun's new agentic MoE — 600B total params, 27B active, 1M-token context — is live on OpenRouter at $1.00/M input and $2.70/M output (cache reads $0.05/M), with up to 64K output tokens. It's tuned for long-horizon coding and document work, so it's a cheap drop-in for repo- and doc-scale agents you can wire up today.
Claude Code 2.1.296 ships
The Oct 9 release adds `autoCompactWindow` so subagents auto-compact earlier than the main conversation, a `CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL` env var to run every workflow agent on one model, and raises the MCP tool-description limit from 2,048 to 4,096 characters. Small knobs, but they directly change how multi-agent and MCP-heavy setups behave.
JetBrains open-sources Mellum2.1
Mellum2.1 is a 12B mixture-of-experts coding model, RL-trained in real coding environments to explore repos, edit files, and verify its own changes — released on Hugging Face under Apache 2.0. A genuinely open, self-hostable agent model you can fine-tune or run locally instead of renting a closed API.
GPT-6.1 Sol gets an Ultrafast tier
OpenAI added an `ultrafast` service_tier to the Responses API for GPT-6.1 Sol, shortening the gap between output tokens for latency-sensitive apps. It's open to all API users (subject to rate limits) with US and EU data residency; Ultrafast pricing sits on the pricing page.
Whistle: speech-to-text in 16.9MB
Cactus Compute's Whistle is a 16.9MB, Apache-2.0 STT model that runs on CPU with no GPU and no dependencies — 4.31% WER on LibriSpeech test-clean (vs Whisper base's 4.9%) and ~11ms to first token on an M4 Pro. It covers 7 languages with word timestamps and keyword biasing, so you can ship offline voice features for phones, wearables, or robots this weekend.