Update: Qwen 125B runs on a 24GB 4090—21 tok/s decode, not 100

The VRAM wall really did crack: a 6B-active MoE at 250K context on one gaming card—what the viral throughput number hides, and how to run it this weekend.

Nowline OCT 5 10:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Update: a datacenter 125B fits a single 24GB 4090

    Qwen 3.8 Flash Next (125B MoE, ~6B active) ran at a 250,000-token context on one 24GB RTX 4090 with no KV-cache quantization—now #1 on Hacker News (664 points) and independently reproduced. Yesterday's Strata pitch was a tool demo; today it's a recipe anyone with the card can follow.

  • The viral 100 tok/s is prefill, not generation

    The headline number is prefill throughput (~364 t/s). Actual generation lands near 21 tokens/sec decode on the 4090—fine for background agents and readable single-stream, but not snappy interactive chat. Budget latency around decode, not the number in the title.

  • Three ways to run it today

    Consumer: GGUF via llama.cpp or Ollama on a 24GB card. Server: vLLM now ships an official serve recipe for H200/H100/GB200/MI355X/RTX Pro 6000. Or use Strata's offload engine. The weekend unlock is a 250K-context, 125B assistant that never leaves your machine.