Update: Strata runs Qwen 125B on a 12GB GPU at 94 tok/s, API-ready
Half the VRAM of the earlier 4090 demo, with drop-in OpenAI and Anthropic endpoints — if you can accept 2-bit quality and one request at a time.

Copy markdown
12GB, not 24GB
Strata v0.1.39 (MIT, released Oct 4) runs Qwen3.8-Flash-Next — a 125B MoE with ~6B active params — on a single 12GB card: ~94 tok/s decode on an RTX 5070, 60 tok/s on an AMD RX 9070 XT. The earlier 4090 story needed twice the VRAM for a fraction of the speed.
Drop-in for your agents
It serves an OpenAI-compatible API at 127.0.0.1:8080/v1, Anthropic Messages at /v1/messages, and — new in v0.1.39 — the OpenAI Responses API at /v1/responses. Point Codex, a Claude client, or your own agent loop at localhost and run frontier-size inference offline with zero API spend.
How it fits in 12GB
A three-tier memory hierarchy makes it work: the GPU caches hot experts, the CPU scores cold ones over AVX-512/AVX2, and an n-gram table lives on SSD. Speculative decoding accepts 2.4–3.2 tokens per pass while staying bit-identical to greedy. Budget ~32–64GB system RAM and ~80GB of disk.
The catch
It's 2-bit (Q2_0/IQ2_XS), so expect quality and vision-accuracy hits; it handles one request at a time, greedy-only, with no KV reuse across turns and slow prefill past 128K context. Most throughput numbers so far are author- or community-reported, not independently audited.