Microsoft Project Zenith: run 30B+ models locally on Windows 11
A ready-to-code Windows build, the dev stack preinstalled, a 64GB floor — plus the token-speed catch, the public config repo, and a reshuffled Index v4.2.

Copy markdown
Run 30B+ models locally, unmetered
Microsoft's Project Zenith (announced Sept 4) is a preconfigured Windows 11 build that promises local, unmetered inference of 30B+ parameter models. It sets a hardware floor — 64GB unified memory and 250 GB/s bandwidth — and launches first on AMD's Ryzen AI Halo silicon.
The token-speed asterisk
"Holding a model says little about how quickly it will respond." A mixture-of-experts model like Qwen3-30B-A3B (~3B active) runs 70–100 tok/s, but a dense 70B at 4-bit crawls near 5 tok/s; dense 7–13B lands around 30–45. Match the architecture to your box before you cut the API cord.
Copy the dev image without the $3,699 box
The image ships with VS Code, GitHub Copilot, Python 3.14+, Node 24+, WSL2/Ubuntu and .NET 10, plus sane defaults (visible file extensions, no forced AI) — and Microsoft published the same setup as its public Windows Developer Configuration repo, so you can apply it today. Hardware starts with Lenovo's ThinkCentre X Ultra ($3,699, Nov 2026); Nvidia's DGX Spark is $4,699.
Elsewhere: Intelligence Index v4.2 reshuffles the board
Artificial Analysis's Sept 4 refresh pushes private, held-out test sets to 40% of the Index weight (double v4.1), retires the saturated GPQA Diamond, and adds a 4,592-page long-context PDF eval. Claude Fable 5.1 leads the composite; GPT-6 Astra sits second, +4pt over 5.6 Sol. Better signal for picking a production model that can't be gamed to the benchmark.