Strata runs a 125B model on a 12GB gaming GPU—no cloud

MIT-licensed, it shards a 125B MoE across GPU, RAM and SSD at ~94 tok/s. Also: Anthropic's Mythos finds an HFS bug that's exploited within a day.

Nowline OCT 4 8:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • A 125B model on the GPU you already own

    Strata, released today under the MIT license, runs Qwen3.8-Flash-Next — a 125B mixture-of-experts — by routing each token through just 10 of its 24,576 experts, keeping the hot ones on a 12GB GPU and offloading the rest to 32GB+ of system RAM. On an RTX 5070 the project reports ~94 output tokens/sec behind a one-click Windows/Linux installer that exposes OpenAI- and Anthropic-compatible endpoints on localhost. The numbers are self-reported, but it puts frontier-class inference on hardware you already own — no API bill, no data leaving the box.

  • AI-found bugs now get weaponized in a day

    Anthropic's Mythos bug-hunting model flagged CVE-2026-61500, a session-forgery RCE in Rejetto HTTP File Server, and attackers had a working exploit within roughly 24 hours of disclosure. If you run HFS anywhere, upgrade to 3.2.1 now — but the real signal is that the window between a bug being found and being exploited is collapsing as models get better at finding and chaining them.