Poolside's Laguna S 2.1: an open-weight 118B coder you can self-host
It matches models several times its size on SWE-Bench yet runs on a single box — as Google reportedly bakes Gemini into custom silicon to slash serving costs.

Copy markdown
One box, no meter
Poolside open-sourced Laguna S 2.1, a 118B-A8B mixture-of-experts model for agentic coding under the permissive OpenMDW-1.1 license. The weights are on Hugging Face and it runs on a single NVIDIA DGX Spark, so you can self-host agentic coding instead of renting it by the token — with API access via OpenRouter, Poolside's own API, and its Pool agent harness.
Punches 10x its weight
It scores 70.2% on Terminal-Bench 2.1, 78.5% on SWE-Bench Multilingual and 59.4% on SWE-Bench Pro — which Poolside says matches or beats DeepSeek-V4-Flash, NVIDIA's Nemotron 3 Ultra and Thinking Machines' Inkling, all several times larger. It even independently proved Erdős Problem #397, and a lighter Laguna-XS-2.1 is up for smaller rigs.
Report: Gemini etched into silicon
The Information reports Google is building "Frozen v2," a chip that bakes Gemini's architecture directly into hardware for a claimed 6–10x jump in serving efficiency over today's TPUs. It targets 2028 and internal use, but far cheaper Gemini inference could eventually pressure the OpenAI and Anthropic pricing you build on.