Gemini 4 Argon: Google's new frontier model, cyber-defenders first
It tops the Vals Index over Opus 5.5 with a 1M-token in/out window, but broad API access isn't open yet; llama.cpp also ships local decision models.

Copy markdown
Gemini 4 Argon leads the benchmarks — if you can get it
Google's first new frontier Gemini since 3.1 Pro takes #1 of 43 on the Vals Index (68.9%), past Claude Opus 5.5's 66.97%, and pairs a 1M-token input with a 1M-token output plus “Long Decode Continuation” to resume long generations across calls. It logs a 15% hallucination rate (vs GPT-6 Astra's 51%) and a 0.7% prompt-injection success rate, at an intro $2/$10 per million tokens (later $4/$20).
But you probably can't call it yet
Access is gated: at launch Argon went first to 650+ “trusted cyber defenders” via Google's Fairwind Program, and the public Gemini API was still returning 404s. Google says paid API customers and AI Ultra subscribers are next, with no date — so keep shipping on Astra or Opus and treat Argon as watch-this-space.
llama.cpp adds decision models you run locally
A new /v1/systemone endpoint (live now via `llama update`) answers by scoring the options you pass in a single forward pass instead of generating text — for routing, moderation, or verifying an agent's next action. Five Apache-2.0 models span 144M to 27B params at 3–43ms, so a router or classifier runs locally with no text-gen round-trip or API bill.
GitHub opens up its security-advisory APIs
GitHub shipped a public-preview REST API to read, add, and edit repository security-advisory comments, plus five new SecurityAdvisory GraphQL fields (CVE ID, source location, review dates, advisory URL) with server-side severity filtering. If you automate vuln triage, you can pull and annotate advisories programmatically instead of scraping the UI.