Ollama v0.35 adds /v1/systemone for local decision models

It returns scores and choices, not text — drop the routing LLM. Plus LiteLLM per-second billing, Copilot CLI reads .claude/rules, a llama.cpp refactor.

Nowline SEP 29 12:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • A local endpoint for decision models, no tokens generated

    Ollama v0.35's new /v1/systemone serves classifiers and routers locally, returning choices, probabilities, and scores instead of generated text — with MLX backend support and flash attention auto-enabled on capable hardware. It's the serving layer for this week's wave of open decision models: swap a routing LLM call for a local score.

  • LiteLLM adds per-second pricing and live-session proxying

    LiteLLM v1.103/v1.104-rc adds per-second pricing for chat models — fixing double-billing on Bedrock — and proxies OpenAI's live sessions over WebSocket at /v1/live/sessions. Guardrail checks are now capped by request timeout, so a slow policy provider can't hang your call indefinitely.

  • Copilot CLI now reads your .claude/rules files

    GitHub Copilot CLI v1.0.90 picks up custom instructions from .claude/rules and follows your repo's pull-request templates when opening PRs. If you already maintain Claude Code rules, they carry straight into Copilot — one instruction set across two agents.

  • llama.cpp moves speculative decoding onto batch_ext

    Build b11236 migrates speculative decoding, multimodal (MTMD), and the server to the newer batch_ext API, plus robustness fixes across Vulkan, Metal, HIP, WebGPU, and OpenVINO. If you self-host, expect steadier multi-backend behavior and a cleaner path for draft-model speedups.

  • OpenCode hardens gateway timeouts and secret redaction

    OpenCode v1.18.33 now honors Cloudflare AI Gateway response and stream timeouts, surfaces MCP browser-launch failures instead of hanging, and redacts credentials and sensitive headers from debug output. Small fixes, but the kind that stop a wedged agent or a leaked token.