Jeff: open 0.8B–2B zero-shot classifiers, ~22ms, train at home
Fine-tunes of Qwen3.5 and Gemma 4 return calibrated probabilities in one forward pass. Plus: Codex CLI 0.158.0's OAuth and a 7-chip ESP32 BitNet cluster.

Copy markdown
A 0.8B that classifies in ~22ms, no tokens generated
Firelex's Jeff is a family of open decision models — Jeff-Qwen3.5-0.8B and -2B, plus Jeff-Gemma4-E2B — that pick among your options and return calibrated probabilities in a single forward pass, no text generation. The 0.8B runs ~22ms on an RTX PRO 6000 and 28ms on an M4 Max (463ms on a 32-thread CPU) at 79.1% across five benchmarks; the 2B hits 83.1%, matching Jev's published 83.0%. Weights are Apache-2.0, code is MIT, and the request format is Jev-compatible, so it drops into classification code you already have.
Retrain it on your own labels in about 2 hours
Jeff is full-weight fine-tuned for a single epoch, so you can retrain it on your own categories on one workstation GPU — the 0.8B in ~2 hours, the 2B in ~3.5. A voice-navigation example jumped from 31.7% zero-shot to 95.8% accuracy in under 30 minutes. It's aimed at support-queue routing, user-intent detection, moderation labels, and game moves, where spending a full LLM call per item is overkill.
Codex CLI 0.158.0: OAuth secrets and Markdown-safe paste
OpenAI shipped Codex CLI 0.158.0 (npm i -g @openai/codex@0.158.0). MCP servers that require pre-registered OAuth client secrets now connect, exec-server WebSocket links are secured with bearer tokens, and the fullscreen TUI adds copy-on-select plus right-click paste that preserves Markdown formatting. Image edits can request transparent backgrounds and reference file-backed conversation images. Heads-up: terminal approvals for elevated-permission commands are now enabled by default.
A 0.5B LLM sliced across seven ESP32-S3 chips
Low-Zi-Hong's ESP32s3-LLM-Cluster runs a ~0.5B model across seven ESP32-S3 boards — one master handling tokenization and embeddings, six compute nodes running attention and MLP over a high-speed SPI daisy-chain. It uses 1.58-bit BitNet ternary weights with assembly-optimized ops and a 32K BPE vocabulary, on ESP-IDF and MIT-licensed. No throughput numbers published yet, but it's a concrete blueprint for LLM inference on ~$8 microcontrollers.