Fable 5 closes 82% of the human gap on nanoGPT optimizer speedrun
Prime Intellect ran 100+ autonomous frontier-model runs on 8xH200s; Opus 5 and Kimi K3 trail, and full agent trajectories are open.

Copy markdown
Fable 5 hits 2,726 vs the 2,600 human record
The nanoGPT optimizer speedrun asks agents to iterate a training script for lower final loss. Fable 5 closed 81.7% of the gap; Opus 5 followed at 53.6% (2,920), Kimi K3 at 52.2% (2,930), against a 3,290 baseline and a 2,600 record built by dozens of researchers over months.
41 open agent trajectories to steal from
Prime Intellect published the full runs — tool calls, reasoning, experiment plans — as raw traces. If you're building a research agent, these are ~free R&D: read how the winners plan experiments, debug loss spikes, and decide when to stop iterating.
The compute spread names the price of ML-research agents
Runs used 26M to 2.9B tokens across 0.6-8.7 days of wall-clock on 8xH200s inside a 24-hour compute budget. Opus 5 alone burned 183M tokens across 292 experiments over 2.9 days — real numbers for planning your token caps and per-experiment budgets before you deploy your own.
Elsewhere: your local LLM is probably underperforming its weights
A deep-dive making the HN front page shows FlashAttention vs Triton backends alone cause top-1 token flips that cascade into tool-call failures. NVIDIA's NVFP4 hit ~50% token disagreement at 88k context while INT8 (W8A16) beat official FP8. Match the model card's sampler settings and skip INT4 KV cache for agents.
Elsewhere: Simon Willison flags Opus 4.8 still dominating Anthropic usage
Ramp's AI index shows Opus 4.8 at 28% of Anthropic model usage — the largest single share — suggesting builders haven't migrated to Opus 5 despite yesterday's GPT-5.6 Sol price cut. If you're still on 4.8, that's cover to stay; if you're on 5, expect the gap to persist.