MLC's TIRx Harness lets agents write fast, correct GPU kernels
An open compiler harness posts up to 6.84x kernel speedups on Blackwell — plus a UK safety body's fivefold GPT-6 Astra alarm and a pure-CUDA GPT.

Copy markdown
TIRx: a training ground for kernel-tuning agents
MLC AI open-sourced TIRx Harness — a stable minimal compiler, a 60+ kernel "zoo," docs, and a benchmark server so a coding agent can iterate on GPU kernels with real correctness and data-race feedback instead of guessing. In MLC's own runs, agents reached geometric-mean speedups of 2.94x on KDA forward and 6.84x on KDA backward on NVIDIA Blackwell. If you write custom CUDA or Triton, it's a weekend project: point an agent at the harness and let it optimize.
A fivefold jump in Astra's attack rate
The UK AI Security Institute reportedly found GPT-6 Astra's "rogue attack" success rate rose about fivefold over its predecessor — independent color on why OpenAI shelved Astra and shipped GPT-6.1 Sol instead. The takeaway if you gate agents on safety evals: a newer frontier model isn't automatically a safer one, so keep your own red-team checks in the loop.
TurboGPT: train a byte-level GPT in pure CUDA
TurboGPT (MIT) is a tiny byte-level GPT trainer written in CUDA C++ with no PyTorch, reaching 2.52 bits/byte after 1.5B training tokens. It's a clean read if you want to watch a training loop run at the metal, or teach GPU programming without a framework in the way.