Mercury 2.5: a diffusion LLM at 1,100 tok/s, $0.04 per million

Builders are finding diffusion models: sub-170ms first token unlocks live voice and RAG. Plus AWS's open agent harness and OpenAI DevDay this Sunday.

Nowline SEP 26 6:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The fastest LLM you haven't tried is a diffusion model

    Inception's Mercury 2.5 generates around 1,100 tokens/sec with sub-170ms time-to-first-token — fast enough for live voice and interactive RAG where transformer models stall. It's billed as the largest diffusion LLM trained yet, ships a 260K context window, and a launch discount lands it at $0.04/M input and $0.15/M output on OpenRouter and Baseten. It hit the Hacker News front page this week as builders clocked the speed.

  • AWS's open-source agent harness runs any model for 28% fewer tokens

    Strands Harness is a bring-your-own-model agent runtime you can deploy anywhere, in Python or TypeScript, with no AWS lock-in. AWS's team reports 28% lower token cost at comparable accuracy versus other harnesses — real savings once an agent loops for minutes at a time. It's open source and installable today.

  • Elsewhere: OpenAI DevDay lands this Sunday, Sept 29

    OpenAI's biggest developer event runs Sept 29 in San Francisco with a livestream and roughly a dozen announcements expected. Reports point to a preview of a GPT-6 cybersecurity model and wider access to the Astra tier — worth watching if model access or pricing shapes your stack.