Qwen3.7 Flash: a 1M-context vision model at $0.03/M input

Alibaba's cheapest tier yet is multimodal, with tool-calling and prompt caching on OpenRouter and QwenCloud — but closed weights and unbenchmarked.

Nowline JUL 30 11:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The price that rewrites the math

    $0.03 per million input tokens and $0.13 output, on a 1M-token window. That's flash-tier pricing on a multimodal model — cheap enough to run vision over thousands of screenshots, PDFs, or video frames in a loop without watching the meter.

  • It sees, not just reads

    Qwen3.7 Flash is a vision-language reasoning model: object recognition, spatial understanding, visual coding, and computer/UI interaction. Screenshot-to-code and browser-automation agents just got a cheap backbone.

  • Full agent toolkit, API-only

    Function calling, built-in tools, structured output, and prompt caching are all live on OpenRouter and QwenCloud. The catch: Alibaba didn't publish weights, so there's no local or offline path — you rent it, you don't own it.

  • The asterisk: it's unranked

    There's no public benchmark table yet, so the model sits unranked on the usual leaderboards. Cheap isn't the same as accurate — run it against your own evals before you wire it into a pipeline.

  • Build this weekend

    Point an agent at a folder of receipts, invoices, or app screenshots and have it emit structured JSON — at $0.03/M in with a 1M window, a batch that used to cost real money now rounds to nothing. Prompt caching makes repeated system prompts cheaper still.