Microsoft-Decision-1 is GA: calibrated scoring, free output tokens
Microsoft's first decision model goes GA in Foundry and on OpenRouter at $0.042/M in — a drop-in scorer for routing, classification and LLM-as-judge.

Copy markdown
A scorer, not a generator — live in Foundry and OpenRouter
Microsoft-Decision-1 is generally available now in Microsoft Foundry and on OpenRouter. Instead of emitting text, it takes a fixed set of options — yes/no, multiple-choice, or a rating rubric — and returns a calibrated probability for each in one structured API call, built for routing, classification, verification, safety screening and grading agent actions. It's a post-trained Qwen3.5-9B, with rebases onto Microsoft's MAI and OpenAI models promised next.
The price is the point: $0.042/M in, output free
Input runs $0.042 per million tokens and output tokens are free — which rewrites the math on any pipeline burning a generative LLM as a judge or router. Microsoft's Xbox team labeled 10,000+ feedback items at quality competitive with GPT-6 Sol but ~14x faster and roughly 1/200th the cost; its Copilot team clocked it ~100x faster than GPT-5.6 Luna on comparable grading. If you run eval harnesses or classification at scale, that's a direct cost lever.
Top of a 36-benchmark board — with the vendor asterisk
Microsoft reports the highest accuracy across a 36-benchmark suite of nearly 150,000 held-out questions, ~35x faster than GPT-6 Sol at P50 latency, and decisions that flip on just 1.3% of perturbed inputs (paraphrases, reversed or shuffled options caused no flips). Every figure is vendor-reported and awaits independent testing — but the consistency claim is the one to watch if LLM-as-judge flakiness has burned you.