Microsoft's Decision-1 Is Built on Qwen 9B

On October 9, Microsoft published Microsoft-Decision-1 in its own technical publication, Command Line. It is a model built for one narrow job, making choices: routing, classification, ranking, validation and workflow control. The byline belongs to Achint Srivastava, a vice president of software engineering in the Office of the CTO at Microsoft. The model is available on Microsoft Foundry and has also been added to OpenRouter.

Its output is deliberately narrow. The caller supplies a fixed set of options, and the model returns a calibrated probability for each one, which software can act on directly. Yes/no questions, multiple choice, scoring and rubric-based grading are all supported. Writing long text and complex reasoning are not part of its remit.

Microsoft's self-reported results

Microsoft says Decision-1 posted the highest accuracy across 36 benchmarks and nearly 150,000 questions, with a P50 latency of roughly 1/35 that of GPT-6 Sol. The article gives an example: in a pipeline that chains 20 decisions, an extra 100 milliseconds per decision adds 2 seconds to the whole chain.

On stability, the same request was perturbed in 8 different ways, and decisions flipped 1.3% of the time on average. When the options were only reworded, reversed or shuffled, the flip rate was zero. Microsoft states the principle behind this test as follows:

"Equivalent inputs should produce equivalent decisions."

Safety testing covered 11 benchmarks and 5,250 requests, including harmful content, jailbreaks and prompt injection, but the article does not give a refusal rate.

Internal trials produced several more figures. The Xbox research team used it on more than 10,000 pieces of user feedback and reported quality on par with GPT-6 Sol, at over 14 times the speed and over 200 times lower cost. The Copilot team measured quality comparable to GPT-5.6 Luna at 100 times the speed. All of this is Microsoft's own testing. The article does not name the 36 benchmarks or list the specific accuracy figures, and no third party has reproduced the results so far. As for which model ranked second and how much slower it was, Microsoft's original post and the Chinese-language retelling do not match, so readers should wait for any official update.

Pricing follows usage

The price is $0.042 per million input tokens, with output free, which IT Home converts to roughly 0.28 yuan. Output from decision-type calls is often only a few tokens, so most of the spend goes on input context, and the pricing fits how the model is actually used.

In an application, a model like this typically sits in front of the large model. It decides which model or workflow a request should go to, or judges whether an agent's output at a given step is acceptable before deciding what happens next. These judgments were often handled by a general-purpose large model as a side task. Swapping in a small, fast specialist saves latency and per-call fees.

The base comes from Qwen

One line in the article carries real weight: Decision-1 was post-trained on Alibaba's open-source Qwen3.5-9B. Microsoft also says it will carry the same method over to other bases, naming Microsoft AI's (MAI) in-house models and OpenAI's models.

For developers in China, this can be read in three ways. First, Qwen's open weights have made it into an official Microsoft product, and Microsoft credits the source in its launch post. Second, the path is now largely laid out for teams that want to build something similar: take an open-source model at the 9B scale and post-train it around "fixed options plus calibrated probabilities," so a single forward pass yields the result. Third, Decision-1 itself does not have open weights, so users in China can only reach it through the Foundry or OpenRouter APIs, while its base, Qwen3.5-9B, can be downloaded directly.

Microsoft has rolled out several in-house MAI models this year. Decision-1 starts from an external open-source base, and whether its results hold after a switch to an MAI base will depend on the numbers in the next release.

Sources: Microsoft Command Line announcement, CocoLoop, IT Home; benchmark count, latency multiple and per-million-token price checked against Microsoft's self-reported figures.