Microsoft-Decision-1 Scores the Choices an Agent Has to Make
Microsoft put a 9-billion-parameter decision-scoring model in Foundry that returns a calibrated probability for each option in a fixed list, and priced it at $0.042 per million input tokens.
TL;DR
- Microsoft-Decision-1 is a small language model that returns a probability score for each option in a list you hand it, so a program can route a request, grade a response or decide what an agent does next.
- Microsoft published the model on 9 October 2026 in Microsoft Foundry, with access through OpenRouter planned.
- The model was built by post-training Qwen3.5-9B for single-pass decision scoring, and Microsoft states it will rebase the model on other systems, including Microsoft AI and OpenAI models.
- Microsoft reports the highest accuracy in a 36-benchmark comparison covering nearly 150,000 questions, and it reports P50 latency about 35x faster than GPT-6 Sol.
- Microsoft names five engineering problems it had to solve: speed, quality that generalises, robustness across paraphrased inputs, probability calibration, and safety.
- Microsoft's internal tests include Xbox Research sorting more than 10,000 pieces of feedback, the Copilot team measuring response quality, and Microsoft Discovery using the model to grade experiments.
- Pricing starts at $0.042 per million input tokens with output tokens free, which is the number that decides whether high-volume routing and grading become affordable.
- A team adopting it needs its own labeled set of tickets, a baseline to compare against, and a look at which errors cost money, because a wrong "proceed" and a wrong "escalate" carry different prices.
What Microsoft shipped in Foundry
Microsoft's launch post on the Command Line publication describes a model built for a narrow job. "Unlike LLMs, which are designed to generate text or reason through complex problems, decision models are purpose-built to deliver structured outputs that software can immediately act on," writes Achint Srivastava, a Microsoft vice president in the Office of the CTO. The model is available in Microsoft Foundry from publication, and the post adds that OpenRouter access is coming.

The shape of the thing explains the use cases. Given a fixed set of answer options, Microsoft-Decision-1 provides a calibrated probability score for each one, which means the application owns the threshold and the model owns the score. Routing a support ticket, checking a proposed tool call against a rubric, and grading an AI response all reduce to the same call: a situation, a closed list, and a number returned for every option on it.
The job is smaller than the one a chat model performs, and Microsoft prices it that way. Windows Report records the rate as $0.042 per million input tokens with output tokens free, and the same report notes that the pricing and the speed figures come from Microsoft's own benchmarking. A rate that low changes arithmetic in ordinary places. A support desk that scores 200,000 inbound messages a month spends under nine dollars on input tokens at that rate, which turns a category of check that teams skip for cost reasons into one they can run on every message.
How a scoring model fits beside a generative one
The architectural reading is the useful part, and it is mine rather than Microsoft's. A pattern already common in production agents puts a cheap judge next to an expensive writer: the large model drafts, and the small model answers one bounded question about what happens next.
Microsoft supplies the argument for the second seat in its own post, with a latency example that needs no benchmark to land. "For example, adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow," the post states. Agents make exactly that many dependent decisions in a support or finance workflow, and each one inherits the delay of the model behind it.
That arithmetic has a second edge that the post leaves for the reader. A general-purpose model asked to score a fixed list still pays for the reasoning it produces along the way, which is the cost a dedicated scoring model removes. The engineering value of a decision model is the removal of generated text from a step where nobody reads the output.
Microsoft's approach to the harness around that step is where our own work sits: we build AI agent harnesses with a swap-ready model layer, and a scoring model that speaks a structured API fits that layer without touching the tools, the retrieval or the approval rules around it.

The five engineering problems Microsoft names
The post lists the difficulties in the order a team would hit them, and the list reads as a specification for anyone building the same thing. Speed comes first, then quality that generalises, then robustness, then calibration, and finally safety.
Quality across unseen tasks is the hardest of the five to buy, and Microsoft describes the method it used to argue the point: dozens of benchmarks kept blind from training, spanning routing, ranking, long context, multilingual and out-of-distribution tasks, plus a set of top public models from the JevBench leaderboard tested across 36 further benchmarks. The claim that follows is a generalisation claim, which is the claim a small model has to earn.
Robustness gets a number that a team can reproduce. "Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled," the post states, describing eight ways of rewriting the same request. For anybody setting an escalation threshold, that rate is the interesting one, and the post frames the calibration goal in operational terms: "Applications use confidence to decide when to act, defer, or ask for review, so a 90% prediction should be right about nine times out of 10 on representative cases."
The internal tests Microsoft published
Vendor benchmarks describe a model on a leaderboard. Internal deployment describes a model inside a workflow, which is a different kind of evidence, and Microsoft publishes four examples of the second kind.
"XBOX Research used Microsoft-Decision-1 to process more than 10,000 open-ended pieces of feedback and reviews from surveys, STEAM, and Twitter/X and sort them into a fixed set of themes established by researchers," the post reports, adding that the team found the model "to be competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive." The Copilot team, measuring chat and agentic response quality, found it "competitive with GPT5.6 Luna and 100 times faster." The on-call engineers use it for knowledge retrieval during live incidents. Microsoft Discovery uses it to grade experiments inside an adaptive replanning loop, where the post reports the model scoring "46 times more consistent than the LLM-based score at three times the speed."

Read those four together and a pattern shows up that is worth copying. Every internal example has a fixed set of labels that predates the model: themes chosen by researchers, a quality rubric, a relevance judgment. The model performs the sort. Somebody else decided what the categories mean.
A test plan for your own tickets
Microsoft has not published a deployment procedure, so the evaluation work stays with the buyer. The plan is short enough to run in a week.
Pick one narrow task with a real cost attached: triage on a support queue, categorisation of customer feedback, or pass-fail grading of an AI draft against a rubric. Build a labeled set of a few hundred of your own cases, because vendor benchmarks describe their edge cases and your tickets carry yours. Run the new model against whatever scores those cases today, and compare accuracy, latency and cost on the same set.
Then read the errors by direction rather than by average. A wrong "proceed" and a wrong "escalate" are different events with different prices, and a threshold set on a headline accuracy number hides that asymmetry. Rephrase the same inputs and watch how often the decision moves, since the vendor's 1.3% figure is a target to reproduce rather than a fact to assume. Keep a person on the consequential actions until your own numbers exist, and log every score next to the outcome it predicted, because that log becomes the calibration data the next decision depends on.
Pricing and the rebase path
Two details in the release shape the planning. The first is the price, $0.042 per million input tokens with output tokens free, which puts a bounded cost on a check that runs on every message or every tool call. The second is the base model. Microsoft states it post-trained Qwen3.5-9B for this job and will soon rebase the model on other systems, including Microsoft AI and OpenAI models.
A rebase is a behaviour change wearing a version number. Scores that a threshold was tuned against can move when the underlying model changes, so a team that writes down its thresholds should also write down the model version those thresholds belong to. That habit costs nothing and it protects the one thing a decision model has to keep: the meaning of a score across releases.

Microsoft put a cheap, calibrated score where an agent's next move is chosen, and published numbers from its own tests next to a price that makes the check affordable at volume. The work left to a buyer is small and specific: your cases, your baseline, your thresholds, and a model version recorded beside them. That is the whole difference between a scoring model in a diagram and a scoring model in production, and the difference is measurable in an afternoon at Netics.
Sources
Source: Introducing Microsoft-Decision-1, our model for fast decision-making — commandline.microsoft.com/microsoft-decision-1-model-foundry/, Microsoft (Command Line), published 9 October 2026, retrieved 2026-10-10 (the availability in Microsoft Foundry and through OpenRouter, the design targets of routing, classification, prioritization, verification and workflow control, the post-training of Qwen3.5-9B with plans to rebase on Microsoft AI and OpenAI models, the calibrated probability score for each option in a fixed set, the support for yes/no, multiple-choice and rating options with rubric-based grading through a structured API call, the highest accuracy in the 36-benchmark comparison across nearly 150,000 questions kept blind from training, the 2.5x and 35x speed comparisons against H2O-Lightning-4B v1.1 and GPT-6 Sol, the P50 latency figure, the 100-millisecond per-decision example with 20 sequential decisions, the 1.3% perturbation result with zero flips for paraphrased descriptions and reordered options, the calibration statement about a 90% prediction, the 5,250 safety requests across 11 benchmarks, and the internal results from Xbox Research, the Copilot team, incident response and Microsoft Discovery).
Source: Microsoft-Decision-1 AI Model Is Here With 35x Faster Performance Than GPT-6 Sol — windowsreport.com/microsoft-decision-1-ai-model-is-here-with-35x-faster-performance-than-gpt-6-sol, Windows Report, published 9 October 2026, retrieved 2026-10-09 (pricing starting at $0.042 per million input tokens with output tokens free, the planned OpenRouter support, the targeting of classification, prioritization, AI response evaluation and agent workflow control, and the statement that the figures come from Microsoft's own benchmarking).
Source: Microsoft's Command Line launch post for Microsoft-Decision-1 — commandline.microsoft.com, 9 October 2026; Windows Report's report of the same launch — windowsreport.com, 9 October 2026. Both retrieved 2026-10-09. Figures: the official launch card and the post capture from commandline.microsoft.com, plus two Netics editorial diagrams rendered from the same post.