Key Notes
- Decision-1 scores predefined choices with confidence estimates.
- Input costs $0.042 per million tokens and output tokens are free.
- Microsoft reports faster decisions in its own benchmark tests.
Microsoft has launched Microsoft-Decision-1, a small AI model built to choose between predefined answers and score its confidence in each option. The October 9 announcement targets the repeated decisions inside software workflows, from routing requests to checking an agent’s next action.
The model is available through Microsoft Foundry and OpenRouter, which lists input pricing at $0.042 per million tokens and no charge for output tokens. Its job is narrower than a conversational assistant’s: return a structured decision that an application can use directly.
How Microsoft-Decision-1 Works
Decision-1 is post-trained from Qwen3.5-9B and scores the supplied choices in a single pass. OpenRouter lists a 32,768-token context window and text input, with probability scores as the output. It explicitly excludes open-ended writing, conversation, translation and summarization from the model’s intended uses.
Consider a support inbox with three queues: billing, technical help and account access. A developer could provide a customer’s message and those options, then use the returned scores to select a queue. If two choices receive similar scores, the application could ask a person to review the case instead of making a confident-looking guess.
The developer still defines the choices and the consequences of selecting them. A classifier cannot route a request to a missing category, and a well-formed response does not prove that the selected action is appropriate. The useful distinction is between recognizing an option and giving software permission to carry it out.
What the $0.042 Price Means
At the listed rate, one billion input tokens would cost $42. As an illustrative calculation, a million requests averaging 1,000 input tokens each would reach that same input bill. This is a token-cost calculation, not a quoted price for running a complete business workflow.
Every request’s instructions, options and supporting material contribute to the input volume. Longer context and repeated checks therefore change the bill, while databases, application hosting and any other models remain separate costs. Free output tokens make the decision response inexpensive; they do not make the surrounding service free.
Microsoft Claims Faster Decisions in Its Tests
Microsoft says Decision-1 led its accuracy comparison across 36 benchmarks covering nearly 150,000 questions. In its latency testing, it reports median performance about 35x faster than GPT-6 Sol and 2.5x faster than H2O-Lightning-4B v1.1. These are Microsoft’s comparisons on structured decision tasks, not evidence of equivalent advantages across general AI work.
The company also tested whether small changes to a request altered the selected answer. Across eight kinds of perturbation, it reports an average decision-flip rate of 1.3%. That addresses a practical concern for software: changing an option’s position or wording should not arbitrarily change the result.
Production performance still needs a workload-specific check. Input length, network delays and the distribution of actual customer requests can differ from a benchmark. Confidence scores also need testing against observed outcomes before a team uses them to decide which cases can proceed automatically and which require review.
Decision Models Become a Separate AI Category
The launch follows Cloudflare’s Clef models, which also score predefined answers. Cloudflare offers both Clef and Clef-flash with open weights, alongside hosted access and business fine-tuning. The shared idea is to give applications a specialized decision component rather than ask a general assistant to generate and explain every selection.
For developers, that creates a more specific comparison: how reliably does each model distinguish the choices their product actually offers, and at what total latency and cost? A low inference price can make frequent checks practical. Whether those checks improve a service depends on the options, the test data and the rules that connect a model’s score to a real action.
Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.