decision.host
首页 / 模型 / Microsoft-Decision-1

Microsoft-Decision-1

Microsoft ChoiceNoulScore systemonemodels.org

Microsoft-Decision-1 is Microsoft's System One model, post-trained from Qwen3.5-9B. It returns a calibrated probability for each fixed answer option to yes/no, multiple-choice and rating questions, through Microsoft Foundry at $0.042 per million input tokens. The weights are closed.

该模型目前没有出现在 S1MB 或 Decision Index 的公开评测结果中。

模型信息

厂商
Microsoft
类别
开源权重
参数规模
未公开
权重
闭源 / 未公开
许可证
未标注 许可不明
输入模态
文本
决策原语
ChoiceNoulScore
延迟
85-125 ms
可微调
未确认
上架状态
Generally available

来自 systemonemodels.org 的详细介绍

What Microsoft-Decision-1 is

Microsoft-Decision-1 is Microsoft’s decision model, and a System One model. Achint Srivastava, a VP of Software Engineering in Microsoft’s Office of the CTO, announced it on 9 October 2026. The Foundry model card gives 8 October as the release date. Microsoft post-trained Alibaba’s open-weight Qwen3.5-9B to score decisions in a single pass. It says it will later rebase the model on other models, including MAI and OpenAI’s. The weights are not distributed. It runs only as a hosted API in Microsoft Foundry.

What it returns

You send a context or state, a question and a fixed set of answer options. It returns a calibrated probability for each option and writes no text, not even an explanation. Microsoft lists yes/no, multiple-choice and rating questions, which match Noul, Choice and Score. It also grades AI responses and agent actions against a rubric. Input is text only, up to 32,768 tokens. Microsoft has not published the request format, the endpoint path, option or question limits, or rate limits as of 10 October 2026. So it is not known whether Foundry accepts Jev’s /v1/systemone request. Through OpenRouter, it takes OpenRouter’s Decisions API request, the same one OpenRouter serves Jev on.

What it is good at

Microsoft aims it at routing, classification, prioritization, verification and workflow control. That covers intent and model routing, confidence-gated actions, LLM guardrails and agent routing and skill selection. Input costs $0.042 per million tokens, the same as Jev, and output is free. The launch post’s accuracy and calibration figures come from Microsoft’s own tests. Across 36 benchmarks and 147,137 questions, it averages 83.5% accuracy, against 81.9% for Quyet-1.0-Large, 79.4% for GPT-6 Luna Decisions and 77.2% for H2O-Lightning-4B. Its calibration score is 92.2 out of 100, second to Quyet’s 93.1. Jev appears only in the latency chart, with no accuracy or calibration score. Microsoft measured a median of 85 ms per request through Foundry. It compares that with JevBench medians from a different harness, such as 240 ms for Jev 1.13.0. On requests altered in eight ways, it changes its answer 1.3% of the time on average, and never when options are paraphrased, reversed or shuffled. Xbox Research sorted more than 10,000 pieces of feedback and reviews with it and found it competitive with GPT-6 Sol at over 14 times the speed. It is not on Benchmark Heaven’s JevBench or the Decision Index as of 10 October 2026.

What it is not for

The model card rules out text generation, open-ended questions, conversation, translation and summarization, and any task without a closed question and fixed options. It reads no images, audio or video. Microsoft says it is tuned for English and may do worse in other languages and in medical, legal and financial work. It should not be the only basis for decisions about people, such as credit, hiring or housing. You cannot self-host it.

Access

It is generally available in Microsoft Foundry as a Direct from Azure model. OpenRouter also lists it as microsoft/microsoft-decision-1, served by Azure at the same price with a 32,768-token context, since 9 October 2026. Microsoft names no SDK for it. For side-by-side detail, see Jev vs Microsoft-Decision-1, Microsoft-Decision-1 vs Clef, Microsoft-Decision-1 vs pplx-decider and Microsoft-Decision-1 vs the OpenAI Decisions API.

Specifications

Question types Choice Score Noul Max Choice optionsNot documented Score levelsNot documented Questions per callNot documented Total context32,768 tokens State budgetNot documented Rate limitNot published as of 10 October 2026. SDKs The Foundry model card gives up to 32,768 input tokens, text only. A request holds a context or state, a question and a fixed set of answer options. Microsoft has not published the request format, the endpoint path, how many questions or options a request can hold, or how many levels a rating can have, as of 10 October 2026.

Versions

Microsoft-Decision-1, 8 Oct 2026, Version 1 in the Foundry catalog. The model card gives 8 October 2026 as the release date and September to October 2026 as the training time. It is a dense decoder-only model built on Qwen3.5-9B, in the 5B to 15B parameter range. Microsoft says it will rebase the model on other models, including MAI and OpenAI's. Release notes

Use cases

What people use Microsoft-Decision-1 for, one page per pattern. Workflow controlIntermediate

Intent and model routing with System One models

One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue. Choice Score Workflow controlIntermediate

Confidence-gated actions with System One models

Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars. Choice Score Safety and qualityIntermediate

LLM guardrails with System One models

Put one Jev request in front of an LLM and one behind it. Yes/no questions return the probability that each hazard holds, a Score rates how much harm complying would do, and your thresholds turn those numbers into pass, review, block, or a crisis path. Noul Score Real-time and agentsIntermediate

Agent routing and skill selection with System One models

An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess. Choice Noul