Rendered at 18:34:02 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
nejch 23 hours ago [-]
This is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider, and a plethora of others released in the last couple of weeks.
Kind of funny how much hype they can all get out of this, but Qwen really is the little engine that could. Great to see open weights (if not open source) driving the whole ecosystem like this though.
sidd0103 20 hours ago [-]
Plus, a lot of IP from a previous startup (that was acquired last year), Pi Labs, was used to help push the model to SOTA performance quickly. The startup was building scoring models long before Jev released
manmal 22 hours ago [-]
My biggest learning after some experiments - a BF16 (unquantized) Qwen beats a Q8 of double its size for decisions. I guess that’s the reason Kev switched to 4B BF16, from the original 8B version. Isn’t it interesting that quantization seems to mess with decision accuracy?
girvo 22 hours ago [-]
That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though
Super cool finding!
ByteAtATime 21 hours ago [-]
Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do
anewhnaccount2 20 hours ago [-]
Yes this is the reason. Transformations like quantisation preserve the rank of outcomes much better than probability mass.
manmal 9 hours ago [-]
FWIW, my tests used only ranking. I used Kev to drive E2E test runs using agent-device, where Kev had to decide the next action - like pressing a button, or scrolling down etc, to reach a certain goal (UI state). It went surprisingly well, but 4B BF16 was significantly better than 8B Q8. The latter made some wrong decisions in every run while the former was almost at Luna precision.
_menelaus 7 hours ago [-]
Fascinating insight. I'm curious, how do you know this?
RussianCow 21 hours ago [-]
I think it also helps that speed and latency aren't as vital for coding and we can afford to let the model think for longer.
giancarlostoro 21 hours ago [-]
Surprised they aren't doing these sorts of one-off models with Microsoft Phi, which is intentionally smaller, but there's no reason Microsoft couldn't try to make a slightly larger Phi model with more capabilities...
throwa356262 11 hours ago [-]
IIRC Phi has a tiny context windows.
And also, it is nowhere as good as Qwen
mjb 19 hours ago [-]
> This is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider
I don't know if Fabio (who's leading work on this model line) would agree, but if I was going to start over I'd probably pick Gemma4 as the base rather than Qwen3.5. But it's not surprising to see a lot of Qwen competition, and I totally agree it's nice to see the work open.
nowittyusername 18 hours ago [-]
I been using gemma 26b moe model as the system one variant for my own uses and its great and far smarter then any alternatives at around 200ms, bigger is better for these things as long as you dont need lower latencies, only caviat is calibration, if you need calibration better use JEV.
nonatofabio 19 hours ago [-]
Yep, Fabio here, and I agree!
NitpickLawyer 22 hours ago [-]
> Great to see open weights (if not open source)
The insistence of naming it open weights as opposed to open source is getting ridiculous, and it's both irrelevant (i.e. no one cares in practice) and factually incorrect.
Weights are source in language models. Apache defines source as ""Source" form shall mean the preferred form for making modifications". That is precisely what's happening here. Everyone is using the preferred form for making modifications to these models (including the model creators themselves). A model is "created" at init time, and then "trained" by modifying the weights.
All these models are open source. What's not open sourced (with qwen et all) is the training code. So open source model, no training code. And that's ok. There are labs that release those as well. Apertus and Olmo series come with open source models, open source training and open datasets. Nemotron comes with open source models, open source training and some open datasets, while others are not published. And that's ok too.
The fact that you see all these models being modified (from AR completion models to "decision models") and re-released should be all the proof you need. That's what a license offers you. The right to inspect, run, modify and re-release a model. A license cannot (and never did) give you any other rights. OpEnWeIgHtS is silly.
teruakohatu 22 hours ago [-]
> Weights are source in language models.
Weights are source in the same way as any x86 binary is source.
You easily modify a x86 binary and change behaviour or examine the machine code instructions. You probably are not aware how easy it is to change the behaviour of a binary executable.
gunalx 19 hours ago [-]
But, binary is not the preferred form.
Also if compiling literary costs millions of dollars, I would prefer the precompiled one.
nejch 22 hours ago [-]
I mean I don't have a strong opinion but if I used the phrase open source models there'd be 5 comments going in the other direction.
I'm happy to have and be able to serve these models and see the ecosystem thrive. And lots of open innovation is outside of weights anyway as DeepSeek repeatedly shows.
mh- 19 hours ago [-]
You used the right term. I don't know what the parent commenter is on about.
Open source has historically meant you could download the source and build your own binary. The appropriate analogue here, IMHO, to the build->binary process is training->weights.
Rohansi 22 hours ago [-]
> OpEnWeIgHtS is silly.
Dictionary.com defines source as:
> any thing or place from which something comes, arises, or is obtained; origin.
chris_money202 24 hours ago [-]
Microsoft is doing things differently with AI. It feels to me they are moving into local inference heavily and see a future where Windows has native AI APIs that run locally or optionally in the cloud/edge.
Nice didn't know that. I was thinking lower APIs similar to directX for gaming
doomroot13 19 hours ago [-]
This is actually the case. Windows ML (https://github.com/microsoft/windowsML) is the inference framework wtih vendor agnostic support for inference on CPU, NPU, and GPU. They also announced quite a bit more including an isolation solution with MXC. There's a lot of marketing fluff in the below link but it covers the recent announcements.
https://blogs.windows.com/windowsexperience/2026/10/07/build...
ampersandwhich 23 hours ago [-]
I hope they can finally make my "Copilot+ PC" infer things locally that are actually useful. Phi Silica for Advanced Paste was a good start, if a bit late. If they got their act together, Microsoft-Decision-1 could have some local potential. Their track record leaves me with some reservations.
withinrafael 23 hours ago [-]
With Copilot+ PC branding already retired, I suspect we won't be seeing much more activity on that front.
doomroot13 19 hours ago [-]
I don't think that's actually the case. There were some rumors flying around about this but in their recent event they actually referred to Copilot+ PCs as generally the line targeting more casual users with support for less powerful local inference. And the new devices based on Nvidia RTX Spark (and likely the more powerful solutions from AMD, Intel, Qualcomm with large unified memory) as a "Builder" class targeting developers and heavy local inference users. They also revealed a quantized version of their coding model designed to run on these devices.
So it does seem they may be somewhat working from the bottom up building smaller models or focusing on capable local inference and balancing with more powerful frontier model access.
This is a lot of marketing speak but covers much of what they revealed.
That’s where Apple is moving to as well. The models doing the implementation work need not be better than opus 4.6. And locally available hardware to run this already exists and likely will be sub 2k of 2026 dollars in a few years time.
jjcm 18 hours ago [-]
MAI-Image-2.6 is a really, really solid image model. I was surprised at how good the models from MS are getting. They're doing good stuff over there.
nxobject 23 hours ago [-]
Honestly, I'm glad that people later to the AI game are exploring niches other than state-of-the-art "smartest" models – I'd love AI applications that tackle the small hassles in life.
a_vanderbilt 23 hours ago [-]
Yeah, because they are desperate to try and justify the investments into Copilot and the NPUs they pushed OEMs into integrating. I'm all for competition, but every one of MS's AI models have just been nothingburgers or relabels of other lab's models. Even their novel high cardinality models are just novelties.
chris_money202 22 hours ago [-]
It depends on what you're doing. None of their models are Opus level, but not everything needs Opus. Their models are targeting cheap and useful for some common things not expensive and useful for anything
17 hours ago [-]
Topfi 18 hours ago [-]
In my minimal suite it was cheaper (by 0.72x) but higher latency (283ms vs 369ms p50) than Jev. Results were very comparable across all scenarios I measure, first of these that I have tested that actually justifies its existence as a commercial release.
Qwen tunes are nice and all but either price or performance makes each example I have tested not viable unless you are able to cheaply self-host and fine-tune further.
chrisandchris 14 hours ago [-]
Will they use it to decide how to name Office, ehm Office 365, ehm Microsoft 365, ehm Microsoft 365 Copilot next?
davidmurdoch 7 hours ago [-]
It's for Clippy. Bring back Clippy!
beng-nl 2 hours ago [-]
Oh no.. careful what you wish for! A clippy maximizer!!
mrbonner 18 hours ago [-]
I just don’t understand the use of a decoder model to make a classifier. Jev model will give you probability scores for each of the classes/selections you ask for. The probability is actual statistical probability that a selection is right.
Using a decoder model for this comes down to picking the most probable class based on the probability what the next token assigned to a class is. They are all probabilities but semantically mean totally different things. Am I right?
mrbonner 18 hours ago [-]
I also didn’t see Ms mentions that they use Qwen as the base model. Somehow I saw a comment here and thought it was the case. Regardless, my parent comment is about other decision models that based on decoder model.
jitbit 6 hours ago [-]
They do mention it in the article
“To build Microsoft-Decision-1, we post trained Qwen3.5-9B”
fredsmith219 20 hours ago [-]
The article talks about using the decision model in code, but could it be used to help indecisive people with everyday life, decision decisions? I know a few and they could really use help.
wolttam 20 hours ago [-]
LLMs already do that but the advantage they have is being able to walk you through a plausible explanation for why they’re offering one position over another.
Using the output of a “decision” model without insight into the reasoning for a given decision seems very trusting.
adilkhanovkz 5 hours ago [-]
[flagged]
prometheus1992 22 hours ago [-]
why wouldn't they benchmark the accuracy against jev too?
sidd0103 20 hours ago [-]
Interestingly, I've seen better performance and similar cost to whats on JevBench
tomalbrc 19 hours ago [-]
Is this an astroturfing account?
sass1caia 21 hours ago [-]
They say they are only benchmarking public models in the blog.
Hmm, so actually I thought it would say that it's not permitted to benchmark or compare to other products, but I can't find such claim?
It does say "develop (or to facilitate the development of) a similar or competing product or service", but I think it would be a long stretch to say that's the case if they would just publish benchmarks. Microsoft legal department might disagree.
simonw 23 hours ago [-]
> To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.
I guess the fear of Chinese models is finally subsiding.
dr_kiszonka 22 hours ago [-]
There isn't much choice, I am afraid. It is Chinese models or Gemma or Llama? (I am skipping a few lesser known ones.)
HarHarVeryFunny 21 hours ago [-]
Looks like yet another non-price-competitive Jev competitor.
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
Well let's see if the Chinese open weight models come in with even cheaper models. Deepseek came out of Algo traders they probably already have small fast decision model they use internally
HarHarVeryFunny 4 hours ago [-]
Open weight model pricing isn't really up to the company that built them (unless they are also serving them, in which case their API price may differ). It's really up to the serving cost and/or pricing of whoever is actually serving the model.
The main market for Jev is going to be business, who may not want to deal with a low cost open weights server like Fireworks AI and are more likely to go with a provider like AWS where pricing is higher.
wkcheng 23 hours ago [-]
I don't see any API documentation for this yet. How can someone actually try it? Did they rush this out for hype?
Maybe I'm just not looking in the right place, but I cannot find what the API shape looks like. I've even deployed this model via Foundry and it doesn't say what to POST or what to expect back.
buredoranna 20 hours ago [-]
clippy! is that you!?
MisterMunchkin 19 hours ago [-]
State: “I shit my pants and now my pants have shit in them”
Question: “Which team should handle this message?”
Result: “Tech Support (85%)”
Yep, sounds about right.
phoghed 16 hours ago [-]
Under 90% you have to fall back to Astra or Fable for a critical business case like this
elzbardico 19 hours ago [-]
In related news TypeSafe AI just raised a ginourmous amount of money.
23 hours ago [-]
bflesch 23 hours ago [-]
While this looks like a contribution from a capable team trying to impress senior leadership, for me personally the Microsoft brand is so badly tarnished I don't even feel negative emotions any more - just pity.
TokenLat 15 hours ago [-]
[flagged]
hulitu 8 hours ago [-]
> Microsoft-Decision-1, our model for fast decision-making
I'm still waiting for the Microsoft Vacuum Cleaner. /s
tencentshill 23 hours ago [-]
But what about when the government's AI skills amount to: "is this DEI, only answer yes or no"
No need to use AI, this was done by ctrl-f on dangerous terms such as "disabilities", "trauma" and "female".
(No, I am not making this up)
hollow-moe 21 hours ago [-]
what in the michaelsoft binbows? micro$oft actually naming a product clearly and concisely? Is the team office hidden in a far building wing that marketing hasn't found yet?
Kind of funny how much hype they can all get out of this, but Qwen really is the little engine that could. Great to see open weights (if not open source) driving the whole ecosystem like this though.
Super cool finding!
And also, it is nowhere as good as Qwen
As of this afternoon, we also have Gemma4-based variants of strands-decider at 2B, 4B, 12B, and 26B: https://huggingface.co/StrandsAgents
I don't know if Fabio (who's leading work on this model line) would agree, but if I was going to start over I'd probably pick Gemma4 as the base rather than Qwen3.5. But it's not surprising to see a lot of Qwen competition, and I totally agree it's nice to see the work open.
The insistence of naming it open weights as opposed to open source is getting ridiculous, and it's both irrelevant (i.e. no one cares in practice) and factually incorrect.
Weights are source in language models. Apache defines source as ""Source" form shall mean the preferred form for making modifications". That is precisely what's happening here. Everyone is using the preferred form for making modifications to these models (including the model creators themselves). A model is "created" at init time, and then "trained" by modifying the weights.
All these models are open source. What's not open sourced (with qwen et all) is the training code. So open source model, no training code. And that's ok. There are labs that release those as well. Apertus and Olmo series come with open source models, open source training and open datasets. Nemotron comes with open source models, open source training and some open datasets, while others are not published. And that's ok too.
The fact that you see all these models being modified (from AR completion models to "decision models") and re-released should be all the proof you need. That's what a license offers you. The right to inspect, run, modify and re-release a model. A license cannot (and never did) give you any other rights. OpEnWeIgHtS is silly.
Weights are source in the same way as any x86 binary is source.
You easily modify a x86 binary and change behaviour or examine the machine code instructions. You probably are not aware how easy it is to change the behaviour of a binary executable.
Also if compiling literary costs millions of dollars, I would prefer the precompiled one.
I'm happy to have and be able to serve these models and see the ecosystem thrive. And lots of open innovation is outside of weights anyway as DeepSeek repeatedly shows.
Open source has historically meant you could download the source and build your own binary. The appropriate analogue here, IMHO, to the build->binary process is training->weights.
Dictionary.com defines source as:
> any thing or place from which something comes, arises, or is obtained; origin.
Supports GPU, NPU and CPU.
https://blogs.windows.com/windowsexperience/2026/10/07/build... https://github.com/microsoft/windowsML
Qwen tunes are nice and all but either price or performance makes each example I have tested not viable unless you are able to cheaply self-host and fine-tune further.
Using a decoder model for this comes down to picking the most probable class based on the probability what the next token assigned to a class is. They are all probabilities but semantically mean totally different things. Am I right?
“To build Microsoft-Decision-1, we post trained Qwen3.5-9B”
Using the output of a “decision” model without insight into the reasoning for a given decision seems very trusting.
Also, section 2.3: https://typesafe.ai/legal/mca
It does say "develop (or to facilitate the development of) a similar or competing product or service", but I think it would be a long stretch to say that's the case if they would just publish benchmarks. Microsoft legal department might disagree.
I guess the fear of Chinese models is finally subsiding.
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
The main market for Jev is going to be business, who may not want to deal with a low cost open weights server like Fireworks AI and are more likely to go with a provider like AWS where pricing is higher.
Question: “Which team should handle this message?”
Result: “Tech Support (85%)”
Yep, sounds about right.
I'm still waiting for the Microsoft Vacuum Cleaner. /s
https://www.whitehouse.gov/presidential-actions/2025/01/endi...
No need to use AI, this was done by ctrl-f on dangerous terms such as "disabilities", "trauma" and "female".
(No, I am not making this up)