Remix.run Logo
▲ nejch 2 hours ago

This is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider, and a plethora of others released in the last couple of weeks.

Kind of funny how much hype they can all get out of this, but Qwen really is the little engine that could. Great to see open weights (if not open source) driving the whole ecosystem like this though.

▲giancarlostoro 18 minutes ago | parent | next [-]

Surprised they aren't doing these sorts of one-off models with Microsoft Phi, which is intentionally smaller, but there's no reason Microsoft couldn't try to make a slightly larger Phi model with more capabilities...

▲manmal an hour ago | parent | prev | next [-]

My biggest learning after some experiments - a BF16 (unquantized) Qwen beats a Q8 of double its size for decisions. I guess that’s the reason Kev switched to 4B BF16, from the original 8B version. Isn’t it interesting that quantization seems to mess with decision accuracy?

▲girvo an hour ago | parent [-]

That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though

Super cool finding!

▲ByteAtATime 5 minutes ago | parent [-]

Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do

▲NitpickLawyer an hour ago | parent | prev [-]

> Great to see open weights (if not open source)

The insistence of naming it open weights as opposed to open source is getting ridiculous, and it's both irrelevant (i.e. no one cares in practice) and factually incorrect.

Weights are source in language models. Apache defines source as ""Source" form shall mean the preferred form for making modifications". That is precisely what's happening here. Everyone is using the preferred form for making modifications to these models (including the model creators themselves). A model is "created" at init time, and then "trained" by modifying the weights.

All these models are open source. What's not open sourced (with qwen et all) is the training code. So open source model, no training code. And that's ok. There are labs that release those as well. Apertus and Olmo series come with open source models, open source training and open datasets. Nemotron comes with open source models, open source training and some open datasets, while others are not published. And that's ok too.

The fact that you see all these models being modified (from AR completion models to "decision models") and re-released should be all the proof you need. That's what a license offers you. The right to inspect, run, modify and re-release a model. A license cannot (and never did) give you any other rights. OpEnWeIgHtS is silly.

▲teruakohatu 43 minutes ago | parent | next [-]

> Weights are source in language models.

Weights are source in the same way as any x86 binary is source.

You easily modify a x86 binary and change behaviour or examine the machine code instructions. You probably are not aware how easy it is to change the behaviour of a binary executable.

▲nejch an hour ago | parent | prev | next [-]

I mean I don't have a strong opinion but if I used the phrase open source models there'd be 5 comments going in the other direction.

I'm happy to have and be able to serve these models and see the ecosystem thrive. And lots of open innovation is outside of weights anyway as DeepSeek repeatedly shows.

▲Rohansi an hour ago | parent | prev [-]

> OpEnWeIgHtS is silly.

Dictionary.com defines source as:

> any thing or place from which something comes, arises, or is obtained; origin.