Remix.run Logo
docheinestages a day ago

Qwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.

embedding-shape a day ago | parent | next [-]

> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable

Yeah, that'd be neat, but that's not what this announcement is about at all:

> With a massive 2.4T parameters

docheinestages a day ago | parent | next [-]

True. It was more of an open letter, with hopes that the Qwen team sees the comments in this thread.

cyanydeez a day ago | parent | prev [-]

dont we all deem the ability to improve large models as the defacto capability to produce small ones?

embedding-shape a day ago | parent [-]

I don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.

cyanydeez a day ago | parent [-]

but that's denying the singularity boostrap theory. Which I don't agree with, but if you can't harness a large model to make a small model, then we're going to have problems brining about the singularity.

If you do think there's some magical singularity, how do you comport?

drob518 a day ago | parent | prev | next [-]

I’d like a “Bonsai 2.8T.” That is, something that is near the Fable/Sol/K3 class, but capable of running locally on consumer hardware.

kingo55 20 hours ago | parent [-]

At 1.5 bits per weight it'll still be over 500gb - that's still not running on consumer hardware.

Best case they release smaller models. 120b class of qwen 3.8 would be incredible - it fits on device for those serious about AI, but without millions of dollars in hardware for terabytes of VRAM

scotty79 a day ago | parent | prev [-]

I can't really blame them that the biggest labs focused on trainig and realeasing huge models.

The niche for small models should be filled with medium sized labs doing distillations of the huge ones into consumer grade hardware runnable models and LORAs for the huge ones.

docheinestages a day ago | parent [-]

I think AI will evolve the same way computers did. We're somewhere in the 80s-90s timeline of the evolution. My prediction is that on-device models will have excellent tool-calling, reasoning, and general skills, but the domain-specific knowledge will be retrieved on-demand from vendors like Google. Rather than downloading models, each device will have a hardware component with weights baked into silicon for maximum efficiency.

gunalx 20 hours ago | parent [-]

Weigths directly in silicon is a bad idea with the way the space is pacing. Just look at chatjimmy.ai it is fast, but on the once good llama3.1-8b but now pretty useless.