Remix.run Logo
kasperni 12 hours ago

Maybe a bit of context for this post? Some people have a life outside of AI.

toshinoriyagi 8 hours ago | parent | next [-]

This model is a preview of Qwen's upcoming Qwen4 architecture. It is a 125B-A6B MoE model, meaning it has 125B total parameters with 6B active at a time, but it also has a 51B parameter engram with it. The engram is basically a lookup table for tokens to my understanding. It allows the model to have access to a much larger amount of info if utilized well.

They said the model is intentionally under-trained since it is mainly for R&D purposes of proving the new architecture. Many people are excited for models in this range as they are a step above the common ~27B models, while not requiring exorbitant sums of money to run like much larger models.

glimshe 12 hours ago | parent | prev | next [-]

I got downvoted yesterday for complaining about the name/brand confusion from all these Chinese models with similar names all claiming they are the best. While I'm an AI enthusiast, it's being hard to keep track.

poincareball 12 hours ago | parent | prev [-]

[dead]