| ▲ | kasperni 12 hours ago | |
Maybe a bit of context for this post? Some people have a life outside of AI. | ||
| ▲ | toshinoriyagi 8 hours ago | parent | next [-] | |
This model is a preview of Qwen's upcoming Qwen4 architecture. It is a 125B-A6B MoE model, meaning it has 125B total parameters with 6B active at a time, but it also has a 51B parameter engram with it. The engram is basically a lookup table for tokens to my understanding. It allows the model to have access to a much larger amount of info if utilized well. They said the model is intentionally under-trained since it is mainly for R&D purposes of proving the new architecture. Many people are excited for models in this range as they are a step above the common ~27B models, while not requiring exorbitant sums of money to run like much larger models. | ||
| ▲ | glimshe 12 hours ago | parent | prev | next [-] | |
I got downvoted yesterday for complaining about the name/brand confusion from all these Chinese models with similar names all claiming they are the best. While I'm an AI enthusiast, it's being hard to keep track. | ||
| ▲ | poincareball 12 hours ago | parent | prev [-] | |
[dead] | ||