Remix.run Logo
toshinoriyagi 8 hours ago

This model is a preview of Qwen's upcoming Qwen4 architecture. It is a 125B-A6B MoE model, meaning it has 125B total parameters with 6B active at a time, but it also has a 51B parameter engram with it. The engram is basically a lookup table for tokens to my understanding. It allows the model to have access to a much larger amount of info if utilized well.

They said the model is intentionally under-trained since it is mainly for R&D purposes of proving the new architecture. Many people are excited for models in this range as they are a step above the common ~27B models, while not requiring exorbitant sums of money to run like much larger models.