| ▲ | sleepyeldrazi 28 minutes ago | |
Speed is (for the most part) active-parameter based, so a 30B-A3B model is roughly 10x the speed of a dense 30B (realistically closer to 8x) in the case when both fit. That's the proposition of MoE and why everyone is trying to make massive models with very few active params, so that they are still fast while having access to a lot of knowledge (at the cost of reasoning, as reasoning ability 'for the most part' comes from active params). You can test this by running this nemo or 35B on the 3090. I have and its very fun (but sadly a worse model than 27B, so I usually keep 27B on my 3090) | ||
| ▲ | XCSme 17 minutes ago | parent [-] | |
I remember running both qwen 30b-a3b and 27b on my 3090, and on the initial test, the 27b was only like 2x slower. | ||