▲ | DiabloD3 5 days ago | |||||||||||||||||||||||||
I don't load all the MoE layers onto my GPU, and I have only about a 15% reduction in token generation speed while maintaining a model 2-3 times larger than VRAM alone. | ||||||||||||||||||||||||||
▲ | EnPissant 4 days ago | parent [-] | |||||||||||||||||||||||||
The slowdown is far more than 15% for token generation. Token generation is mostly bottlenecked by memory bandwidth. Dual channel DDR5-6000 has 96GB/s and A rtx 5090 has 1.8TB/s. See my other comment when I show 5x slowdown in token generation by moving just the experts to the CPU. | ||||||||||||||||||||||||||
|