| ▲ | zozbot234 2 hours ago | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
You can definitely offload n-gram embeddings to storage; they're very sparsely used (only a few KB fetched per token) so this is quite effective. Loading to DRAM only becomes necessary if they are a bottleneck to overall performance (which might happen if you're doing very wide batches and everything else uses super fast VRAM/HBM). | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | verdverm 2 hours ago | parent [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I was looking at the qwen-next-flash, and the weights would fill my OEM Spark on their own, before the n-gram. I'm unclear if offloading to disk can work here, is that what you are implying is possible?! | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||