| ▲ | kadoban 4 hours ago | |||||||
You can run the ~4 bit quant(s) on 24gb, if you're not _too_ picky on context size. This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks. | ||||||||
| ▲ | spijdar 4 hours ago | parent [-] | |||||||
I run Unsloth's UD-Q4_K_S on 20 GB of VRAM (RX 7900 XT) and I get ~90k tokens of context without quantizing KV cache. With 8-bit quantization, I get about a 134k token context window. That's with only one slot, but for me, it works pretty darn well, with 20-35 tok/s depending on how full that window is. | ||||||||
| ||||||||