| ▲ | falsaberN1 2 days ago | |
With llama.cpp (CUDA) and a 5060ti (16GB) I get 60t/s with 128K token space. Odd you got 7t/s, did you verify all the model was loaded in VRAM? (--gpu-layers all) | ||
| ▲ | Pragmata 2 days ago | parent [-] | |
it was but i max out on context so it doesn't all fit with kv cache etc... | ||