Remix.run Logo
falsaberN1 2 days ago

With llama.cpp (CUDA) and a 5060ti (16GB) I get 60t/s with 128K token space. Odd you got 7t/s, did you verify all the model was loaded in VRAM? (--gpu-layers all)

Pragmata 2 days ago | parent [-]

it was but i max out on context so it doesn't all fit with kv cache etc...