| ▲ | pixelesque 2 hours ago | |||||||
What quant are you using, and what tps are you getting with K3? | ||||||||
| ▲ | redrove 2 hours ago | parent [-] | |||||||
The FP8 version from DeepSeek themselves [0], around 1800 tps prefill and 45 tokens per second decode. I’ve been running a custom VLLM image with b12x as well as nvfp4_ds_mla. I would say it’s quite fantastic in day to day, I use it mostly in Hermes and sometimes for coding. I have qwen 3.6 27b on an rtx 6000 pro as well so I use that as a workhorse in pi with DS as a reviewer/planner. [0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 Edit: I think you may have misread my post. k3s is NOT kimi k3, and I did mention I was running deepseek. | ||||||||
| ||||||||