| ▲ | joey64 19 hours ago | |||||||
So, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8? | ||||||||
| ▲ | vanillax 19 hours ago | parent | next [-] | |||||||
you cant. The best you can do is Qwen 3.6 27b with a 24gig ( or cumaltive gpus ) to get to 24gb vram. ala 3090, mac with 36gb ram, amd cards, halo strix amd, dgx spark etc. Lots of youtube videos out there. | ||||||||
| ▲ | svachalek 16 hours ago | parent | prev | next [-] | |||||||
I haven't seen either of these running outside their creator's services yet, but typically you can watch services like openrouter or nano-gpt for it to show up at a (usually small) discount. | ||||||||
| ▲ | Alpha3031 19 hours ago | parent | prev | next [-] | |||||||
Well, if you're happy with around (as in within an order of magnitude or two of) 0.1 tokens per second... I believe that's around what people are getting when loading MoE weights from NVMe. | ||||||||
| ||||||||
| ▲ | drnick1 17 hours ago | parent | prev [-] | |||||||
There will be smaller versions in the 10-30B parameters range that can run on consumer GPUs. | ||||||||