Remix.run Logo
AussieWog93 4 hours ago

Someone did this exact thing recently, but running GLM-5.2 with something like 16GB of DRAM, a standard desktop CPU and nVME SSD.

I think they got something like 10 _seconds per token_ (not tokens per second).

EDIT: it was 25GB of ram and up to 20 seconds per token! https://github.com/JustVugg/colibri

3eb7988a1663 3 hours ago | parent [-]

That's incredible. Sure, not practical for most applications, but if you really want a local top tier model, you can run it on anything as long as you are patient.

As someone with a healthy amount of RAM, but just a 16GB GPU, I am wondering what kind of work I could queue up for overnight runs. I thought the best models were fully out of reach, but the 128GB CPU only test had a 1.8 tokens/second. While not speedy, you could probably do something with that given extensive coffee breaks. This speed simulator[0] demos what it looks like.

[0] https://shir-man.com/tokens-per-second/?speed=1.8