| ▲ | chrishynes 6 hours ago |
| Why can't this scale to run much larger models on CPU backed by flash with good access patterns? |
|
| ▲ | AussieWog93 4 hours ago | parent | next [-] |
| Someone did this exact thing recently, but running GLM-5.2 with something like 16GB of DRAM, a standard desktop CPU and nVME SSD. I think they got something like 10 _seconds per token_ (not tokens per second). EDIT: it was 25GB of ram and up to 20 seconds per token!
https://github.com/JustVugg/colibri |
| |
| ▲ | 3eb7988a1663 3 hours ago | parent [-] | | That's incredible. Sure, not practical for most applications, but if you really want a local top tier model, you can run it on anything as long as you are patient. As someone with a healthy amount of RAM, but just a 16GB GPU, I am wondering what kind of work I could queue up for overnight runs. I thought the best models were fully out of reach, but the 128GB CPU only test had a 1.8 tokens/second. While not speedy, you could probably do something with that given extensive coffee breaks. This speed simulator[0] demos what it looks like. [0] https://shir-man.com/tokens-per-second/?speed=1.8 |
|
|
| ▲ | Rohansi 4 hours ago | parent | prev [-] |
| My guess is because the ESP32's flash is only ~1/4 the bandwidth of the internal SRAM. If you do this on a more powerful system not only is the gap much wider but you also have much more compute you need to keep fed with bandwidth to be efficient. |
| |
| ▲ | monocasa 2 hours ago | parent [-] | | It's also mapped into the address space so there's very little extra latency in grabbing the embedding as opposed to something like nvme that will have to setup a command list, submit it to the drive's microcontroller, wait for the op to be processed, etc. |
|