| ▲ | prometheus1992 11 hours ago | |||||||
It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat | ||||||||
| ▲ | Balooga 10 hours ago | parent | next [-] | |||||||
Now I'm feeling pretty good about getting 10-11 tokens/sec running Qwopus 3.6-35B-A3B Q6_K on an old Mac Pro 2013 (trashcan) with 128GB RAM (DDR3), 12 core Xeon, dual D700s. Arch Linux and llama.cpp. | ||||||||
| ||||||||
| ▲ | trollbridge 8 hours ago | parent | prev | next [-] | |||||||
Anything smaller than a 16” runs into serious thermal problems; even an identically equipped 14” just can’t dissipate enough heat. | ||||||||
| ▲ | carloslfu 10 hours ago | parent | prev [-] | |||||||
interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out! | ||||||||