| ▲ | AbsurdCensor an hour ago | |
I have had an impossible time getting 120B or better models running on Strix Halo (especially under Windows) with any large context windows. And 30-40 tokens/second is fine, but not the fastest. For the most part lately I have been sticking with Qwen 3.8 27b and that thing will easily suck up 64gb of ram. Add in docker with some additional programs running and it's really easy to eat up 128gb of ram. | ||
| ▲ | SwellJoe an hour ago | parent [-] | |
I found a couple of different 4-bit quantizations of Laguna S 2.1 that run pretty well with pretty big context (also quantized, to 8 bits, I think). Unfortunately, Laguna isn't better than Qwen 3.8 27B, which I'm able to run at roughly the same speed on my desktop machine, so I don't use Laguna or the Strix Halo very much, lately. (It's also too hot for me to be running heaters for inference. It's been ~110F most days for the past few weeks.) | ||