| ▲ | spider-mario 2 hours ago | |||||||
> Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it. Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB. | ||||||||
| ▲ | SwellJoe 2 hours ago | parent | next [-] | |||||||
Even at 30 t/s, 3.8 thinks so long, even on medium, it still takes 3x or more longer than any cloud model, in my testing. | ||||||||
| ||||||||
| ▲ | Xeoncross 2 hours ago | parent | prev | next [-] | |||||||
Nice, which model quantization is this? Is it on huggingface? | ||||||||
| ▲ | bellowsgulch an hour ago | parent | prev [-] | |||||||
Thanks, man! I’ll go use that now that I know. llama-server the last time I used it for inference with this model wasn’t able to produce work fast enough to reach those numbers. | ||||||||