Remix.run Logo
spider-mario 2 hours ago

> Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it.

Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB.

SwellJoe 2 hours ago | parent | next [-]

Even at 30 t/s, 3.8 thinks so long, even on medium, it still takes 3x or more longer than any cloud model, in my testing.

lowbloodsugar 14 minutes ago | parent [-]

I've got an M1 Max 64GB too. It's just not an LLM-class workstation. Give it a year and buy an M7 and you'll be laughing. Right now is a really bad time to invest in anything - using the cloud is the cheapest option, especially for open weight models.

Xeoncross 2 hours ago | parent | prev | next [-]

Nice, which model quantization is this? Is it on huggingface?

bellowsgulch an hour ago | parent | prev [-]

Thanks, man! I’ll go use that now that I know. llama-server the last time I used it for inference with this model wasn’t able to produce work fast enough to reach those numbers.