Remix.run Logo
smcleod an hour ago

The smarter 27B is so fast with MTP I've found I really don't need the 35B-A3B. You get around 70tk/s on a M5 Max lowering to around 40tk/s at higher context sizes.

seanmcdirmid an hour ago | parent | next [-]

I've benched 3.8 27B being significantly slower and less quality than 3.6 35B-A4B (both 4-bit quant, MTP, both using turboquant 4-bit served by oMLX), to the point that I'm not even using it right now (on an M3 Max). What's your use case and what did you observe? I might be missing something.

smcleod an hour ago | parent [-]

I believe you mean 35B-A3B, there was no such thing as A4B. I use 27B and other models for software development, and quite a few research or similar agents. I cannot imagine a world where the old 35B-A3B model is smarter / more capable than 3.8 27B - the difference is night and day for coding at least. Where 35B-A3B was fast and felt like a Haiku model, 27B feels like a strong Sonnet when given the right tools. I don't use turbo quant so can't comment on that, but with the A3B model you're using you probably won't get much from using MTP with small MoE models like that.

xscott 14 minutes ago | parent [-]

People over-quantize things, muck with the temperature and other settings based on superstitions or results from models they think are similar. There's lots of ways to make 3.8 27B dumber.

quinncom an hour ago | parent | prev | next [-]

I only get ~4 tok/sec on a M1 Pro with MTP.

Muromec an hour ago | parent | prev [-]

Does 27b mean it fits one 32GB GPU?

smcleod an hour ago | parent [-]

It doesn't mean that, but yes it would (with a 5 bit quant).