| ▲ | smcleod an hour ago |
| The smarter 27B is so fast with MTP I've found I really don't need the 35B-A3B. You get around 70tk/s on a M5 Max lowering to around 40tk/s at higher context sizes. |
|
| ▲ | seanmcdirmid an hour ago | parent | next [-] |
| I've benched 3.8 27B being significantly slower and less quality than 3.6 35B-A4B (both 4-bit quant, MTP, both using turboquant 4-bit served by oMLX), to the point that I'm not even using it right now (on an M3 Max). What's your use case and what did you observe? I might be missing something. |
| |
| ▲ | smcleod an hour ago | parent [-] | | I believe you mean 35B-A3B, there was no such thing as A4B. I use 27B and other models for software development, and quite a few research or similar agents. I cannot imagine a world where the old 35B-A3B model is smarter / more capable than 3.8 27B - the difference is night and day for coding at least. Where 35B-A3B was fast and felt like a Haiku model, 27B feels like a strong Sonnet when given the right tools. I don't use turbo quant so can't comment on that, but with the A3B model you're using you probably won't get much from using MTP with small MoE models like that. | | |
| ▲ | xscott 14 minutes ago | parent [-] | | People over-quantize things, muck with the temperature and other settings based on superstitions or results from models they think are similar. There's lots of ways to make 3.8 27B dumber. |
|
|
|
| ▲ | quinncom an hour ago | parent | prev | next [-] |
| I only get ~4 tok/sec on a M1 Pro with MTP. |
|
| ▲ | Muromec an hour ago | parent | prev [-] |
| Does 27b mean it fits one 32GB GPU? |
| |