| ▲ | kevin42 20 minutes ago | |
What hardware do you run? I have a first-gen mac studio, and I just run cmake and build with no special options. Same thing with llama-server, I just specify the model and use the built-in web UI. For reference, I get ~26 tok/sec with the new Muse 30B model. | ||
| ▲ | dofm 15 minutes ago | parent [-] | |
An M1 Max MBP manages roughly 10 tok/sec without the Dflash speculative draft support so that tracks; the M1 Max apparently has trouble actually saturating its memory bandwidth. | ||