| ▲ | itake 3 hours ago | |||||||
does your comment depend on the OS? I thought MLX has better performance on MacOS than llama.cpp | ||||||||
| ▲ | quantumleaper 3 hours ago | parent [-] | |||||||
The gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF ecosystem provides a better selection of quants, in my experience Unsloth ones are excellent. | ||||||||
| ||||||||