Remix.run Logo
ollamer a day ago

https://github.com/ollama/ollama/issues/11772

A year and still no implementation for such a basic need as offloading MoE layers onto the CPU selectively. On llama.cpp I can get models like Qwen 35BA3B running partially on gpu/cpu with 40t/s on a laptop thanks to --n-cpu-moe but on this VC funded joke it would be simply unusable. I can't quite understand how you make a wrapper so much worse than the code you're ripping out.

>This funding is fuel for what’s ahead. Ollama sits front and center in the open model ecosystem

No.