Remix.run Logo
purplemoonx 26 minutes ago

Still no easy way to get up and running performantly with Llama.cpp

Nobody tells you how, they just act like you're an idiot. So I always say fuck this and install Ollama.

Then everyone goes BLASPHEMY "just use llama.cpp"

Yeah I did, and it's slow as hell. It doesn't work well. Idk why.

"You aren't doing it right"

Okay tell me how

"NO"

-----

For this reason, Ollama is the superior solution. I know, downvote, everyone hates Ollama here but until llama.cpp gets their shit together on developer experience it doesn't exist as far as I'm concerned

kevin42 19 minutes ago | parent | next [-]

What hardware do you run? I have a first-gen mac studio, and I just run cmake and build with no special options. Same thing with llama-server, I just specify the model and use the built-in web UI.

For reference, I get ~26 tok/sec with the new Muse 30B model.

dofm 14 minutes ago | parent [-]

An M1 Max MBP manages roughly 10 tok/sec without the Dflash speculative draft support so that tracks; the M1 Max apparently has trouble actually saturating its memory bandwidth.

drittich 20 minutes ago | parent | prev | next [-]

There are certainly challenges. When setting up a new model, I get AI to walk me through the commands using llama-benchmark that determine the best parameters for my particular configuration and needs. Once you've got that it's pretty easy to port those parameters to llama-server. It takes me about an hour to run through this process. It would be great if there was a registry of hardware, models, configuration parameters, and resulting tokens per second. Maybe one day we'll get there.

kevin42 16 minutes ago | parent | next [-]

What kind of parameters do you end up changing, and how much difference does it make. Perhaps I am missing something and get more tok/sec, but I usually just do a git pull, then rebuild the latest whenever I get a new model.

In the past, I had to play with chat templates for some models to work with agents for tool calling. But I've never had to do anything other than specify the model, and tweaking the context size in some cases.

purplemoonx 15 minutes ago | parent | prev [-]

> It takes me about an hour to run through this process

Yeah not doing that

unglaublich 25 minutes ago | parent | prev | next [-]

I think people generally throw Claude or Codex at the configuration challenge, so they don't know either.

purplemoonx 19 minutes ago | parent [-]

Maybe the llama.cpp dev loved webpack as a child, or just loves making the most simple thing complicated as hell for no reason lol

NamlchakKhandro 5 minutes ago | parent | prev [-]

No