Remix.run Logo
OutlawHusbando 2 days ago

People in my server are running it on Strix Halo 128GB using RoCmFP4 and reporting 35tok/s, without much optimization, with proper MTP, better kernel, expecting about 50-60tok/s.

walrus01 2 days ago | parent | next [-]

The largest unsloth published gguf also fits and runs just fine on a cpu-only machine with 128GB RAM, using llama-server PR 27742

https://github.com/ggml-org/llama.cpp/pull/27742

manmal 2 days ago | parent | prev [-]

How do they like it, compared with 3.8 and DS4?