Remix.run Logo
lnenad 2 days ago

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

springtimesun 2 days ago | parent | next [-]

I’ll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that’s at Q8 and a 4090 doing pre fill so it could be pushed up.

The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will turn down speed to try to stay below 80.

Happy to send my llama.cpp config settings if you want it.

lnenad 2 days ago | parent | next [-]

I am getting 10t/s on unsloth's Q3kxl with 2x3090s@250w. It's enough for me for now. I will probably upgrade the GPUs down the line. DDR5 would have made the price of the machine double and I just wasn't prepared to pay that much.

Temp wise, no throttling, surprisingly cool.

pdntspa 2 days ago | parent | prev | next [-]

I was running one of the older llamas (3.1 I think?) at slow-ish (10-20 tok/sec at Q4?) but OK speeds on 12 year old DDR3 ECC Xeon machine

springtimesun 2 days ago | parent [-]

I find 10 to be very usable. It’s not (that) interactive but it chews through tasks. I let Kimi churn away at 4 overnight and it gives good results that are ready for me in the morning.

pixl97 2 days ago | parent | prev [-]

Typically computers with these larger memory amounts have fans that scream like a banshee trying to move impossible amounts of air over the memory and CPU. Getting something both cool and quite can be a bit difficult.

springtimesun a day ago | parent | next [-]

Yes, I thought when I was starting that 1u and 2u form factors were to save space. Maybe they are, but they also have the advantage of moving air front to back very effectively through and over the components. Though there still must be some need because I see even those boxes have optional manufacturer built memory shrouds to try to force airflow between the DIMMs.

I had some 120x38mm fans from another server box that I pulled out because they were too loud and I didn't need the static pressure they were giving. They went in here. That 13mm (and the extra 1k rpm) moves so much more air.

mafuy a day ago | parent | prev [-]

I built a dual epyc server with 64 cores and 1 TB of DDR4. Draws around 800W or so under load. I used off the shelf liquid cooling. It is audible but not noisy.

The trick is to turn on the cooler's RGB in your 6000€ server to get a free speed boost. I am not liable for sysadmin's heart attack upon reading this.

guybedo 2 days ago | parent | prev | next [-]

I have a dual epyc + 1TB RAM. I could push glm 5.2 to 7 tok/s CPU only.

0x457 2 days ago | parent | prev | next [-]

Depending on which Epyc you got it might be slower than 1/5 of the speed.

lnenad 2 days ago | parent [-]

48c 7643. I'm getting about 10tps @Q3kxl with 2x3090s.

nazgulsenpai 2 days ago | parent | prev | next [-]

Curious about that price, if you don't mind sharing a ballpark

lnenad 2 days ago | parent [-]

About 5k with RAM and GPUs bought used. Eastern Europe.

fsuts 2 days ago | parent | prev | next [-]

It’s not unified ram? I.e VRAM so it will struggle

lnenad 2 days ago | parent [-]

I'm getting about 10tps @Q3kxl with 2x3090s.

jchw 2 days ago | parent | prev [-]

Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.

Philpax 2 days ago | parent | next [-]

The fastest I was able to get my Threadripper 3960X + 2x 3090s + 256GB DDR4-3200 to run a 2-bit quant of GLM-5.2 was 8 TPS. I would expect to be in seconds-per-token territory for a pure-CPU 4-bit quant.

jchw 2 days ago | parent | next [-]

One thing I'd like to try is MoE offloading: I have 2x32 GiB of VRAM and 128 GiB of DDR5 running at 4800 MT/s (only 2 channels though). I've seen people post difficult to believe MoE offloading results albeit a decently long time ago with older models. Maybe there is a quant that would fit with MoE offloading?

That said, I am guessing my problem is not enough RAM - but this poor consumer platform struggles to do memory training with 128 GiB as it is.

Now I surely regret not having gotten Threadripper and 256 GiB of RAM in the before-times.

Philpax 2 days ago | parent [-]

My measurement was with MoE offloading, but there's only so much you can keep on-GPU with a 200GB quant and 48GB of VRAM. It's hard to overcome the CPU/RAM bottleneck.

For what it's worth, all of my hardware was used; I think, all-in, I'm probably at around 3k-4k USD? Not cheap, but also not the worst for something relatively versatile.

jchw 2 days ago | parent [-]

Ah, I see - so MoE offloading is no savior. A shame but no surprise either.

snerbles 2 days ago | parent | prev [-]

With a 4-bit quant of GLM-5.2, I can get about 0.8-1.1 tok/s on an underclocked dual Xeon E5-2698 v4 with 512GiB of DDR4-2400. I think it was specifically a Q4_K_M quant. Of course, the time-to-first-token is absolutely atrocious.

Which is completely insane for a ten year old configuration.

lnenad 2 days ago | parent | prev [-]

What model are you interested in? DS Flash 0731@Q4KXL I'm about 25-30tps. Same as the new Qwen3.8 Flash Next. The new GLM 5.3Q3KXL at 10tps. I've got 2x3090s which I didn't mention in the original message.