| ▲ | walrus01 2 days ago |
| One could also run it locally on a used dual xeon (or amd-equivalent) server with 512GB RAM, albeit slower, if you have a useful workflow for it that's like "take this day's efforts and run it through various analysis agents", combined with giving it one-shot tasks/modules to build overnight. You would want a place like a garage or basement to put the server because it'll be loud. |
|
| ▲ | peri-cl 2 days ago | parent | next [-] |
| > "dual xeon" Does inference make full use of the memory bandwidth in a NUMA system? |
| |
| ▲ | walrus01 2 days ago | parent [-] | | Memory throughput performance is better than many equivalently priced (talking about the cheaper end of the used/refurb market here) single socket systems because each CPU is talking to its own banks of RAM. Using an example of llama-server on a 256GB (each CPU is attached to 128GB) system running qwen 3.8-flash-next Q8, which takes up about 188GB with the model and context, it's not that the inference engine 'knows' about NUMA, but if you use --load-mode mlock to load the whole thing into memory at once at the time llama-server launches, the kernel will distribute it relatively evenly across the memory connected to both CPUs. |
|
|
| ▲ | dataplumb3r 2 days ago | parent | prev [-] |
| You'd also likely spend far more in electricity than the API cost of processing the prompt(s) |
| |
| ▲ | walrus01 2 days ago | parent [-] | | yes, though for some uses, not sending data anywhere to third parties has its own value which is harder to measure. |
|