| ▲ | herf an hour ago | |||||||
I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti). Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as: set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0 | ||||||||
| ▲ | anerli 31 minutes ago | parent [-] | |||||||
Thanks for reporting the issue. Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? https://github.com/magnitudedev/magnitude/issues As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination. | ||||||||
| ||||||||