| ▲ | NitpickLawyer 11 hours ago | |||||||
There's also the 2x spark way, which should be ~8k eur? Someone down the thread reported ~60tps for 2x sparks. That's totally usable for local inference. You can also do 2x 6kPRO in a workstation, for ~20k. | ||||||||
| ▲ | matrik 10 hours ago | parent | next [-] | |||||||
For the same performance, one could even go about 50% cheaper with 16 channel ddr4 + a rtx3090 for prompt processing. But still, even for mid level projects API is orders of magnitude cheaper, since you don't need to set it up and maintain it. | ||||||||
| ||||||||
| ▲ | spwa4 10 hours ago | parent | prev [-] | |||||||
Currently the 3bit (and 2 bit) quant on DGX spark (on one of them) and the M5 Max should just start. Right now. (I'm hoping to get an M5 Max delivered on monday, let's see if it happens this time. It's 2+ months since I ordered now) The 4 bit quant technically fits (there's a 127 GB version) but ... obviously that's not going to work. It is so close though, surely someone will a way to do it. | ||||||||