| ▲ | scosman 6 hours ago |
| So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon.... Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB. |
|
| ▲ | sourcecodeplz 15 minutes ago | parent | next [-] |
| i am scared for the PRO model maybe it is Fable level |
|
| ▲ | luckydata 4 hours ago | parent | prev | next [-] |
| I would like to see your "home" |
| |
| ▲ | cmrdporcupine 4 hours ago | parent [-] | | Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest. But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons. | | |
| ▲ | ycui7 4 hours ago | parent [-] | | the rational in one’s mind is similar to buying expensive supercar but no driving it daily. owning a few GPUs is a lot cheaper than supercars. | | |
| ▲ | cmrdporcupine 3 hours ago | parent [-] | | I dunno. I bought the Spark in January and it has led indirectly to paid work. I don't use it for local inference so much. I use it to learn. I also use it as my daily driving Aarch64 development system. Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it... | | |
|
|
|
|
| ▲ | segmondy 5 hours ago | parent | prev [-] |
| Q8 is ~ 151gb |
| |