| ▲ | andy_ppp 4 hours ago | |||||||||||||
Yes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes. | ||||||||||||||
| ▲ | structural 4 hours ago | parent | next [-] | |||||||||||||
Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts. Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower. It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it. | ||||||||||||||
| ||||||||||||||
| ▲ | Caracas288 4 hours ago | parent | prev | next [-] | |||||||||||||
Man wouldn’t it be cool to be able to slot a massive ROM AI chip into the external AI drive of the pc… | ||||||||||||||
| ||||||||||||||
| ▲ | porphyra 4 hours ago | parent | prev | next [-] | |||||||||||||
Also right now Sol 5.6 Max is super slow but if it were way faster on a chip (like Taalas' Llama 8b demo) then it would be an extreme value multiplier. But the model is so large that "baking it onto a chip" doesn't seem straightforward. | ||||||||||||||
| ▲ | redox99 4 hours ago | parent | prev [-] | |||||||||||||
That'd be ungodly expensive. | ||||||||||||||