| ▲ | jodacola 18 hours ago | |||||||||||||||||||||||||||||||
Setting aside annoyance at the downtime, I'm really curious about the reasons for the failures, because I have to imagine there are some novel failure modes when serving these giant models that I haven't experienced with the kind of work I've done. Anyone out there working in this space who can elucidate us on interesting failure scenarios unique to the space? | ||||||||||||||||||||||||||||||||
| ▲ | martinald 17 hours ago | parent [-] | |||||||||||||||||||||||||||||||
It's (mostly?) compute shortages. Right now it seems there is an issue in the SpaceX datacentres, so they will have less compute than normal. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||