| ▲ | Lerc 4 hours ago | |||||||||||||||||||
It would not be a particularly wide ranging hack. There is a strong likihood of the weights being on the actual machine that is running the model, because duh. It is something that I have wondered about with models like chatgot. How many physical locations are needed to serve a model on that scale. Do they have a huge number of sites running inference. My suspicion is that the ability to provide inference to that many people is mutually exclusive to having a security level sufficient to stop a state actor wandering off with a copy of the wrights. At the very least if they want to provide inference affordably. | ||||||||||||||||||||
| ▲ | numpad0 38 minutes ago | parent | next [-] | |||||||||||||||||||
I think it's more likely that the model gets pulled from a SAN into NVIDIA pods, and agents/harnesses would run on a separate random Xeon box or something on the same subnet, using the pod through OAI v1 API. That's easier to maintain overall. | ||||||||||||||||||||
| ▲ | valleyer 4 hours ago | parent | prev [-] | |||||||||||||||||||
"because duh"? OpenAI et al. have extensive infrastructure for running the model on a different machine from the one the harness is being run on, because... that's their main product. I would be absolutely shocked if the model were being run on the same machine as the harness. | ||||||||||||||||||||
| ||||||||||||||||||||