| ▲ | amelius 7 hours ago | ||||||||||||||||||||||
> Without training cost you can infer only the marginal cost of serving this kind of models. Which is by far the most interesting number of the two. > Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?) If you get close in output quality, then does that matter? | |||||||||||||||||||||||
| ▲ | embedding-shape 7 hours ago | parent | next [-] | ||||||||||||||||||||||
> If you get close in output quality, then does that matter? When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter? | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | vb-8448 5 hours ago | parent | prev [-] | ||||||||||||||||||||||
> Which is by far the most interesting number of the two. Only if you don't have to continuously train new models, and you are not at a runway risk. | |||||||||||||||||||||||