Remix.run Logo
arcanemachiner a day ago

> I dont know how they make money here

I assume it's a subsidy to get more training data.

EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?

tedsanders a day ago | parent [-]

By default, OpenAI does not train on API data. I promise you that Luna's low pricing is not a subsidy to get more training data. We've been lowering prices for years.

(I work at OpenAI.)

arcanemachiner a day ago | parent [-]

Wait, so you guys don't anonymize the user data, then train on it after it's been sanitized? I thought this was done to some degree or another.

So what is the value prop then? Just basic supply and demand?

FWIW I have definitely noticed OpenAI's emphasis on efficiency and value in the last year, so that part isn't new to me... I just thought there was more to it then that.

tedsanders a day ago | parent [-]

API: By default, no training (opt in).

ChatGPT enterprise: By default, no training (opt in).

ChatGPT personal: By default, training (opt out).

shostack 3 hours ago | parent [-]

Ted can you confirm your choice of words here to be precise for an audience who is familiar with the nuances, when you say "no training" or "training (opt out)" for personal... Is that inclusive of "sanitized" (or pseudonymized) data?

Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."

People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."

But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.