Remix.run Logo
beering 8 hours ago

That is addressed in the article.

floatrock 8 hours ago | parent [-]

OpenAI's position:

> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).

biophysboy 8 hours ago | parent | next [-]

Why is it unlikely?

karmasimida 41 minutes ago | parent | next [-]

Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.

tristanj 3 hours ago | parent | prev | next [-]

Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.

It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.

We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.

biophysboy 2 hours ago | parent [-]

I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math

enraged_camel 4 hours ago | parent | prev [-]

Because OpenAI says so, obviously!

cute_boi 8 hours ago | parent | prev [-]

I thought openai don't use any user data if we opt out of training and via api?

andrewguenther 8 hours ago | parent [-]

That is correct. It is possible they didn't opt out and given the timeline and anonymization of data unclear whether a particular conversation would have made it into the training set if they hadn't.

luke5441 6 hours ago | parent [-]

Easy to ask for the account used to see if its usage went into training data. Also easy to say "Knowledge cut-off of the used model was date X".

That they don't is telling.

ImaCake 3 hours ago | parent [-]

Maybe they didn't ask or the anthropic people said no? Plenty of other explanations here.