Remix.run Logo
nehan 8 hours ago

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."

I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.

pfisch 8 hours ago | parent [-]

If they could then it wouldn't be de-identified data...

dfdydx 6 hours ago | parent | next [-]

Well you could search for elements similar to the proof / problem in the training data, even if it's de-identified, right? OpenAI can probably do better than Ctrl-f "Navier Stokes".

paxys 4 hours ago | parent [-]

There are probably thousands of serious academics taking a crack at millennium problems using AI every day. All those attempts are in the training data. And in fact the two researchers benefited from those attempts as well.

ImPostingOnHN 4 hours ago | parent | prev [-]

The researcher could share a string from one of their conversations and OpenAI can confirm whether it exists in their training data.

Or OpenAI could just look at their code and say what it does (maybe have their AI do it if they're having so much trouble with this?)