Remix.run Logo
highfrequency 6 hours ago

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models

This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.

But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.

MichaelDickens 5 hours ago | parent [-]

> But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game.

Just because something is legal and permitted by terms of service doesn't mean it's morally right.

Jtariiiii 4 hours ago | parent [-]

>Just because something is legal and permitted by terms of service doesn't mean it's morally right.

What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data? Are they supposed to manually review all their data to make sure competing mathematicians didn't accidentally leave the "submit prompts" toggle on?

Or were they supposed to not try to solve Navier-Stokes, or were they supposed to just not tell anyone that they had solved it?

plaidfuji 4 hours ago | parent | next [-]

To me it’s morally ambiguous… if you hand parts of your thinking over to a tool like this (knowing full well the terms of service), of course the tool makers will want to claim some credit, and they do deserve it. But the bigger question to me is the scientific one: did their new model arrive at this result because it had closely-related training data from a human, or did it extrapolate to this line of thought on its own? The answer says a lot about how valid their claims of “AGI” are vs. a very fortuitously cherry-picked example.

It would actually be a really interesting study, if they would ever be willing to be transparent about this, how the result differs with and without his conversations in the training set. How quickly it arrives at the result, whether it takes the same approach, etc.

vemacs 4 hours ago | parent | prev | next [-]

> Are they supposed to manually review all their data t

Yes. They should determine if training data included this teams data. Consider the money they spent, the press release and the purpose of their publication.

Since they failed to answer this question they shouldn't have published.

tristanj 3 hours ago | parent [-]

[dead]

nozzlegear 4 hours ago | parent | prev [-]

> What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data?

Personally, I would expect them to have a little class, to KYC, and to manually turn off training for known competitors using their service so as to avoid any unforced goofs like this.