| ▲ | Do you think it happened? Research stolen from their Codex private chats(reddit.com) | ||||||||||||||||||||||
| 77 points by KingOfMyRoom a day ago | 25 comments | |||||||||||||||||||||||
| ▲ | QuadmasterXLII a day ago | parent | next [-] | ||||||||||||||||||||||
If one of the researchers forgot to check "don't train on this" a single time, the pretrain is going to know what they were working on. The question of whether they remembered to check "don't train on this" every single time is super concrete, and embarrassing to _both_ sides if the answer is "no, he didn't read the eula." These models will absolutely remember a brilliant insight that appeared a single time in the pretrain corpus, because if they couldn't-- they would get a slightly worse loss. I don't know what this wishy-washing "well maybe we trained on it but we didn't read it" is supposed to mean. (Example of Fable knowing the content of a deeply unimportant LessWrong post I wrote: https://claude.ai/share/a907b46c-bf7b-4fca-9c71-8582cf8507cc a working Navier Stokes solution would be way more salient) | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | mucha a day ago | parent | prev | next [-] | ||||||||||||||||||||||
It looks like clicking "don't train" might not matter. Per Mark Chen, the Chief Research Officer at OpenAI, "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | Havoc a day ago | parent | prev | next [-] | ||||||||||||||||||||||
If true it could have quite an impact. Everyone I think assumes training in the general “it affects the distribution” sense. But if it fishes precise novel insights out as alleged then it makes these products dramatically less valuable. At least the non zdr ones | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | ProllyInfamous 13 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
Quote from one of the human scientist's papers [†]: >I can say the first LLM-generated-proof Levent sent me was the most horrendous I have ever read; [once] we verified it ... we have been working around the clock to understand this proof and turn it into something readable. How incredible that the smartest humans are showing their limits of comprehension. How human that their smugness still gets in their way, in the presence of actual brilliance (i.e. the LLM writes so correctly that humans cannot understand the depth, only seeing "horrendous" until confirmed). [†] <https://cims.nyu.edu/~tristanb/statement.pdf> [pg1, edited for comprehension] ---- >"The important thing is ... the significance that a mathematician and an LLM model [sic] can now do all this work in a month." [†] >"THE SIGNIFICANCE OF THIS WITH RESPECT TO THE WAY WE TRAIN STUDENTS, ASSIGN CREDIT, REFEREE, AND DECIDE WHAT IS WORTH ONE HUMAN LIFE'S ATTENTION CANNOT BE UNDERSTATED." [†] >"The community needs to have serious and unhurried discussion about where to go from here." [†] | |||||||||||||||||||||||
| ▲ | jrflo a day ago | parent | prev | next [-] | ||||||||||||||||||||||
It's possible that OpenAI was mining prominent researcher's chats for inspiration to tackle these problems, but it's also entirely possible the leak came from his collaborator's end as he works at Anthropic and I'm sure there's plenty of corporate espionage going on between those two. That would also explain why they didn't want to comment on the source of the prompt. | |||||||||||||||||||||||
| ▲ | naniel a day ago | parent | prev | next [-] | ||||||||||||||||||||||
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things. What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit. "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/ | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | willtemperley a day ago | parent | prev | next [-] | ||||||||||||||||||||||
They are currently being sued for trade secret theft so it seems likely. | |||||||||||||||||||||||
| ▲ | quicklywilliam a day ago | parent | prev | next [-] | ||||||||||||||||||||||
The fact that this link is a Reddit thread should probably be taken as evidence that we don't have a clear picture of what happened yet, and speculation is rampant. | |||||||||||||||||||||||
| ▲ | khelavastr a day ago | parent | prev | next [-] | ||||||||||||||||||||||
I'm surprised solutions weren't trained off their Office365 or Google Docs.. | |||||||||||||||||||||||
| ▲ | ChrisArchitect a day ago | parent | prev | next [-] | ||||||||||||||||||||||
[dupe] Discussion: https://news.ycombinator.com/item?id=49605915 And currently: https://news.ycombinator.com/item?id=49613262 | |||||||||||||||||||||||
| ▲ | Ydarbleoj a day ago | parent | prev | next [-] | ||||||||||||||||||||||
Yes. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | slipperybeluga a day ago | parent | prev | next [-] | ||||||||||||||||||||||
[dead] | |||||||||||||||||||||||
| ▲ | josefritzishere a day ago | parent | prev [-] | ||||||||||||||||||||||
100% their data was stolen. You can't enter any private data into AI and expect it to remain private. | |||||||||||||||||||||||
| |||||||||||||||||||||||