| ▲ | tedsanders 4 hours ago | |||||||||||||||||||||||||||||||||||||||||||||||||
If they opted out of training, then we definitely did not train on them. If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof. Reasons for my doubt: - I know most of our training recipes - Our model's proof is very different from theirs - The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question) - This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here. Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying. If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful. | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | mucha 4 hours ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||
That's not what your Chief Research Officer, Mark Chen, says on X: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | lambda 4 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||
But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being. Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set. | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Imnimo 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||
I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it also seems at least possible that a key idea or a particular step found its way into training data. It wouldn't mean OpenAI stole their proof - clearly the model developed its own approach. Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this. | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | calf 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||
Your perspective is not helpful until you read and reflect on Tristan's letter stating serious grievances. Your remarks here have minimized his complaints and that is a sign of bias. Do not then pre-accuse HN commenters of being convinced when there reasonable skepticism such biased behavior showing itself in this very thread, saying things that amount to "my tribe/company would never be so egregious and if you think that then it is bad faith". That's the projection. If the word prejudice means anything to you then please do the work of attending to that instead of using the platform to reinforce such biases. If you are not a PhD yourself maybe your are not culturally qualified to assess and expound on the overall situation anyways. | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | lossolo 4 hours ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||||||||||||||
> If they opted out of training, then we definitely did not train on them. Can't you guys just check their account settings so the public knows what was set? EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation. | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||