|
| ▲ | mucha 4 hours ago | parent | next [-] |
| That's not what your Chief Research Officer, Mark Chen, says on X:
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." https://x.com/markchen90/status/2097400166554993041 |
| |
| ▲ | sebzim4500 3 hours ago | parent | next [-] | | Can you explain what part of his post you believe is inconsistent with that quote? | | |
| ▲ | mucha 3 hours ago | parent [-] | | "If they opted out of training, then we definitely did not train on them." Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me. | | |
| ▲ | derangedHorse 3 hours ago | parent [-] | | > Per OpenAI's privacy policy, they use de-identified data to improve their products That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products. | | |
| ▲ | mucha 2 hours ago | parent [-] | | It appears than you can only opt-out from having OpenAI train models on your data. There isn't an option for opting to exclude your de-identified data from being used to improve OpenAI products. | | |
| ▲ | hellohello2 an hour ago | parent [-] | | Are you certain of this? I would be inclined to believe you but it would be nice to know decisively. |
|
|
|
| |
| ▲ | vemacs 3 hours ago | parent | prev [-] | | Does OpenAI think de-identified data is no longer user data? Wild take for OpenAI and certainly not industry standard. |
|
|
| ▲ | lambda 4 hours ago | parent | prev | next [-] |
| But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being. Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set. |
|
| ▲ | Imnimo 3 hours ago | parent | prev | next [-] |
| I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it also seems at least possible that a key idea or a particular step found its way into training data. It wouldn't mean OpenAI stole their proof - clearly the model developed its own approach. Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this. |
| |
| ▲ | derangedHorse 2 hours ago | parent [-] | | > It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this. Fall “victim” to what? Having their responses in the training data if they fail to opt out? That is what will happen. If you’re referring to falling “victim” to OpenAI scooping a problem discussed in training, this also wasn’t the case. They chose the problem based off human-spread rumors. |
|
|
| ▲ | calf 2 hours ago | parent | prev | next [-] |
| Your perspective is not helpful until you read and reflect on Tristan's letter stating serious grievances. Your remarks here have minimized his complaints and that is a sign of bias. Do not then pre-accuse HN commenters of being convinced when there reasonable skepticism such biased behavior showing itself in this very thread, saying things that amount to "my tribe/company would never be so egregious and if you think that then it is bad faith". That's the projection. If the word prejudice means anything to you then please do the work of attending to that instead of using the platform to reinforce such biases. If you are not a PhD yourself maybe your are not culturally qualified to assess and expound on the overall situation anyways. |
|
| ▲ | lossolo 4 hours ago | parent | prev [-] |
| > If they opted out of training, then we definitely did not train on them. Can't you guys just check their account settings so the public knows what was set? EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII.
If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation. |
| |
| ▲ | derangedHorse 2 hours ago | parent [-] | | I don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior. | | |
| ▲ | lossolo 2 hours ago | parent [-] | | Yeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all.
Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different. |
|
|