| ▲ | trescenzi 14 hours ago |
| If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know. |
|
| ▲ | yturijea 14 hours ago | parent | next [-] |
| It might be a lost cause regardless, because this is under the assumption that we can trust OpenAI to be honest about their own investigation, which is unlikely. |
|
| ▲ | fractorial 14 hours ago | parent | prev | next [-] |
| Only under gross negligence would it be unprovable: Did you use a model whose training set included user data? Did the transcripts of any of the agents include a tool call whose result including user data? |
| |
| ▲ | heaney-555 14 hours ago | parent | next [-] | | >Did you use a model whose training set included user data? OpenAI, as with all AI companies, openly admits that it trains on user data unless the user opts out. But the mathematicians have not said whether or not they opted out. >Did the transcripts of any of the agents include a tool call whose result including user data? They have already explicitly denied this. | | |
| ▲ | croon 11 hours ago | parent [-] | | >>Did the transcripts of any of the agents include a tool call whose result including user data? > They have already explicitly denied this. AFAICT they only denied accessing data through a request targeting a user, not that they accessed (users) data through targeting (an extremely niche) topic, which is the relevant part here. So, unless I'm mistaken, a "result including user data" is very much still in the air. |
| |
| ▲ | red75prime 14 hours ago | parent | prev [-] | | Or under strict privacy measures, maybe? When you are forbidden to connect a user and the user's data. |
|
|
| ▲ | xxs 14 hours ago | parent | prev | next [-] |
| It's the 'conversations' of some mathematicians with OpenAI. So the question is: did anyone feed the conversation(s) back as training data. |
|
| ▲ | worldsavior 14 hours ago | parent | prev | next [-] |
| What? Can they just look if they fetched certain documents/conversations and feeded them into the training loop? |
| |
|
| ▲ | heaney-555 14 hours ago | parent | prev [-] |
| This is exactly the issue. What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question! |
| |
| ▲ | cmiles8 14 hours ago | parent | next [-] | | Well that assumes OpenAI respects that setting. Given some of the company’s ethical challenges to date that’s not something folks are willing to just assume is happening. | | |
| ▲ | heaney-555 14 hours ago | parent | next [-] | | If you could prove that, it would be a gargantuan class-action lawsuit. And there would almost certainly be at least one whistleblower. | | |
| ▲ | cmiles8 11 hours ago | parent [-] | | See other post today about where OpenAI keeps flipping the toggle back to “share.” Even if it was technically turned on, the dark patterns that are clearly trying to override the obvious intent of the user are deeply unethical. |
| |
| ▲ | brookst 14 hours ago | parent | prev [-] | | How does OpenAI’s respect (or lack thereof) for the setting change whether the people involved could say whether they had opted out of training? They comment you replied to noted that none of them had shared that info. How are they blocked from doing so? |
| |
| ▲ | socialcommenter 14 hours ago | parent | prev | next [-] | | Andreas opted out on June 29th[0]. As discussed elsewhere on HN[1] [0] https://mathstodon.xyz/@andreasthom/117240535270608201 [1] https://news.ycombinator.com/item?id=49638353 | | |
| ▲ | heaney-555 14 hours ago | parent [-] | | Then none of his work after June 29 will be included in the training data. What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question. | | |
| ▲ | croon 11 hours ago | parent [-] | | > Then none of his work after June 29 will be included in the training data. What are you basing this on? The linked threads already acknowledged his opting out, and they don't look as settled as you claim. |
|
| |
| ▲ | asimpletune 14 hours ago | parent | prev | next [-] | | One of them said they turned it off in June, back in the original mastodon thread. | | |
| ▲ | heaney-555 14 hours ago | parent [-] | | Then none of his work after June will be included in the training data. What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question. |
| |
| ▲ | aenis 14 hours ago | parent | prev [-] | | And its reasonable to assume that if they did, in fact, opt out, they'd make it clear. Most people do not understand that the main reason for the subscriptions is to give OpenAI and Anthropic the priceless, unique data that shows how the models are used, what people are building, how they are building, which solutions they consider OK, which they consider bad -- they purchase this data with cheap tokens. This is their only moat, really. If some really proprietary IP gets swept in the training data set its not really OpenAI's fault -- its the researchers'. Have something secretive? Dont fricking paste this into chatgpt. Duh! (I'd definitely not think OpenAI/Anthropic ignore the opt outs, or ZDRs. All it would take is one whistleblower to get them into terminal troubles. And why would they do it? They are not in the business of scooping unique IP -- they are in the business of understanding how AI is used across a variety of mundane, day to day work of individuals and companies. Useless math problem is good (or bad, as in this case) PR, but otherwise entirely worthless for the labs. |
|