Remix.run Logo
square_usual 6 hours ago

I think this is stupid, for three reasons:

1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.

2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.

3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)

solenoid0937 6 hours ago | parent | next [-]

> in this case too they didn't actually have the solution

Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.

I would almost expect training to overweight conversations with novel scientific and mathematical implications.

> the only reason they can't definitively say no is that for privacy reasons

They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.

Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.

> that opted-out user data was used for training

Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.

letmevoteplease 6 hours ago | parent [-]

You quoted the OP saying "in this case too they didn't actually have the solution" and responded with the totally unrelated, "Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data."

Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.

>They could 100% definitely say no, if they know they did not train on user data.

No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.

>They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.

They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.

faangguyindia 5 hours ago | parent | prev [-]

If the mathematicians are using ChatGPT, then they themselves are benefiting from the work of other ChatGPT users, so ChatGPT using their work is not wrong!