Remix.run Logo
mewse-hn 8 hours ago

"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?

WarmWash 8 hours ago | parent | next [-]

Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.

If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".

nozzlegear 4 hours ago | parent | next [-]

> If you need privacy, then you are going to have to pay full price for those tokens (API).

At this point, how can we even trust that they aren't accidentally training on those tokens too?

jdm2212 4 hours ago | parent [-]

It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.

nozzlegear 3 hours ago | parent [-]

> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.

One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.

Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.

WarmWash 3 hours ago | parent [-]

Neither of those other two examples would stop corporations from giving you money. Bold faced lying to them would.

Boardrooms run businesses, not bookstore ethics clubs.

14u2c 5 hours ago | parent | prev | next [-]

You can also pay for their business plan, which includes data controls and starts at $50/mo (2 seats). Not exactly a high bar.

spruce_tips 5 hours ago | parent | prev | next [-]

what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?

magicalhippo 4 hours ago | parent [-]

It's quite well explained here[1], which is linked from the Privacy section of their plan overview[2].

Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.

You'd have to take their word, but that goes for anything in life.

[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...

[2]: https://chatgpt.com/pricing/

TZubiri 5 hours ago | parent | prev | next [-]

>They only fuck over the poor ones, I can pay the expensive prices so this is not a problem.

perching_aix 7 hours ago | parent | prev [-]

There's literally an opt out toggle even pesky peons like me can peruse, actually.

lima 5 hours ago | parent [-]

They may still train on it if you submit feedback or flag a safeguard. The terms are a bit fuzzy on this.

nradov 8 hours ago | parent | prev | next [-]

Is it spying? I think this usage is disclosed in their terms of service.

gowld 8 hours ago | parent [-]

If it happened it's plagiraism. Consent to see data isn't consent to claim priority.

red75prime 7 hours ago | parent | next [-]

Establishing plagiarism requires sufficient similarity between works. Training data changing a model’s weights in some direction, and the model then producing a different solution, hardly qualifies.

But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.

brainwad 6 hours ago | parent | prev [-]

I mean... none of these humans have priority. The result is due to the team of LLM agents.

dash2 8 hours ago | parent | prev | next [-]

If they had agreed to let OpenAI train on their data, it wouldn’t be spying.

8 hours ago | parent [-]
[deleted]
sinuhe69 6 hours ago | parent | prev | next [-]

More like helped improve our work (the disproof)

elwell 6 hours ago | parent | prev | next [-]

Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)

ImaCake 3 hours ago | parent | next [-]

Yes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.

taylorfinley 4 hours ago | parent | prev [-]

It's a bit disingenuous to preface a disclosure like this with an unsubstantiated assessment of its likeliness. It is a press release, I'm not sure we owe it credulity.

vessenes 6 hours ago | parent | prev | next [-]

If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.

All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.

kypro 3 hours ago | parent | prev | next [-]

I think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny).

The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.

I get the scepticism, but I feel some of the accusations here are bad faith.

jimbob45 6 hours ago | parent | prev [-]

What does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.