Remix.run Logo
postalcoder 6 hours ago

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).

People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:

  1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)

  2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.

  3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.

  4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).

edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.

ahsg17 6 hours ago | parent | next [-]

> Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

Yes folks, please moderate yourselves and talk meekly like the academics on Mastodon, so that the IPOs aren't in danger and nothing will ever change.

larodi 5 hours ago | parent | prev | next [-]

> If it is found that OpenAI and other labs are not respecting the training opt out

HOW?? how precisely do we/them/us find this, given said companies are 100% non-auditable by external parties. how? if not by blaming them with evidence, anecdotal if it can be. no really, how do we find it out, surely not by lashing out at teach other on HN!

taylorfinley 3 hours ago | parent | next [-]

Maybe we can create some extremely low probability sentences and make sure to include them in our chats. If a future model can re-create the very low probability sequence, we have proof of our "private" conversation being trained on or accessed.

postalcoder 4 hours ago | parent | prev [-]

What do we need to audit? The researcher in question here did not opt out of training until a few months ago.

chunky1994 5 hours ago | parent | prev | next [-]

Why are we being so charitable to trillion dollar organizations here? If OAI keeps re-enabling the train model toggle on every app update to codex, does it also fall under "common sense opsec" to re-disable this toggle every time?

Arguably you expect that unless you are explicit about providing permissions to these labs to use your data for training then your data is yours, and not theirs. Especially on a paid account (let alone an enterprise one). Why is the opt-out supposed to be "common sense opsec" rather than the opt-in should be common sense regulation?

BeetleB 2 hours ago | parent | prev | next [-]

> Anyone remotely familiar with CC over the years

years?

SpicyLemonZest 6 hours ago | parent | prev | next [-]

This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.

I understand why the nature of AI products makes it harder to avoid this category of issue, nearly impossible to prove that it didn't happen if it could have, and easy to stumble into it without any human being intending harm. But those factors are exactly what people have in mind when they say OpenAI "steals" intellectual property! If OpenAI doesn't want people to be nasty to them, they'll have to find better solutions.

cma 6 hours ago | parent | next [-]

I think Google does train on anything you put into docs if you aren't careful with the Gemini integration?

SpicyLemonZest 5 hours ago | parent [-]

Yes, this is a problem with modern AI systems in general. It's not just OpenAI, and if you know any artists you know this is why they're pretty vehemently opposed to all AI.

lowbloodsugar 5 hours ago | parent | prev [-]

That’s … Googles entire reason for making these “you don’t pay with money” tools. Did you not understand that?

HDThoreaun 2 hours ago | parent | next [-]

Google docs exists as a competitor to microsoft office. Google gives it away free to consumers for the same reason AI labs sell subscriptions for 10% of the price of the api. They hope businesses will switch to what employees know how to use.

SpicyLemonZest 5 hours ago | parent | prev [-]

What? I don't understand how you even came up with this idea, much less consider it so obvious to condescend about it. Do you have even a single example of a research project that got scooped because the Google Docs team forwarded their private documents to someone?

lowbloodsugar an hour ago | parent [-]

Lol. That's not what I said nor what TFA is claiming. The claim is that private data is used for training.

thevillagechief 5 hours ago | parent | prev | next [-]

You know, I don't think I've ever accused anyone of being a shill. I've thought about it maybe a few times (daringfireball). This is going to be as close as I get. I don't know the facts in this case but I cannot believe the argument being made here with a straight face. Is it common sense that tools you use and pay for steal your work and profit off of it at your expense and without recognition? If this isn't the textbook definition victim blaming, I don't know what is.

mittensc 6 hours ago | parent | prev | next [-]

Imagine OpenAI Astra model weights were made public because the datacenter they use had T&C that allows them to make them public

Would that be ok in your mind?

Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did)

Nobody would care if they provided published research that author made public same as a google search would offer that.

aurareturn 5 hours ago | parent [-]

  Would that be ok in your mind?
It would in my mind. Hopefully companies have looked through the agreement.
mittensc 3 hours ago | parent [-]

well then fingers crossed someone does that

1294827 6 hours ago | parent | prev [-]

[flagged]