Remix.run Logo
jwr 2 hours ago

I always thought it was enough to switch off the "Improve the model for everyone" setting on chatgpt.com:

"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy."

But apparently there is also an entire completely different route "Do not train on my data"?

Does this mean that before I submitted the "Do not train on my data" request, my data was used for training in spite of "Improve the model for everyone" being turned off?

We are getting to facebook/meta-levels of privacy settings obfuscation.

jsw97 an hour ago | parent | next [-]

The last time I checked there was a loophole — if you provide feedback in-session (responding to “how are we doing” or “which prompt is better”) then they can use that feedback + relevant context. Relevant context for chatgpt might include memories / other sessions. That may not be the only loophole.

That in itself is a dark, dark pattern. There should at the very least be explicit warnings for users who have checked “do not train”; or they should not be presented with such dialogs.

teiferer 2 hours ago | parent | prev [-]

That's because what people enter into LLMs is the last gold there is out there. Everything else is already scraped or ensloppified.

Maybe next step is to filter your input client side through an unknown number of obfuscators where you ask LLMs to rephrase your question (onion router idea) such that no single provider can be certain that this is human input and not some slop feedback loop.