Remix.run Logo
profsummergig 15 hours ago

Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).

How was I not aware of this before?

vaylian 15 hours ago | parent | next [-]

AI is also trained on your HN posts. And lots of other things you post on the internet.

profsummergig 13 hours ago | parent [-]

Public posts on the internet are acceptable (to me).

For my (private) prompts, I need a warning telling me they may be used for training.

vaylian 9 hours ago | parent | next [-]

Facebook and other services are happy reading your private chats as well.

rramadass 10 hours ago | parent | prev [-]

> Public posts on the internet are acceptable (to me).

Everybody needs to rethink this again.

Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"?

Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get).

I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at.

PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.

alansaber 13 hours ago | parent | prev | next [-]

Everything. Your prompts, your conversation as a whole, public data, private data, usage metadata. It all goes into the big data machine.

cleaning 14 hours ago | parent | prev | next [-]

Good question, this was very well known. Do you have an answer?

profsummergig 13 hours ago | parent [-]

There is no fine-print (let alone a loud banner) on the chat thread page that tells me my prompts can be used for training.

asdff 36 minutes ago | parent | next [-]

Every single internet connect piece of software there is probably collects telemetry at this point. Why would this be any different? You know google logs your search data as well right? Not just the companies scan it but law enforcement too.

kzrdude 9 hours ago | parent | prev [-]

But the very fact that you go to "chatgpt.com" and write to them; "Dear Diary, today I thought.."; there is no reason they would not receive and process your data, unless explicitly promising not to (which also requires us to trust them).

The fundamental rule in this case is that if we offload our data to a cloud provider we can assume they read it, if they can, unless they promised very clearly they will not.

ga_to 14 hours ago | parent | prev | next [-]

Because you have not been paying attention to the discourse regarding AI for the last couple years? That AIs unethical train on data wherever they may get it from has been in the news basically weekly.

madethemcry 13 hours ago | parent | prev [-]

Don't make this our fault. I would even ask how is this not off by default or why aren't we asked upfront about it if they really care. It's disguising data collection as good faith. I don't even understand how this is legal under GDPR/EU given how much of PII they receive through chats.