Remix.run Logo
▲ AI companies leak data to advertisers [pdf](jorgegarciaherrero.com)
223 points by damaru2 4 hours ago | 58 comments
▲delis-thumbs-7e 2 hours ago | parent | next [-]

In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.

We have all become Milhouse now.

▲Avicebron 2 hours ago | parent | next [-]

Inequality has eroded trust in society in ~50 years, a lot of the old models (heh) of how we see the world aren't relevant. It's hard to exist when everything around us is adversarial.

▲flipping_beacon an hour ago | parent | prev [-]

This can't be real I am go gonna ask Willie

▲Coeur 3 hours ago | parent | prev | next [-]

"multiple providers disclose sensitive conversation-derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation."

Not good at all.

▲postalcoder 3 hours ago | parent | next [-]

My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy.

Perplexity does this. Visiting a past perplexity search url exposes your full conversation.

▲albert_e 2 hours ago | parent | next [-]

Security by obscurity -- such an age old anti-pattern!

I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link

Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped

There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines

This is shockingly lax approach to data security and privacy by design

▲kevindamm an hour ago | parent | next [-]

The same assumptions are true about giving any human that shareable link. They could pass it on to anyone, screenshot it, paste it into their own session. This has been true since before "share with link" permissions on Docs and elsewhere.

If you click "provide a shareable link" you should decide (and behave) as though that made it public.

I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."

▲postalcoder 2 hours ago | parent | prev [-]

Chat UIs are a minefield of “if you accidentally click this your data will be shared or trained without you realizing it!”

▲msdz 2 hours ago | parent | prev [-]

Genuinely asking: If you don’t share the UUID-based URL yourself, what makes it not privacy-friendly?

It’s not like someone’s gonna guess that URL… right?

▲postalcoder 2 hours ago | parent [-]

Yes, technically, guessing a url is impossible. But browser histories are stored in cleartext and trivially accessible to sketchy actors.

I also accidentally paste random stuff into input boxes all the time.

▲msdz an hour ago | parent [-]

Good points, thanks.

Although I think at the point of some on-device program reading your browser history against your will, you’re gonna have bigger problems.

▲kjs3 an hour ago | parent [-]

You already have bigger problems then. Most people have given any number of plugins, etc., access to peek at browser history, clipboards, etc and didn't realize it. Run an ad blocker? VPN? Check out the permissions those sorts of things have on your device.

▲Aboutplants 2 hours ago | parent | prev | next [-]

“Screenshots”

The amount of sensitive information that accidentally gets left on screenshots is pretty large. This is a pretty massive security issue

▲baggachipz 13 minutes ago | parent | prev [-]

I'm shocked... SHOCKED! that these companies would sell this data to advertisers and violate the privacy of their users. Who could have seen this coming??

▲drywater2 2 hours ago | parent | prev | next [-]

No, they don't "leak data", data is sold. Leaking data requires a mistake. This is intentional.

▲kjs3 an hour ago | parent [-]

Thank you. I keep saying this and the spin doctors keep using words that imply "oh, my...I am so sorry we had that tiny little privacy issue...we'll get right on that in the next sprint...". And winning the message war.

It's not a 'leak'. It's in the T&Cs you agreed to. This is the business plan.

▲segmondy 10 minutes ago | parent | prev | next [-]

Prior to this new AI age, your data was calculated and what was inferred about you was "shallow", but as of today. The sort of profile and things that can be known about you is scary especially if you are constantly engaged with cloud AI. IMO, the number one risk of using cloud AI is loss of privacy and loss of freedom. With AI and capabilities, more controls can be placed on people and the more you put yourself out there, the more you are going to lose.

For example, we now have self driving cars, we have cameras everywhere. Based on your chat with a cloud AI, you can automatically trigger an automatic monitoring event that follows and tracks you in the real world with the fleet of cameras, cars, GPU, cell signal. Your tracking due to AI has moved into the real world and eventually, a self driving car will take you in to be "processed" against your will, not even for what you posted in a public forum, but for your private ribbing and chatting with some cloud AI.

So definitely put local AI into the mix and keep personal stuff and thoughts local only.

▲pbasista an hour ago | parent | prev | next [-]

Tangential:

I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.

This partial prompt data might potentially be used to "pre-warm" some kind of cache.

But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.

▲j4k0bfr 2 hours ago | parent | prev | next [-]

This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!

My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.

Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).

▲amarcheschi 2 hours ago | parent | next [-]

I'm taking an onboarding process for an Ai company helping other (much) bigger Ai companies and the amount of vibecoded platforms and documentation is staggering. Like, training process so broken that the platform just doesn't load sometimes, things that have never even been tried are published and you have to use them and they suck so much because it is apparent that no human ever touched that and probably wouldn't want to

▲Forgeties79 an hour ago | parent [-]

Doesn’t sound like your new company is going to be very helpful for other AI companies lol

▲amarcheschi 33 minutes ago | parent [-]

It's a part time job hiring contractors, I'm doing this only because they pay would be nice but I know very well I can be fired anytime or when the project is done...

Anyway I'm a student so I'm not looking for something stable. BTW now it's getting better but if I were the client paying for those services I'd be pretty fucking furious for what's going on - guess they'll never know tho

▲alansaber 2 hours ago | parent | prev | next [-]

AI companies rushing an implementation? Surely not :).

▲Forgeties79 an hour ago | parent [-]

No no it’s hyperscaling

▲internet_points 11 minutes ago | parent [-]

hyperleaking

▲mrweasel 2 hours ago | parent | prev [-]

So I only read the abstract, but the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).

I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.

▲kjs3 an hour ago | parent | next [-]

the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.

▲duskdozer 2 hours ago | parent | prev [-]

Well, some data is deliberately provided at least. It's not a situation of other apps managing to grab the data like the facebook-localhost exploit

>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.

▲kdaniel_03 an hour ago | parent | prev | next [-]

It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't. Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.

▲20k 2 hours ago | parent | prev | next [-]

A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you. They don't care about the agreements you've signed. You're just a pile of cash to them

▲int3trap 2 hours ago | parent | next [-]

> A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you

Who thinks OpenAI or any big company for that matter give a crap about them? This isn't a popular sentiment at all, it's just patently false.

▲oblio an hour ago | parent [-]

All those people with AI friends, boyfriends, girlfriends.

▲Mistletoe an hour ago | parent | prev [-]

That’s not fair, we are also more free training and refining of their models.

▲vivekpolavarapu an hour ago | parent | prev | next [-]

Ad-tech spent 20 years trying to infer intent from clickstreams. Chat apps now hand over an AI-written one-line summary of intent, labeled and keyed to a cookie. Are we sure the ads business model in AI is about ads in the chat, and not the chat as the targeting signal?

▲bix6 a minute ago | parent | prev | next [-]

Hard to read on mobile. Is there a tldr of how bad each provider is?

▲dhanushnehru 13 minutes ago | parent | prev | next [-]

It’s not a leak if it’s the business model.

▲Traster 2 hours ago | parent | prev | next [-]

I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.

First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.

▲delis-thumbs-7e an hour ago | parent | next [-]

> First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!

> But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.

▲jeltz 2 hours ago | parent | prev [-]

> a whole bunch of their execs come from Meta, and all those guys learned the hard way.

Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.

▲0xcrypto 2 hours ago | parent | prev | next [-]

This is why I built my own chat interface. https://ai.ivx.run/

Access through: https://ai.ivx.run/chat/

▲DrMandalay 3 hours ago | parent | prev | next [-]

The word is "sell" not "leak". This title takes away all agency from the thieves selling private data to advertisers.

▲otabdeveloper4 2 hours ago | parent [-]

The personal data fell off the back of a truck, man.

▲pluc 2 hours ago | parent | prev | next [-]

You thought... they didn't?

▲gagan2020 3 hours ago | parent | prev | next [-]

All sells but I saw Chinese models are upfront about that most of the time.

▲classified 3 hours ago | parent | prev | next [-]

Is it still called a leak if it was the whole point and purpose of the deal?

Someone should have to investigate, but I suppose it's all "legal"?

▲lava_pidgeon 3 hours ago | parent [-]

In the US.

In EU law it is very likely against GDPR.

▲okokwhatever an hour ago | parent | prev | next [-]

I never thought this could happen (XD)

▲nwhnwh 2 hours ago | parent | prev | next [-]

I am very surprised.

▲folkrav 3 hours ago | parent | prev | next [-]

Insert surprised pikachu meme

▲reedf1 3 hours ago | parent | prev | next [-]

my first guess is always Gboard.

▲Grimeton an hour ago | parent | prev | next [-]

Oh no, what, what happened?

What happened? Oh no!

How terrible! That’s just, that’s just awful!

How terrible! Oh no!

▲charcircuit 3 hours ago | parent | prev | next [-]

The paper doesn't say when the app sends the conversion artifact.

▲robertclaus 2 hours ago | parent | prev | next [-]

Hanlon's Razor given that these tools are almost certainly vibe coded at this point?

▲jeltz 2 hours ago | parent [-]

Incompetence and carelessness is rampant in our industry abd with vibe coding it has only gotten worse.

▲999ziyadej 2 hours ago | parent | prev | next [-]

data is sold i think

▲damaru2 4 hours ago | parent | prev [-]

Entire conversations via exposed permalinks. For Grok: trackers receiving the conversation URL could access the full chat because the link lacked access controls.

Screenshots of conversations. TikTok received screenshots of Grok chats during sharing, exposing the actual visible conversation content.

Conversation-derived content tied to persistent identifiers, including prompts and automatically generated chat titles revealing sensitive facts. "Salary 85k NYC: mortgage 280–350k".

▲Traster 2 hours ago | parent [-]

That is a staggering level of incompetence.

▲ex1fm3ta 16 minutes ago | parent [-]

what do you want it's the era of vibecoding.