Remix.run Logo
alansaber 7 hours ago

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

calvbak 6 hours ago | parent | next [-]

I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.

defmacr0 4 hours ago | parent | next [-]

They're almost certainly pin-pointing high-quality conversations and giving them a special weighting. Seems stupid to not do that.

alansaber 4 hours ago | parent [-]

Oh they for sure classify conversations by type (cybersecurity, other guardrail proximates?) and quality.

clbrmbr 4 hours ago | parent | prev [-]

my understanding is that a sufficiently large model will memorize the training data once enough representations are built up. Opus 4 scale seems to have been sufficient. cf NYT vs OAI.

sigbottle 5 hours ago | parent | prev | next [-]

In general, a lot of moral invariants that natural selection has rendered as "intuitive" to us are no longer intuitive or possible. These natural brakes are not braking.

utopiah 6 hours ago | parent | prev [-]

This is such a naive position though.

The most successful companies of the last decade have precisely been ... selling usage data.

Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.

alansaber 5 hours ago | parent [-]

Yeah but marketing companies are aggressively fingerprinting and stalking you to sell you snacks from japan, or oscilloscopes because they figured out you work in a lab, etc. Not to fuck you over by stealing your livelihood (which is what is happening to these mathematicians). It's on a whole new scale.