Remix.run Logo
calvbak 7 hours ago

I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.

defmacr0 5 hours ago | parent | next [-]

They're almost certainly pin-pointing high-quality conversations and giving them a special weighting. Seems stupid to not do that.

alansaber 4 hours ago | parent [-]

Oh they for sure classify conversations by type (cybersecurity, other guardrail proximates?) and quality.

clbrmbr 5 hours ago | parent | prev [-]

my understanding is that a sufficiently large model will memorize the training data once enough representations are built up. Opus 4 scale seems to have been sufficient. cf NYT vs OAI.