| ▲ | calvbak 7 hours ago | |||||||
I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad. | ||||||||
| ▲ | defmacr0 5 hours ago | parent | next [-] | |||||||
They're almost certainly pin-pointing high-quality conversations and giving them a special weighting. Seems stupid to not do that. | ||||||||
| ||||||||
| ▲ | clbrmbr 5 hours ago | parent | prev [-] | |||||||
my understanding is that a sufficiently large model will memorize the training data once enough representations are built up. Opus 4 scale seems to have been sufficient. cf NYT vs OAI. | ||||||||