| ▲ | moritzwarhier an hour ago | |
They still scrape code, I'd guess, e.g. from GitHub? And there's tons of Claude-generated code there. Also, I'd guess this is not just to prevent any AI-generated code in the training data, but specifically their own. Percentage of users who put out their code on the web and also have a plan where Anthropic promises not to train on their data is problem also low. So not excluding own code could be a real issue, since it would be impossible to deduplicate the training and RILHF data from their sessions with the code accessible elsewhere, and written by the very same users. | ||