| ▲ | yorwba 2 days ago | |||||||
GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection. | ||||||||
| ▲ | spwa4 2 days ago | parent [-] | |||||||
You mean very likely the Anna's archive torrent dump because it's MUCH better quality than the general internet and beyond a certain amount of input data (which is a lot, but much less than the internet) the only thing that matters in training is the quality of the data, to the point that now many labs have thousands of people just making and improving essentially school exercises full time? Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad. | ||||||||
| ||||||||