| ▲ | spwa4 2 days ago | ||||||||||||||||
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course. | |||||||||||||||||
| ▲ | yorwba 2 days ago | parent [-] | ||||||||||||||||
GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection. | |||||||||||||||||
| |||||||||||||||||