| ▲ | wat10000 2 days ago | |||||||||||||||||||||||||
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders. But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs? | ||||||||||||||||||||||||||
| ▲ | cyberpunk 2 days ago | parent [-] | |||||||||||||||||||||||||
you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them. distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work. so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||