Remix.run Logo
▲ gregw2 an hour ago

I am not the earlier poster, but I believe copyright infringement at the heart of their business model is what is referred to.

If I download a bunch of music from a torrent without a license, even if I don't listen to it, I'm liable, but if OpenAI or other LLMs gets content by some other unlicensed means (Anna's Archive, t), they are somehow not liable for the copy they made however temporary (but not so temporary if they leave it around to train a second model)?

And/or derivative works? And/or contributory copyright infringement when they regurgitate that copyrighted text when given certain prompts?

I get there is some nuance to copyright law, (four prongs), some utility to the outcome, and some legal (but not plausible) deniability. But there is no way they have clean hands on the copyright front for at least some actions they have taken. If there were, it would be in all their marketing and they would be pushing regulators to bind their competitors, onshore or offshore, more tightly in this regard.

I was around when search engines took advantage of similar ambiguities in copyright that took many many years to get litigated for similar reasons.