Remix.run Logo
dhx an hour ago

The US government's official position on LLMs is (very simply paraphrased) that LLMs are sufficiently transformative and do not hamper the potential market of authors of training material, therefore, copyright claims arising from training material should not be successful.[1] For original and creative training material, for example, a Harry Potter novel, seemingly the US government is asking the courts to set aside some previous questionable findings such as copyright existing very loosely in the likeness of fictional characters (impacting the likes of fan fiction). Can a human -- or LLM -- create a story about children travelling on a train from New York to a school of magic in the "wild west", with many loose similarities to Harry Potter for those familiar with those books? The US government appears to be saying this is OK, especially with the view that the market for Harry Potter is not diminished by a "wild west magic school" book in its likeness.

However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.

[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwzqxzbpw/...

triceratops 40 minutes ago | parent [-]

> LLMs are sufficiently transformative and do not hamper the potential market of authors of training material

And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.

Furthermore inference indirectly affects the market for the artist's future work. Don't need the writers and artists the LLM trained on anymore, when it can do similar work for free.