Remix.run Logo
▲ prathje 4 hours ago

What about IP and copyright?

We got asked this question two weeks ago when we conducted a workshop on effective and responsible use of AI tools at a local conference. But how can you argue about about IP and copyright if the breakthrough of LLMs is potentially based on circumventing or breaking IP and copyright in the first place?

How do you feel about all of this?

▲ben_w 2 hours ago | parent | next [-]

> What about IP and copyright?

For training the models, and assuming that the content was itself acquired without other acts of infringement? That was ruled legal by the judge in the case I actually (skim) read the judgement of.

At least two companies engaged in acts of infringement to get training data. This is not lawful, and what Anthropic settled out of court for.

▲prathje an hour ago | parent [-]

Of course, if the data/knowledge is open and accessible, why not?

▲GJim 38 minutes ago | parent [-]

open and accessible != free from copyright

▲ben_w 19 minutes ago | parent [-]

As per court ruling, training is fine when they have lawful access.

i.e. open and accessible on the general web is fair game, torrents from the pirate bay is not.

Feel free to argue that copyright law should be changed; this wouldn't be the first time it needed a significant update because new technology made it cheap to do at industrial scale something that was previously so hard that even being able to pull it off made you look legit.

▲sambeau 3 hours ago | parent | prev | next [-]

I feel for Aaron Swartz, his poor mother, and what the US government put him through.

▲duskdozer 3 hours ago | parent | next [-]

His big mistake was clearly giving away what he got for free instead of trying to make a lot of money off it

▲Arkhaine_kupo 3 hours ago | parent | prev [-]

Seeing Mark Zuckerberg use THE EXACT SAME papers that Aaron was murdered for to train Llama and him being rewarded with a stock tick that got him out of the metaverse hole will never not make my blood boil.

▲JimDabell 2 hours ago | parent | prev | next [-]

> how can you argue about about IP and copyright if the breakthrough of LLMs is potentially based on circumventing or breaking IP and copyright in the first place?

Easy answer: it’s not. Learning isn’t copyright infringement. Never has been and hopefully never will be.

▲dimbletimbers an hour ago | parent | next [-]

This is another case of anthropomorphic language for AI failing us. Who could be opposed to “learning?”

If I learn from a textbook I stole, some people maybe be thrilled I learned something, but stealing is still a crime and hopefully always will be.

▲prathje an hour ago | parent | prev | next [-]

Well, of course learning isn't the problem and shouldn't be. But what about reselling/commercialising that knowledge? That is the big (commercial) opportunity. But if AI models learn by fitting their parameters to the given data, then overfitting on said data could result in the IP or copyrighted being "slopped" out verbatim?

▲semiquaver 25 minutes ago | parent [-]

The knowledge isn’t copyrighted and, thank god, cannot be.

I thought we were supposed to be hackers. Information wants to be free and all that. What a bunch of hypocrites, willing to betray their ideals as soon as it’s convenient.

▲LunaSea an hour ago | parent | prev | next [-]

Learning isn't but using that knowledge is.

▲bluefirebrand an hour ago | parent | prev [-]

> Learning isn’t copyright infringement. Never has been and hopefully never will be.

For humans.

Why should computer systems owned by corporations have anything remotely similar to the same freedoms as humans?

▲bko 3 hours ago | parent | prev [-]

It's no different than humans taking inspiration from IP and copyrighted works in their own creative endeavors. Imagine if programmers were unable to learn from open source. Or artists were unable to mimic styles and storylines. Musicians can't riff on things that others created.

▲snarfy 2 hours ago | parent | next [-]

It is different. The scale makes it different. It's different when it is all copyrighted works and IP. It's different when you solve Navier-Stokes by taking "inspiration" from a researcher's private work.

▲devsda 2 hours ago | parent | prev [-]

If it's "learning" is to be seen in the same context as a human, then the punishment/ repercussion for hacking and causing real harm should also be similar to what a human might recieve.

It cannot opportunistically morph between a human for benefits and a machine for liability.

▲ben_w 2 hours ago | parent [-]

> It cannot opportunistically morph between a human for benefits and a machine for liability.

Indeed. Though here it would be morphing between human and machine even just for pure liability, which is even worse.

We have no way to tell if a machine has an experience (in general and not only) when being positively or negatively rewarded for what it produces; even if we assume it does, we have no way to tell if punishments are experienced like mild scolding (and rewards like mild praise), or if the punishments feel like being burned to death at the stake (while rewards feel like a religious experience combined with an orgasm).

We are just beginning to scratch the surface of what these kinds of question even look like in mechanistic terms; all philosophical discussions before it about p-zombies and the Chinese room and so on, they are all no more useful than any other non-expert armchair experts discussing things.

Even the current research, such as it is, is probably only at the level analogous to humoral theory when we want something at the level of germ theory.

https://en.wikipedia.org/wiki/Humorism