Remix.run Logo
▲ rossy 5 hours ago

There are three things that have become very clear to me after the rise of generative AI:

1. The sheer amount of material on the internet that is "free to view but not free to use for any purpose" is the greatest resource of our time, and despite it being easy for individuals to take advantage of it (no one will take you to court for printing a newspaper comic and pinning it to your corkboard,) it's historically been difficult for corporations to exploit it (their best idea pre-AI is to encourage people to post it on social media walled-gardens where they can surround it with ads.)

2. The reason behind the impressive results of generative AI is because it exploits the above "free" resource, which is the greatest resource of our time. The reason behind the industry-wide push for AI and the insane amount of investment in it, is that they know it's their first real chance to exploit the greatest resource of our time. This is the gold rush.

3. Anthropomorphism is the wool that AI labs are pulling over legislators eyes so they can pull off this heist. If you see training and inference as a black box, a process that consumes a copyrighted work (among others) and produces something very similar to the original work that also competes directly with it, is clearly something that's against the spirit of copyright. But if you (afraid of being judged a luddite) see AI as a little man inside the computer who is "learning" and "creating," how could you deny him? Especially if it would deny your jurisdiction access to the above gold rush. A lot of scientific-sounding AI communication is propaganda for this way of thinking, like the Anthropic J-space stuff, which stops just short of claiming AI is conscious, despite leading the reader to that conclusion.

▲derektank 2 hours ago | parent [-]

You don’t need anthroporphism to draw the conclusion that using copyrighted material in model training is fair use. It extends pretty directly from existing fair use doctrine surrounding transformative uses of technology, in particular from Authors Guild v. Google, where the 2nd circuit ruled that creating a searchable database of copyrighted material (i.e. Google Books) was transformative, as long as the material presented to the end user was limited snippets of text and not the entirety. Bartz v. Anthropic explicitly cites that case as precedent

▲michaelt a minute ago | parent | next [-]

In this instance the training data was New Yorker cartoons signed by Brendan Loper, and the material presented to the end user is New Yorker-style cartoons signed by Brendan Loper.

To me that seems a exceedingly broad definition of fair use.

▲franga2000 35 minutes ago | parent | prev [-]

Google: scans all books, people can search for them, find little snippets, see the title and author, buy the book to read it

AI: scans all art, people can ask it to produce art they're searching for, no reference to the original or its author, people no longer pay artists/designers/...

I think there's a very obvious distinction there. The whole purpose of copyright is ensuring the financial viability of creating works. AI slop is a direct market substitute for the originals it ate up for training, so it goes directly against that. Meanwhile, search actually improves the reach of works and, well, snippets and summaries are somewhere in between...