Remix.run Logo
rich_sasha 4 hours ago

I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.

Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.

brainwad 4 hours ago | parent | next [-]

Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.

Planktonne 3 hours ago | parent | next [-]

> Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.

It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources.

I'm not convinced that a competitive use at one remove should be treated as not competitive.

brainwad 3 hours ago | parent [-]

I think that's more true in image generation than in text? At least, all the money is in LLMs that write code, not LLMs that write O'Reilley-style coding books.

wongarsu 2 hours ago | parent | next [-]

If you have a websites that offers guides, how-tos or tutorials, LLMs directly compete with you. StackOverflow would also have a really good case

After all LLMs don't just code, they also answer questions and give step-by-step instructions. In terms of total userbase those features are used a lot more than writing code

_flux 3 hours ago | parent | prev [-]

I thought LLM-authored books were rampant in Amazon, though?

nixpulvis 3 hours ago | parent | prev | next [-]

I think LLMs providers pretty directly compete with content they scrape like Wikipedia and SO...

bayindirh 3 hours ago | parent | next [-]

Remember the golden rule of the golden rules:

Who has the gold makes the rules.

atemerev 3 hours ago | parent [-]

If a society operates under a rule like this, it is no better than Russia or any other tyranny where law is for me but not for thee. This is not how it should work in a supposedly free and lawful country.

bio_hacker 2 hours ago | parent [-]

That’s how works everywhere

bluefirebrand 3 hours ago | parent | prev [-]

No kidding. I dunno what kind of traffic loss Wikipedia has had but SO is really dead these days

If they weren't competing with AI then why is AI killing it?

brainwad 3 hours ago | parent [-]

There's a difference between creating a market for something better, so that nobody wants the old thing, and competing _in_ the market for the old thing by copying it directly.

immibis2 32 minutes ago | parent | next [-]

If two things are competing, they are in the same market by definition.

wongarsu 2 hours ago | parent | prev | next [-]

If LLMs only made SO redundant by writing code and solving my technical problems autonomously so I never have to think about it, I would agree. But often I do ask LLMs technical questions, and they answer in great detail. And that part is a very direct SO competitor

darkwater 3 hours ago | parent | prev | next [-]

And what would be a read-only version of X like XCancel compete against, exactly? Ads impressions? That would be the only possible thing yet they don't add any ads.

brainwad 3 hours ago | parent [-]

It's depriving X of impressions that they could monetise, no? Xcancel doesn't have to make money itself, it just has to impair the rights of the copyright holder. Otherwise piracy would also be legal as long as it were non-profit...

darkwater 2 hours ago | parent [-]

> Otherwise piracy would also be legal as long as it were non-profit...

Which is in a few jurisdictions, or at least is not prosecuted if it's for personal use. Also, according to your definition, the creator of uBlock Origin or any other adblock system should be sued in the same way, because they are depriving $ADS_CORP of their precious impressions.

brainwad 2 hours ago | parent [-]

Well, adblockers don't copy the copyrighted content. They just control how it's rendered on the user's machine. Copyright cares about making copies and especially distributing them.

darkwater 2 hours ago | parent [-]

You have a point on this, I recognize, but it still seems a very thin line to walk (for X) - at least morally, because I don't think they are actually loosing real big money to anyone.

bluefirebrand 2 hours ago | parent | prev [-]

I don't really buy this.

It's like saying "toaster oven/air fryer combos" don't actually compete with toaster ovens or air fryers because they are creating a market for something better

Of course they complete.

Toaster ovens compete with toasters. Microwaves compete with toaster ovens.

Just because it's not the exact same product doesn't mean it's not competing

Would I be allowed to steal LG's designs for a microwave and make a "superwave" that does laundry and heats food? Would you claim those products don't compete because the superwave is "something better"?

brainwad 2 hours ago | parent [-]

It's not illegal to write a similar book or song to an existing one. Copyright only protects existing works from literal copying (possibly in part).

immibis2 28 minutes ago | parent [-]

And derivative works, like AI makes

bonsai_spool 3 hours ago | parent | prev | next [-]

> Precedent is pretty clear:

What cases are you citing when you say this?

brainwad 3 hours ago | parent | next [-]

Bartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM.

lesuorac 3 hours ago | parent | next [-]

Bartz is an author though.

Is X claiming ownership of the posts people make because pretty much every single social media site doesn't so they have section 230 protection.

brainwad 3 hours ago | parent [-]

They can just round up some friendly users and sue under their names. Starting with their own corporate accounts?

lesuorac 8 minutes ago | parent | next [-]

That would definitely limit damages to strictly those accounts.

I'm not even sure he can use his own account as one of them. The SEC might be pretty friendly to him but I'm not sure that limiting access to a location where material information about Tesla/SpaceX is provided won't become a problem.

But I'm not even sure what damages the accounts are suffering as revenue sharing is going away [1]. With Bartz the damage is a loss of sale. With X the damage is $0 per post to the poster.

There is a newer Original Content Rewards program [2] but it seems to split revenue from X Premium and presumably people that have X Premium are not using XCancel so the damages would be 0.

[1]: https://help.x.com/en/using-x/creator-revenue-sharing

[2]: https://help.x.com/en/using-x/original-content-rewards

bonsai_spool an hour ago | parent | prev [-]

These statements would suggest that the precedent is not, in fact, clear

3 hours ago | parent | prev [-]
[deleted]
immibis2 3 hours ago | parent | prev [-]

Perhaps https://en.wikipedia.org/wiki/Warner_Bros._Entertainment_Inc....

ratelimitsteve 2 hours ago | parent | prev [-]

How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from)

bonsai_spool an hour ago | parent [-]

> How clear is it when I google a recipe and get an AI-generated recipe

Recipes can’t be copyrighted

Here’s one discussion about this https://www.nycbar.org/reports/secret-ingredients-how-to-pro...

ratelimitsteve an hour ago | parent [-]

fair, but recipes aren't the only thing where AI summaries at the top of search results are simultaneously transformative and competitive. Any information that is scraped from a website and then summarized by AI in a way that prevents that website from getting views is both transformative and competitive.

embedding-shape 4 hours ago | parent | prev | next [-]

Transformation.

Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.

Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company

bluefirebrand 3 hours ago | parent [-]

So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed?

Because that's stupid. These laws are stupid.

brainwad 3 hours ago | parent | next [-]

That's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.

bluefirebrand 2 hours ago | parent [-]

That's stupid, these courts are stupid

It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't.

That shouldn't satisfy anyone.

immibis2 27 minutes ago | parent | next [-]

> they can produce copyrighted works, they've just been tuned so they don't.

In other words... they can't. The court is not stupid, and will consider this fact.

2 hours ago | parent | prev [-]
[deleted]
dismalaf 2 hours ago | parent | prev [-]

The point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it).

htrp 3 hours ago | parent | prev [-]

except a bunch of paywalled stuff did end up in training corpora