| ▲ | zqwt3k 4 hours ago |
| Thanks Elon, now we know that scraping is illegal! Very good to clarify that for future proceedings against the AI thieves. |
|
| ▲ | immibis2 3 hours ago | parent | next [-] |
| Everything is both legal and illegal until a lawsuit happens. Then it collapses depending a little bit on the facts and mostly on who has the better lawyers. I suspect Nitter's first round with lawyers pointed out that scraping is legal, but now they have been threatened with something else than scraping - Elon claims something else the way Nitter runs is illegal, such as the use of fake accounts to circumvent an access control device (DMCA 1201). |
| |
| ▲ | dannyw 3 hours ago | parent | next [-] | | At the end of the day, as an individual or a team or a company, regardless of the statue and case law, you have to perform the calculus on your monetary and legal resources versus your counterparty. Obviously ungrounded and frivolous cases tend to be easier to defend asymmetrically, but if I was X's legal team, there's no shortage of semi-plauisble claims I could throw at the wall and see what sticks. As an example of this imbalance in action, BrightData is a 'gray area' company that basically does this exact kind of scraping. They have somehow won against Meta Platforms suing them, and even got X's lawsuit against them for scraping -- identical (?) activity to XCancel -- dismissed. According to Wikipedia: > In May 2024, a federal judge dismissed the suit, ruling that Bright Data did not violate X's terms of service or copyright by scraping publicly accessible data.[24] The judge emphasized that such scraping practices are generally legal and that restricting them could lead to information monopolies But does XCancel have the resources of a company like Bright Data, that's funded and used by companies like Deloitte and Moodys? | | |
| ▲ | tehwebguy 2 hours ago | parent | next [-] | | If the name Bright Data is ringing a bell to anyone, it’s probably because they are a (the?) primary offender running the LG TV “residential proxy” (botnet) | | |
| ▲ | immibis2 an hour ago | parent | next [-] | | Which, btw, was the least bad thing about those TVs. I support residential proxying as a means to liberate public data that's being held captive by people like Elon. | |
| ▲ | ranger_danger an hour ago | parent | prev [-] | | And the company's previous name was, ominously... Luminati. |
| |
| ▲ | immibis2 3 hours ago | parent | prev [-] | | Notice it says "terms of service or copyright". If X's lawyers have any intelligence, they'll have a reason why XCancel is not identical to Bright Data. Perhaps this time, instead of claiming it's a copyright violation, they'll claim it's wire fraud because multiple accounts are used. |
| |
| ▲ | Klonoar 2 hours ago | parent | prev [-] | | Scraping is legal, but IIRC scraping authenticated contents is not. Isn't Nitter abusing account sign-in for this? | | |
| ▲ | CamperBob2 4 minutes ago | parent | next [-] | | To fix this: 1) Nitter offers a deal: download our browser extension, sign up for Twitter, and we'll give you some kind of perk (Amazon credits, whatever). 2) Browser extension surfs Twitter on the user's behalf, scraping and sending copies to Nitter. 3) We find out how serious the legal system is about prosecuting scraping. | |
| ▲ | immibis2 an hour ago | parent | prev | next [-] | | Anyone can make an account, though, and instantly access that content, so Elon doesn't really have any leg to stand on by claiming they're private. It's not the same as Cambridge Analytica scraping stuff you had to have certain privileges to see by tricking the system into granting those privileges. | | |
| ▲ | Klonoar 16 minutes ago | parent [-] | | Elon unfortunately may have a leg to stand on, because it doesn't functionally matter whether "anyone can make an account and see it". I don't believe any legal ruling to date has been fine with that distinction. |
| |
| ▲ | ktm5j 2 hours ago | parent | prev | next [-] | | Yes, that's what they went on to point out. | |
| ▲ | ranger_danger an hour ago | parent | prev [-] | | Aurora Store does the same thing to facilitate downloading Play Store apps. |
|
|
|
| ▲ | rich_sasha 4 hours ago | parent | prev | next [-] |
| I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”. Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled. |
| |
| ▲ | brainwad 4 hours ago | parent | next [-] | | Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content. | | |
| ▲ | Planktonne 3 hours ago | parent | next [-] | | > Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content. It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources. I'm not convinced that a competitive use at one remove should be treated as not competitive. | | |
| ▲ | brainwad 3 hours ago | parent [-] | | I think that's more true in image generation than in text? At least, all the money is in LLMs that write code, not LLMs that write O'Reilley-style coding books. | | |
| ▲ | wongarsu 2 hours ago | parent | next [-] | | If you have a websites that offers guides, how-tos or tutorials, LLMs directly compete with you. StackOverflow would also have a really good case After all LLMs don't just code, they also answer questions and give step-by-step instructions. In terms of total userbase those features are used a lot more than writing code | |
| ▲ | _flux 3 hours ago | parent | prev [-] | | I thought LLM-authored books were rampant in Amazon, though? |
|
| |
| ▲ | nixpulvis 3 hours ago | parent | prev | next [-] | | I think LLMs providers pretty directly compete with content they scrape like Wikipedia and SO... | | |
| ▲ | bayindirh 3 hours ago | parent | next [-] | | Remember the golden rule of the golden rules: Who has the gold makes the rules. | | |
| ▲ | atemerev 3 hours ago | parent [-] | | If a society operates under a rule like this, it is no better than Russia or any other tyranny where law is for me but not for thee. This is not how it should work in a supposedly free and lawful country. | | |
| |
| ▲ | bluefirebrand 3 hours ago | parent | prev [-] | | No kidding. I dunno what kind of traffic loss Wikipedia has had but SO is really dead these days If they weren't competing with AI then why is AI killing it? | | |
| ▲ | brainwad 3 hours ago | parent [-] | | There's a difference between creating a market for something better, so that nobody wants the old thing, and competing _in_ the market for the old thing by copying it directly. | | |
| ▲ | immibis2 32 minutes ago | parent | next [-] | | If two things are competing, they are in the same market by definition. | |
| ▲ | wongarsu 2 hours ago | parent | prev | next [-] | | If LLMs only made SO redundant by writing code and solving my technical problems autonomously so I never have to think about it, I would agree. But often I do ask LLMs technical questions, and they answer in great detail. And that part is a very direct SO competitor | |
| ▲ | darkwater 3 hours ago | parent | prev | next [-] | | And what would be a read-only version of X like XCancel compete against, exactly? Ads impressions? That would be the only possible thing yet they don't add any ads. | | |
| ▲ | brainwad 3 hours ago | parent [-] | | It's depriving X of impressions that they could monetise, no? Xcancel doesn't have to make money itself, it just has to impair the rights of the copyright holder. Otherwise piracy would also be legal as long as it were non-profit... | | |
| ▲ | darkwater 2 hours ago | parent [-] | | > Otherwise piracy would also be legal as long as it were non-profit... Which is in a few jurisdictions, or at least is not prosecuted if it's for personal use.
Also, according to your definition, the creator of uBlock Origin or any other adblock system should be sued in the same way, because they are depriving $ADS_CORP of their precious impressions. | | |
| ▲ | brainwad 2 hours ago | parent [-] | | Well, adblockers don't copy the copyrighted content. They just control how it's rendered on the user's machine. Copyright cares about making copies and especially distributing them. | | |
| ▲ | darkwater 2 hours ago | parent [-] | | You have a point on this, I recognize, but it still seems a very thin line to walk (for X) - at least morally, because I don't think they are actually loosing real big money to anyone. |
|
|
|
| |
| ▲ | bluefirebrand 2 hours ago | parent | prev [-] | | I don't really buy this. It's like saying "toaster oven/air fryer combos" don't actually compete with toaster ovens or air fryers because they are creating a market for something better Of course they complete. Toaster ovens compete with toasters. Microwaves compete with toaster ovens. Just because it's not the exact same product doesn't mean it's not competing Would I be allowed to steal LG's designs for a microwave and make a "superwave" that does laundry and heats food? Would you claim those products don't compete because the superwave is "something better"? | | |
| ▲ | brainwad 2 hours ago | parent [-] | | It's not illegal to write a similar book or song to an existing one. Copyright only protects existing works from literal copying (possibly in part). | | |
|
|
|
| |
| ▲ | bonsai_spool 3 hours ago | parent | prev | next [-] | | > Precedent is pretty clear: What cases are you citing when you say this? | | |
| ▲ | brainwad 3 hours ago | parent | next [-] | | Bartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM. | | |
| ▲ | lesuorac 3 hours ago | parent | next [-] | | Bartz is an author though. Is X claiming ownership of the posts people make because pretty much every single social media site doesn't so they have section 230 protection. | | |
| ▲ | brainwad 3 hours ago | parent [-] | | They can just round up some friendly users and sue under their names. Starting with their own corporate accounts? | | |
| ▲ | lesuorac 8 minutes ago | parent | next [-] | | That would definitely limit damages to strictly those accounts. I'm not even sure he can use his own account as one of them. The SEC might be pretty friendly to him but I'm not sure that limiting access to a location where material information about Tesla/SpaceX is provided won't become a problem. But I'm not even sure what damages the accounts are suffering as revenue sharing is going away [1]. With Bartz the damage is a loss of sale. With X the damage is $0 per post to the poster. There is a newer Original Content Rewards program [2] but it seems to split revenue from X Premium and presumably people that have X Premium are not using XCancel so the damages would be 0. [1]: https://help.x.com/en/using-x/creator-revenue-sharing [2]: https://help.x.com/en/using-x/original-content-rewards | |
| ▲ | bonsai_spool an hour ago | parent | prev [-] | | These statements would suggest that the precedent is not, in fact, clear |
|
| |
| ▲ | 3 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | immibis2 3 hours ago | parent | prev [-] | | Perhaps https://en.wikipedia.org/wiki/Warner_Bros._Entertainment_Inc.... |
| |
| ▲ | ratelimitsteve 2 hours ago | parent | prev [-] | | How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from) | | |
| ▲ | bonsai_spool an hour ago | parent [-] | | > How clear is it when I google a recipe and get an AI-generated recipe Recipes can’t be copyrighted Here’s one discussion about this https://www.nycbar.org/reports/secret-ingredients-how-to-pro... | | |
| ▲ | ratelimitsteve an hour ago | parent [-] | | fair, but recipes aren't the only thing where AI summaries at the top of search results are simultaneously transformative and competitive. Any information that is scraped from a website and then summarized by AI in a way that prevents that website from getting views is both transformative and competitive. |
|
|
| |
| ▲ | embedding-shape 4 hours ago | parent | prev | next [-] | | Transformation. Taking something someone else made and showing it as-is, bypassing their own restrictions: No no. Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company | | |
| ▲ | bluefirebrand 3 hours ago | parent [-] | | So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed? Because that's stupid. These laws are stupid. | | |
| ▲ | brainwad 3 hours ago | parent | next [-] | | That's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works. | | |
| ▲ | bluefirebrand 2 hours ago | parent [-] | | That's stupid, these courts are stupid It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't. That shouldn't satisfy anyone. | | |
| ▲ | immibis2 27 minutes ago | parent | next [-] | | > they can produce copyrighted works, they've just been tuned so they don't. In other words... they can't. The court is not stupid, and will consider this fact. | |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
|
| |
| ▲ | dismalaf 2 hours ago | parent | prev [-] | | The point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it). |
|
| |
| ▲ | htrp 3 hours ago | parent | prev [-] | | except a bunch of paywalled stuff did end up in training corpora |
|
|
| ▲ | dismalaf 3 hours ago | parent | prev | next [-] |
| Pretty sure it's legal when you do it for your own use (same as browsing a website) but it's illegal to redistribute web scraped results. |
| |
| ▲ | burnte 2 hours ago | parent [-] | | It's not that cut and dry or else search engines wouldn't be legal. It depends on how much is used, for what context, etc. This very well may wind up being fair use. | | |
| ▲ | dismalaf 2 hours ago | parent [-] | | Search engines "modify" it ie. show snippets + direct to the actual site. In general fair use pretty much always requires it to be transformative and/or point to the source. Simply scraping it to prevent people from going to X isn't free use in any definition I've heard. | | |
| ▲ | gosub100 2 hours ago | parent [-] | | Remember ~10 years ago when Google had "cached" versions of the websites? I believe they removed that feature for this reason. |
|
|
|
|
| ▲ | gosub100 2 hours ago | parent | prev | next [-] |
| Now I want to know, what happens if you redistribute an "AI summary" of the copyright material? |
|
| ▲ | blitzar 4 hours ago | parent | prev | next [-] |
| Thats "Billionaire Use" -- its like a "Fair Use" exemption but for billionaires. |
| |
|
| ▲ | znpy 3 hours ago | parent | prev | next [-] |
| scraping content is mostly legal, redistributing content is not. if you started doing the same to, say, instagram content both meta and individual creators would sue you as well. sites like archive.ph are in a similar bucket btw, and yet nobody's complaining (except websites seeing people evading their paywall). but at the end of the day it's not really fair to apply laws differentially on the basis of whose political ideas we like more. |
| |
| ▲ | immibis2 26 minutes ago | parent [-] | | Meta has no exclusive rights to the content on Instagram, and X has no exclusive rights to the content on X. They have a non-exclusive license to republish it, etc. |
|
|
| ▲ | wahnfrieden 4 hours ago | parent | prev [-] |
| It’s the rehosting of the content not the scraping Edit: not a moral stance |
| |
| ▲ | jm4 4 hours ago | parent | next [-] | | I love how the content belongs to them when someone else reposts it but it belongs to the user if the content is illegal. Such a double standard with these social media and AI companies. Why do we put up with it? | | |
| ▲ | embedding-shape 4 hours ago | parent | next [-] | | > Why do we put up with it? You let it happen. Once people stop letting it happen, it'll stop. But social media is apparently the new "opium of the masses" so here we are and no one wants to do anything. | | |
| ▲ | flaburgan 3 hours ago | parent [-] | | I actually do want to do something. I started to scrape Twitter to make it freely available. The web should stay open. | | |
| ▲ | immibis2 23 minutes ago | parent [-] | | I am also doing something, by hosting one of the public instances of Nitter. Nitter is useful for sporadic random access to tweets, but for public feeds like municipal authorities etc. it would be useful if someone scraped the feed and re-hosted the feed from their own server, without being hobbled by rate limits. Is that what you're doing? |
|
| |
| ▲ | simianparrot 4 hours ago | parent | prev [-] | | Did the users of X consent to XCancel copying their posts to their servers? | | |
| ▲ | conception 4 hours ago | parent | next [-] | | Everyone who goes to twitter copies the posts to their computers. | | |
| ▲ | 27183 3 hours ago | parent [-] | | Spot on. This is where a lot of these "terms and conditions" break down logically. Viewing some content on the internet is literally copying it. So is the distinction that xcancel served the content? But when I run mtr xcancel.com
I see a bunch of hops between me and them. Every one of those hops is literally copying and retransmitting all the content. Are they not also serving it? | | |
| ▲ | immibis2 3 hours ago | parent [-] | | No, this is where programmers rules-lawyer in ways that actual lawyers don't and then get law stuff hilariously wrong. No judge thinks that viewing an HTML page is downloading it, because downloading means saving a copy to your computer, not just looking at it. Even having an internet cache folder doesn't count as downloading. Even copying the file from the internet cache folder to somewhere might not count as downloading, although it'd still be a copy. Same as when LG said their TVs don't record you and then Hacker News said "how can they detect voice commands if they don't record your voice"... facepalm. | | |
| ▲ | 27183 3 hours ago | parent [-] | | I don't pretend to understand law, mostly it just doesn't make sense at all. | | |
| ▲ | immibis2 3 hours ago | parent | next [-] | | It makes more sense when you remember it's not a computer program and the things that are written in the law are not the things that will actually happen in the way that "if(foo) bar;" makes bar happen if foo is true. It's more like a book of excuses you could use for why you didn't do your homework. Then the other side also has to bring an excuse for why you were supposed to do it, and if the principal thinks their excuse is better than yours, you get detention. If you tell the principal "I don't have to do my homework because work means employment and it's illegal to employ a minor" you'll get detention for not doing your homework and extra detention for being a smartass. | | |
| ▲ | rimunroe 3 hours ago | parent [-] | | And this example is not just due to people not taking the trouble to write fully specified rules. I don't think such rules could even be written. You can just do your best to cover the cases you can think of. The complexity of society is incomprehensibly vast and constantly changing, and the law has to have wiggle room to account for it. |
| |
| ▲ | rimunroe 3 hours ago | parent | prev [-] | | Could you elaborate in what way you find the law mostly doesn't make sense? It has to be flexible in order to work with actual humans. Why should visiting a page on your computer count as copying? Usually when we talk about copying it's someone making a duplicate so it can be accessed later. Only a very technical user is going to be diving into their cache to view that content after the fact. The vast majority of people don't understand that the browser is storing anything on their computer, much less how to access it before it's purged. | | |
| ▲ | sekh60 3 hours ago | parent | next [-] | | I can't remember the court case, but Blizzard did argue and win in court that WoW Glider's producers violated copyright law. If I recall correctly violating the TOS meant that an unauthorized copy made by executing the file chasing it to load WoW into RAM was created. | | |
| ▲ | rimunroe an hour ago | parent [-] | | It looks like that was MDY Industries, LLC v. Blizzard Entertainment, Inc., which relied on MAI Systems Corp. v. Peak Computer, Inc. for the relevant part of the ruling. The person I was responding to was saying that anytime you viewed copyrighted content with a browser you’d necessarily be committing copyright infringement. I’m not a lawyer but I can imagine that the reasoning there would be slightly different from someone simply viewing a post in a browser as part of the intended use of the site. | | |
| ▲ | sekh60 33 minutes ago | parent [-] | | Oh yeah, I understood your point, but given MDY Industries, LLC v .Blizzard who knows what the "right" judge would rule? With IP laws these days we're really getting into weird places. |
|
| |
| ▲ | 27183 3 hours ago | parent | prev [-] | | > Why should visiting a page on your computer count as copying? Because there's no physical mechanism for the information to be transmitted over a computer network other than by copying the bytes. Note this is distinct from broadcast systems like analog television or radio. Packet switching networks only function by copying information and storing multiple copies around the internet, including in your computer's RAM (and disk, if cached). So a legal definition that says "this kind of copying is copying but that other kind of copying isn't copying" makes no sense at all. Like many other legal definitions--it's all about what has been successfully snuck past a jury at one point or another in the past, without any heed for how things actually work. | | |
| ▲ | amiga386 2 hours ago | parent | next [-] | | It's not about "how things actually work", the law is there to regulate human activity. The law tends to call these copies on the wire, in RAM, in caches, etc. "transient copies", which is fine until a human starts using them as non-transient copies, e.g. saves them for later. You could argue that your MP3 of Enjoy the Silence is actually just a big number, and you can XOR it with 0xFF and it's a completely different big number, and you just happen to XOR it with 0xFF when you want to listen to it. The courts would look past that, and instead determine if you created that "big number" by MP3-encoding the track from a CD you owned (legal), versus obtaining it from some file-sharing network (not legal) Classic essay about techies not understanding the law: What Colour Are Your Bits? https://ansuz.sooke.bc.ca/entry/23 | |
| ▲ | rimunroe 3 hours ago | parent | prev [-] | | > Because there's no physical mechanism for the information to be transmitted over a computer network other than by copying the bytes. Your response seems to ignore everything in my comment other than the second sentence. I was asking why that detail should matter as far as the law is concerned, and I gave some reasons I don't think that would be good or practical. | | |
| ▲ | card_zero 2 hours ago | parent [-] | | There's the matter of linking to copyrighted works: https://en.wikipedia.org/wiki/Copyright_aspects_of_hyperlink... If your link is set up to make the image display immediately (that is, you wrap it in image tags, or as in one case, embed Instagram posts) then you may be violating copyright. What's more, in Europe, just a hyperlink to a copyrighted work violates copyright. Conclusion: copyright is not about copying, it's about access. | | |
| ▲ | rimunroe 2 hours ago | parent [-] | | Sure, but that seems different from what I was addressing. The person I was responding to was saying that the law as a whole usually doesn’t make sense. They were saying that in the context of arguing that if the law didn’t consider viewing a page of copyrighted copying as involving copying due to the technical basis of it having to transfer bits to your computer then the law didn’t make sense. My point was that laws don’t have to encompass or fully specify all edge cases, and that the ways laws are written can be open to interpretation. I think I removed a sentence before posting about the purpose of finders of facts in the US system like juries or judges in bench trials. |
|
|
|
|
|
|
|
| |
| ▲ | embedding-shape 4 hours ago | parent | prev | next [-] | | I think users implicitly consent to their public data being public data when they put it in public together with other public data. How I view that public data they decided to make public data, is none of their business. | |
| ▲ | bakies 4 hours ago | parent | prev | next [-] | | i just screenshotted this comment without your consent | |
| ▲ | ranger_danger 37 minutes ago | parent | prev | next [-] | | How do you know data is copied to their servers? > If you serve as a mere conduit for automatic transmission of user communications, there are no other qualifications or obligations you need to meet. If you serve a caching function, in addition to the two requirements above, you must maintain comply with the notice-and-takedown process. https://www.copyright.gov/512/ https://internetcases.com/2024/02/12/dmca-subpoena-to-mere-c... | |
| ▲ | roosterIllusi0n 4 hours ago | parent | prev [-] | | Why can't they distill twitter when AI companies distill everything including twitter? Distilling is copying and redistributing. | | |
| ▲ | akerl_ an hour ago | parent | next [-] | | Are they distilling? Distilling isn’t copying and redistributing, for the same reason that you reading a story and then writing your own story based on ideas you learned is different from you reading a book, writing all the words down verbatim, and then publishing it as your own. | | |
| ▲ | immibis2 20 minutes ago | parent [-] | | That's not what distilling is either. Distilling is training your AI to exactly copy someone else's AI. |
| |
| ▲ | wahnfrieden 2 hours ago | parent | prev [-] | | No it’s not |
|
|
| |
| ▲ | zqwt3k 4 hours ago | parent | prev | next [-] | | That is what all LLMs could do in 2023, verbatim, before they trained it out of them in order to keep up the pretense that there is no plagiarism. Now they all obfuscate the original or refuse to cite. | |
| ▲ | echelon 4 hours ago | parent | prev | next [-] | | Grok can do that if you give it an HN thread or other websites. | |
| ▲ | HumblyTossed 4 hours ago | parent | prev [-] | | There's a button on X that allows me to repost someone else's content. | | |
| ▲ | bel8 4 hours ago | parent | next [-] | | there's a setting to disable that | |
| ▲ | wahnfrieden 4 hours ago | parent | prev [-] | | That’s not rehosting Edit: iPhone autocorrected my OP which meant to say rehosting not reposting |
|
|