| ▲ | peri-cl 2 days ago |
| Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference. |
|
| ▲ | gruez 2 days ago | parent | next [-] |
| >Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever. >and continue doing it Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them. |
| |
| ▲ | bix6 2 days ago | parent | next [-] | | > From the start, Anthropic “ha[d] many places from which” it could have purchased
books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and
chief executive officer Dario Amodei put it (see Opp. Exh. 27). https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz... Sounds pretty pro piracy to me. | |
| ▲ | wolvoleo 2 days ago | parent | prev | next [-] | | They stopped because they already had everything in their model anyway. | |
| ▲ | austhrow743 2 days ago | parent | prev | next [-] | | Yes there are people who are both for and against copyright. | |
| ▲ | sam_lowry_ 2 days ago | parent | prev [-] | | [flagged] | | |
| ▲ | yonran 2 days ago | parent | next [-] | | > This is an unethical as a company may behave, short of killing people. This is hysterical. No, format shifting old unwanted books is not unethical. The books still exist, in an internal digital library. If copyright law were to change to allow sharing orphan works some day, Anthropic could share them. But under current law, the books are preserved digitally and used for transformative uses that all Claude users benefit from. | |
| ▲ | empiricus 2 days ago | parent | prev [-] | | did not follow all the details, but my understanding is that some form of copyright law nudges in the direction of destroy after scan? |
|
|
|
| ▲ | spwa4 2 days ago | parent | prev | next [-] |
| In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once. One wonders if they're still doing it. |
| |
| ▲ | ungut 2 days ago | parent | next [-] | | OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models.
There is just not enough non-copyrighted data out there. | | |
| ▲ | yorwba 2 days ago | parent | next [-] | | You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models
without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation. | | |
| ▲ | spwa4 2 days ago | parent | next [-] | | I'm pretty sure we would know if they did that. And we don't. Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit! | | |
| ▲ | yorwba 2 days ago | parent [-] | | AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742 And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway. | | |
| ▲ | spwa4 2 days ago | parent [-] | | But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course. | | |
| ▲ | yorwba 2 days ago | parent [-] | | GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection. | | |
| ▲ | spwa4 2 days ago | parent [-] | | You mean very likely the Anna's archive torrent dump because it's MUCH better quality than the general internet and beyond a certain amount of input data (which is a lot, but much less than the internet) the only thing that matters in training is the quality of the data, to the point that now many labs have thousands of people just making and improving essentially school exercises full time? Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad. | | |
| ▲ | yorwba a day ago | parent [-] | | No, I mean the WebText corpus whose construction from 45 million Reddit post with at least 3 karma is described in section 2.1 of the PDF I linked. They did remove all Wikipedia documents. Anna's Archive didn't exist in 2019. |
|
|
|
|
| |
| ▲ | tygon 2 days ago | parent | prev | next [-] | | I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate. | | |
| ▲ | yorwba 2 days ago | parent | next [-] | | There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly. | | |
| ▲ | tygon 2 days ago | parent [-] | | This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way. It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude. |
| |
| ▲ | ungut a day ago | parent | prev | next [-] | | According to the US Chamber of commerce the US exports over 270 billion dollars.
https://www.uschamber.com/intellectual-property/unlocking-cr... | |
| ▲ | ungut 2 days ago | parent | prev [-] | | Doesn't really matter. The incentive structure to steal clearly exists, so why would they even go through the trouble? |
| |
| ▲ | ungut 2 days ago | parent | prev [-] | | Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place. I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days. | | |
| ▲ | yorwba 2 days ago | parent [-] | | I'm not aware of any successful accusations against OpenAI for illegally obtaining copyrighted material, in contrast to Anthropic, who settled for $3000 per work and then still had to buy legal copies to keep using them (likely for much less). Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy. | | |
| ▲ | ungut a day ago | parent [-] | | You are kind of right, but you also did not look very hard. They deleted huge datasets in anticipation of lawsuits, at least that much is known.
Of course plaintiffs were unable to depose their internal lawyers (who apparently know why they were frantically deleted) due to 'attourney client priviledge' further refusing to provide any kind of transparency. But yeah, I guess they are they are better at covering their tracks and destroying evidence. Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them. I think the real issues here are two-fold: Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment. Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man. |
|
|
| |
| ▲ | deadbunny 2 days ago | parent | prev [-] | | What tosh. It's copywrited material, they pay to access it like everyone else. |
| |
| ▲ | outside1234 2 days ago | parent | prev | next [-] | | Of course they are. They have just put on their Swiss Banker suit now and have all sorts of deflection techniques in place such that, of course, "the money has the stamps that says its clean" (when it it really blood money hidden behind a pretty wall). | |
| ▲ | IncreasePosts 2 days ago | parent | prev [-] | | I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fine |
|
|
| ▲ | hintymad 2 days ago | parent | prev | next [-] |
| > they can afford the penalties and continue doing it. I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag. |
| |
| ▲ | Gud 2 days ago | parent [-] | | That's what they're doing now, when they are established. But when it was a proof of concept, they were using pirated data. Just like Spotify did. | | |
| ▲ | treszkai 6 hours ago | parent [-] | | Sorry I'm out of the loop: how did Spotify use AA or other pirated data? |
|
|
|
| ▲ | vaylian 2 days ago | parent | prev [-] |
| > Just like they want it to be illegal to run local ML inference. Citation? |
| |
| ▲ | peri-cl 2 days ago | parent | next [-] | | https://news.ycombinator.com/item?id=49076057 ("Our position on open-weights models (anthropic.com)", 1812 comments) | | |
| ▲ | a day ago | parent | next [-] | | [deleted] | |
| ▲ | vaylian a day ago | parent | prev [-] | | I see. The following two statements (headlines) stand out: * We should not sell powerful chips or chipmaking equipment to China * We should crack down on industrial-scale distillation operations That's not quite a ban of local ML inference, but it basically says that he doesn't want companies in China to create their own state of the art models. | | |
| ▲ | peri-cl a day ago | parent [-] | | No; it's the other headline, mandatory US government certification, that would ban Americans from running their own Chinese large models. |
|
| |
| ▲ | bookofjoe 2 days ago | parent | prev [-] | | Over the past decade I've noticed on HN the following order of frequency in choice of words, most common to least: 1. Citation 2. Source 3. Reference Long ago in a career based on original research, I/we ONLY used "reference." | | |
| ▲ | tygon 2 days ago | parent | next [-] | | While it is definitely over a decade at this point (over two in fact), some of this likely comes from the term [citation needed], that originated on Wikipedia, as a cynical backhanded response to unsourced claims. It has become a catch-all. Language and how it evolves is a pretty interesting subject. | | |
| ▲ | bookofjoe 2 days ago | parent [-] | | You're right. Has to be of Wikipedia use origin. Thanks! | | |
| ▲ | vaylian a day ago | parent [-] | | OP here. You came to the right conclusion. This is inspired by Wikipedia. | | |
| ▲ | genxy 14 hours ago | parent [-] | | Pretty sure (citation needed) predated wikipedia. | | |
| ▲ | bookofjoe an hour ago | parent [-] | | Perhaps, but I'm betting "reference" significantly antedated the Wikipedia-induced ubiquity of "citation [needed]." |
|
|
|
| |
| ▲ | vaylian a day ago | parent | prev [-] | | I could have written more words, but they would have not conferred more meaning. |
|
|