| ▲ | trompetenaccoun 39 minutes ago |
| It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well. I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI? |
|
| ▲ | qarl 36 minutes ago | parent | next [-] |
| > the largest copyright theft operation in human history Many people think that it was fair use: training is akin to reading, not copying. Especially the courts. |
| |
| ▲ | Trusteando 8 minutes ago | parent | next [-] | | No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights, and they obtain profit from them. Noone can compete which copying and pasting (in encoding from) from copyright protected material. | | |
| ▲ | qarl 5 minutes ago | parent [-] | | > No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights Exactly analogous to a human reading the material. |
| |
| ▲ | trompetenaccoun 26 minutes ago | parent | prev [-] | | That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands? The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material. | | |
| ▲ | qarl 20 minutes ago | parent [-] | | > That's not been legally established 100% of the rulings agree with me. The piracy is not in question. It is unarguably copyright violation. But that's not what anyone means in this context. Training is what everyone means. > The law is the law, there can't be different law for corporations with billions in backing. I didn't say otherwise. That's a straw man. | | |
| ▲ | trompetenaccoun 15 minutes ago | parent [-] | | So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here? | | |
| ▲ | qarl 9 minutes ago | parent [-] | | > So we agree they have violated copyright at a much larger scale than LibGen No. |
|
|
|
|
|
| ▲ | TeMPOraL 32 minutes ago | parent | prev [-] |
| They didn't believe it was sketchy. They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that. Judging by how AI threads look like for the past year, they were absolutely right to be worried. > largest copyright theft operation in human history In fact, you're doing exactly that right here. |
| |
| ▲ | probably_wrong 17 minutes ago | parent | next [-] | | The comment you're replying to is citing almost verbatim [1] Microsoft’s director of Applied Science, Brent Hecht, who called OpenAI's data collection practices "the largest theft of labor in human history" in an internal memo. [1] https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-s... | | | |
| ▲ | discreteevent 20 minutes ago | parent | prev | next [-] | | > frame it in a dumb, manipulative way like that
> In fact, you're doing exactly that right here. You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong. Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that? | |
| ▲ | latexr 17 minutes ago | parent | prev [-] | | > They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that. You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not. Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve on libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from. |
|