| ▲ | mrngld 2 days ago | |||||||||||||||||||||||||||||||||||||||||||
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court hasn't said so. The open web is open, after all. And they don't seem to be stealing books, they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book. They also pay big bucks for commercially curated data and training sets. Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead. | ||||||||||||||||||||||||||||||||||||||||||||
| ▲ | rsstack 2 days ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||||||||
> they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book There’s still the open question on learn vs copy/mimic/repeat. As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission. IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees. | ||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jchw 2 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||
There really isn't any moral argument against distillation, which is itself pretty goddamn benign. It's pretty simple. Someone pays for Claude access. Claude outputs tokens that are not copyrighted. Then you train on those tokens, which doesn't create a derivative work in the first place. Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped. Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden. Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information. I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials. | ||||||||||||||||||||||||||||||||||||||||||||
| ▲ | riedel 2 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||
While many things might be legal (or hasn't been ruled clearly illegal yet) /under different jurisdictions, we are still in the process of figuring out what we accept as ethical. As the output of models can't be easily copyrighted, destillation is equally disputed. Particularly if the primary model interaction was not destillation (as in this case) IMHO it will be legally quite difficult to restrict secondary use for training. In the end we have to find a legislation and probably even international treaties that account for the fact that classical copyright is beginning to become an obsolete concept. | ||||||||||||||||||||||||||||||||||||||||||||
| ▲ | esafak 2 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||
If American AI labs can learn from the open web why can't Chinese AI labs learn from American ones?? The argument is that learning and distillation are transformative and legal, is it not?? What's good for the goose is good for the gander. | ||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||
| ▲ | wat10000 2 days ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||||||||
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders. But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs? | ||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||