| |
| ▲ | trhway 3 days ago | parent [-] | | The judge found their use is fair use. They are paying not for their use of the content, they are paying for using illegal copies of the content. The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled. To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter. | | |
| ▲ | usef- 3 days ago | parent [-] | | Fair, but isn't "illegal" access what they're talking about in OP? It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international. | | |
| ▲ | trhway 3 days ago | parent | next [-] | | >It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for. >It's a cost that American open models will seem to have to pay but not international. It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc. | | |
| ▲ | usef- 3 days ago | parent [-] | | > International distillers doesn't use that premium content, so they don't pay for it. This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone. > American companies and their business are protected by American laws and have access to the market protected by those laws Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going. |
| |
| ▲ | Bratmon 3 days ago | parent | prev [-] | | > It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI). Do you have a source? | | |
| ▲ | usef- 3 days ago | parent [-] | | There's been many news stories of it over the past year(s) as they signed each one. Here's the first result I could see with a rundown of many of them (am on mobile). https://digiday.com/media/a-timeline-of-the-major-deals-betw... | | |
| ▲ | Bratmon 3 days ago | parent [-] | | Those are licenses for API access to data too new to be in the training data (for use by agents), not for the training itself. I don't really understand why you think they're relevant, given that this conversation is about the training itself. | | |
|
|
|
|
|
| |
| ▲ | giaour 3 days ago | parent | next [-] | | > I doubt these will have worse protection than software does, which has far better protections than copyright Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable. | | |
| ▲ | trhway 3 days ago | parent | next [-] | | Distillation isn't a copy. Distillation is more akin to "clean room" implementation. Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training. By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that. There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like. | |
| ▲ | breppp 3 days ago | parent | prev [-] | | Software is protected by the DMCA, patents, licenses, EULAs, all of those aren't there for books. I doubt new laws won't be written for model outputs. Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply | | |
| ▲ | bigiain 3 days ago | parent | next [-] | | The C in DMCA stand for Copyright. All (I think?) software licenses are underpinned and made legally enforceable by copyrights. EULAs are underpinned by licenses which are founded on copyright. Patents are the only one of those protections that are not based on copyright, and there are lots of very good arguments against at least most software patents (all software patents of the form "Do {well known and obvious thing} with a computer" should, in my opinion, be immediately revoked and potentially have every company who's enforced payments from such patents investigated for fraud). | |
| ▲ | giaour 3 days ago | parent | prev | next [-] | | You may recall that the DMCA was originally written to protect music and movies. It does in fact apply to creative works. If you have ever purchased an MP3, eBook, or streaming movie, you will also be aware that you purchased a license to the underlying IP. This is also true of physical media, but the license agreement you have to accept when obtaining a digital work makes this explicit. I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article. | | |
| ▲ | vel0city 3 days ago | parent [-] | | You do patent ideas, but the actual words written in a book describing that idea would only be protected by copyright at best. FWIW, the exact words describing the idea being patented are technically public domain; that's the whole point. You're free to go look up that patent, print it out, make whatever copies of it you want. Take any of the drawings in patents, put them on t-shirts, and sell them. No problem. Implementing the ideas those words represent is a different story. For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea. |
| |
| ▲ | wasfgwp 2 days ago | parent | prev | next [-] | | How is “model output distillation” different to using outputs (which are legally copyrightable) for any other purpose? | |
| ▲ | queenkjuul 3 days ago | parent | prev [-] | | Afaik (and ianal) there's nothing stopping anyone from attaching a EULA to a physical book |
|
| |
| ▲ | vel0city 3 days ago | parent | prev | next [-] | | Let's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP. Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen? If the model output is owned by the person prompting it and paying for the tokens, what's the problem here? If the model output is owned by the trainer of the model, that's a big nasty can of worms. | |
| ▲ | preg_match 3 days ago | parent | prev | next [-] | | Why would this be the case. Why would software output from a model magically have greater protection than the software the model trained on. | |
| ▲ | wasfgwp 2 days ago | parent | prev [-] | | LLM outputs are not copyrightable. At least that’s the current established legal precedent in the US. The only question is whether the user owns the copyright without significantly transforming the output but that’s not really relevant in those specific situation. I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models.. |
|