| |
| ▲ | usef- 3 days ago | parent [-] | | Fair, but isn't "illegal" access what they're talking about in OP? It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international. | | |
| ▲ | trhway 3 days ago | parent | next [-] | | >It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for. >It's a cost that American open models will seem to have to pay but not international. It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc. | | |
| ▲ | usef- 3 days ago | parent [-] | | > International distillers doesn't use that premium content, so they don't pay for it. This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone. > American companies and their business are protected by American laws and have access to the market protected by those laws Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going. |
| |
| ▲ | Bratmon 3 days ago | parent | prev [-] | | > It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI). Do you have a source? | | |
| ▲ | usef- 3 days ago | parent [-] | | There's been many news stories of it over the past year(s) as they signed each one. Here's the first result I could see with a rundown of many of them (am on mobile). https://digiday.com/media/a-timeline-of-the-major-deals-betw... | | |
| ▲ | Bratmon 3 days ago | parent [-] | | Those are licenses for API access to data too new to be in the training data (for use by agents), not for the training itself. I don't really understand why you think they're relevant, given that this conversation is about the training itself. | | |
|
|
|
|