Remix.run Logo
kelnos 6 hours ago

I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.

I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.

I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.

torginus 5 hours ago | parent | next [-]

With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.

Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.

This could mean every potential serious customer would have no option but to seek alternatives to these online services.

stymaar 5 hours ago | parent | prev | next [-]

This. Distillation “attacks” are a made up concept. It's as if I claimed that Anthropic made a “training attack” when training on my internet writing.

giancarlostoro 5 hours ago | parent | prev | next [-]

Abolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain.

LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.

TZubiri 3 minutes ago | parent | prev | next [-]

>I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.

With what knowledge are you claiming this? If it turns out companies are using IP proxy networks would you change your mind?

What if the IP Proxy networks were used by criminals for similar attacks like DDoS or plain cyber attacks?

What if the source of the IP proxy networks were residential addresses to avoid detection?

What if the way these IPs were acquired were through pwned devices?

What if the credit cards used do not identify the company that carries the attack? What if they use the employee's personal credit cards? What if it's family members of employees? What if it's a network of personal credit cards where cc owners get a payment for making a purchase on their name? What if they are stolen ccs?

Not just a hypothetical btw, I believe almost all of these are true.

godwinson__4-8 5 hours ago | parent | prev | next [-]

If the leading private labs attempt to use the government to pull up the ladder under the pretense of "safety" then the response of the people should be to take such questions out of private hands and nationalize the leading labs.

Or they could abide by the precedents they set and learn to compete. They shouldn't be allowed to have it both ways.

soundworlds an hour ago | parent | prev | next [-]

100% - the work came from the people, it should go back into the hands of the people.

I also think if Anthropic and OpenAI had been releasing Open models along the way, people wouldn't be nearly as suspicious of them.

Aurornis 5 hours ago | parent | prev | next [-]

There is nothing illegal about training on traces from frontier models.

However the frontier labs don’t have to serve customers who are farming the service for distillation purposes. That’s their choice and they’re free to make it if they detect distillation happening.

throwawayk7h an hour ago | parent | prev | next [-]

"Strip-mine" is not correct. The commons are all still there and you can still train on them just like the frontier labs did. Of course, it may be illegal to do so, but that's not any different than before.

larodi 32 minutes ago | parent | prev | next [-]

Given (A) :

> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models. They famously ingested plenty of copyrighted material without the permission of those intellectual property holders.

And many people's shared opinion (B):

>> I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.

...

It is very difficult to actually say NO to the fact that (A) was done, which then leads logically to conclusions as (B). But also we should remember that if these two hold (and (A) is an axiom more or less now), then it comes as no surprise that then also all opensource licensing is immediately rendered void and null, as keeping it would contradict (A) and would go against the very common and consequential logic in (B).

Copyright is so dead. And it was not me killing it with a cynical post on HN. Dunno why so many people still fail to face it. There is no way it can exist in its current form, because then immediately (A) happens and (B) follows.

pj_mukh 2 hours ago | parent | prev | next [-]

I wonder if along with “Pacing the frontier”, we can get the frontier labs to Share the raw data.

I’m sure the labs claim that their real innovation is in the RLHF, training and architecture. Keep that and just share the raw data somewhere.

impossiblefork 3 hours ago | parent | prev | next [-]

Morally I agree, but since there's probably a lot of LLM text in the training data, distilling on another model will probably make your model copy the values encoded into the other model as well, even in cases where you only distill on value-neutral stuff.

By copying their programming style, you'll move the model towards that way of writing, which will move the model towards the values expressed in those documents.

I feel that Deepseek v4 got so claudified at the end that it was like Claude.

Barbing 5 hours ago | parent | prev | next [-]

All correct, just help me get over the idea of an open-weight Mythos where one or a dozen of us eight billion does something stupid on the bioweapon front. Smart people who’ve exhausted possibilities for what they can do with books and web search and today’s Kimi/GLM.

Figure we’ll have to reckon with this next year in any case, guess we’ll see.

mobelkh 5 hours ago | parent | prev | next [-]

why can't I use the tokens i paid for anyway?

an hour ago | parent | prev | next [-]
[deleted]
knollimar 5 hours ago | parent | prev [-]

I'm sure they put some BS in their TOS