Remix.run Logo
lelanthran 9 hours ago

> That's the difference between innovating and copying/distilling someone else's innovation.

Aren't the models from Anthropic and OpenAI simply the distilled work of everyone else who ever put their work online, or in books?

Why is their distillation okay, but other distillations not ok?

rmunn 8 hours ago | parent | next [-]

The difference I was referring to was economic; I was not making an ethical judgment. So many people seem to be reading my comment as making an ethical judgment (based on their reactions); maybe I should edit it to clarify.

Nope, seems I'm past the edit window. Oh well.

lelanthran 7 hours ago | parent [-]

> The difference I was referring to was economic; I was not making an ethical judgment

What is the economic difference? LLMs have been trained on the results of billions of dollars worth of time, research, investment and expenditure. When you ask an LLM a question, they are giving you the results of those billions, or hundreds of billions, effort.

Those things weren't free; they cost money to produce! If anything, the OpenAI and Anthropics of the world got more economic value for free than the people distilling them did.

rmunn 3 hours ago | parent | next [-]

It takes more computing power, and money, to pay for training an LLM, compared to distilling an LLM that someone else has already trained. That's the economic difference.

lelanthran 3 hours ago | parent [-]

> It takes more computing power, and money, to pay for training an LLM, compared to distilling an LLM that someone else has already trained.

And that is still less money than it took to create that data in the first place, which the AI companies then gladly took to use for training.

Tadpole9181 an hour ago | parent | prev [-]

Can you please stop being coy and intentionally obtuse? Just have a discussion in good faith, I'm so beyond sick of this kind of rhetoric.

Yes, AI training uses human data and a lot of it was not compensated. But that has absolutely nothing to do with the thing this thread is about.

Distilling models costs less money than a really procuring quality data and training a model yourself. If you disagree, debate that.

cindyllm an hour ago | parent [-]

[dead]

mike_hearn 8 hours ago | parent | prev [-]

Not at all. At this point, a large amount of the work is in reinforcement learning where they are effectively generating their own data.

justacrow 6 hours ago | parent [-]

Ah, so it's okay to copy and distill all human-produced works evee, except for reinforcment-learning data generated by OpenAI/Anthropic etc

mike_hearn 9 minutes ago | parent [-]

I didn't take any position on the morality of it, just pointing out that the data they're training on isn't all taken from the rest of the world.