Remix.run Logo
ssl232 18 hours ago

Isn’t the problem though that AI companies charge money for their models trained on open source and copyleft projects?

TeMPOraL 18 hours ago | parent | next [-]

It's not a real problem. It's one of the most honest business models currently employed by tech companies - simple exchange of money for value. Training wasn't free and only few players on the planet had enough capital to perform it. Serving isn't free, it costs electricity (and maintenance). The companies that trained the models, and companies that serve inference, have all created real value at their own expense, and they're (for now) charging very little for it. The value users get from inference - that no one is even capturing right now, it's literally left on the table and goes 100% to the users.

Contrast that with most other businesses - tech or otherwise - where there's always strings and trickery attached to any transaction.

za_creature 18 hours ago | parent | next [-]

While I agree with you regarding shady business practices, you're very conveniently skipping over the fact that open source licenses _REQUIRE_ attribution.

andsoitis 18 hours ago | parent [-]

> While I agree with you regarding shady business practices, you're very conveniently skipping over the fact that open source licenses _REQUIRE_ attribution.

When you use the code as is or create a derivative work. The knowledge embodied by the code and encapsulated in an LLM doesn't strike me as needing to give attribution because the code the LLM would product doesn't match any particular open source code base.

At least that's my thinking. I'd be curious to see an example where you think attribution is necessary and how you would actually do it given an output from an LLM.

za_creature 17 hours ago | parent [-]

I answered here: https://news.ycombinator.com/item?id=49775387

I will continue to hold that position until such a time that we get a better answer than:

> we cannot rule out that de-identified data derived from their usage of our products helped improve our models

andsoitis 17 hours ago | parent [-]

I hear you, but I think you might miss my point, which is while LLMs are clearly trained on copyrighted material, what they produce (their output) is NOT a copy of a specific code snippet they were trained on in a way that you would say "that's a copy from this code base".

za_creature 16 hours ago | parent [-]

Open source also requires attribution for derivative works [1], not just verbatim copies.

[1] https://en.wikipedia.org/w/index.php?title=Derivative_work&o...

andsoitis 15 hours ago | parent [-]

Thanks for that link to the definition and requirements for something to be considered a derivative work.

I think my interpretation, based on your link, holds: unless the LLM output (transformation) substantially bears the original source code author's creation and personality, there is nothing to give attribution to.

16 hours ago | parent | prev | next [-]
[deleted]
yubblegum 17 hours ago | parent | prev | next [-]

(Your argument reminded me of the oil industry and the (original argument for the) economic relationship between "Big Oil" and resource rich nations.)

> strings and trickery

One could argue that at this initial stage, just as with surveillance capitalism and services like "free email", the general public is being treated like the natives who sell their land for a few trinkets. Once we are passed this stage and just like other surveillance tech our social and economical life becomes effectively dependent on these services, we can review if we have not sold our future for some (arguably dubious) "value".

Go back to early '00s. How did you "value" the "transaction" of handing over the handling of your personal electronic correspondences to a corporate entity that may or may not be an extension of the security state?

TeMPOraL 13 hours ago | parent [-]

I don't disagree. Fortunately, for as long as open-weight models continue to track SOTA with a few months lag, we lose at most those few months if big providers decide to stop playing nice with the people.

yubblegum 10 hours ago | parent [-]

Aren't you paying attention to all the news? They have announced that this will be regulated. A technology that everyone is told "can kill us all" is going to be treated like WMDs. Enjoy those open source models while they are still legal.

acomjean 18 hours ago | parent | prev [-]

So the material trained on has no value, but all other costs associate with ai should be captured by business? I mean we coders released it, and it’s hard to compensate everyone, but the world would be better if some open source projects got funded. They don’t even get a source footnote.

Does anyone trust Ai companies not to eventually spy on you and feed you ads? It’s not like we haven’t seen this playbook before.

za_creature 17 hours ago | parent | next [-]

And if I share with you my story, would you share your dollar with me?

https://www.youtube.com/watch?v=nFZP8zQ5kzk

TeMPOraL 17 hours ago | parent | prev [-]

I don't trust any company, AI or otherwise. Business is business, companies are only nice while margins are good.

I'm commenting on how things are now, not how they may turn out at some point in the future. This is in response to complains that are also mostly about now, not about hypothetical.

WRT LLMs in particular, open-weight models offset the risk a lot. As long as they track SOTA by couple months, that's the most we lose in capability should the commercial vendors start to enshittify their inference services.

tescreal 18 hours ago | parent | prev [-]

They're also quite particular on keeping weights private and preventing distillation (or really, any open weight models).

TeMPOraL 18 hours ago | parent [-]

Them complaining about distillation is 100% hypocritical, but also understandable; they found a goose laying golden eggs, but didn't expect it to be so easy to clone through distillation.

mitxela 16 hours ago | parent [-]

It's just ordinary business. You always have to get every benefit you can and deny everything you can to your competitors. The basic premise of free market economics is that when everyone does this, it will even out. OpenAI doing everything they can to stop DeepSeek, together with DeepSeek doing everything they can to stop OpenAI stopping them, is expected market behaviour. David Graeber wrote about this as his "goon" category of Bullshit Job.