Remix.run Logo
linkregister 3 days ago

It matters because the closed-source frontier labs spend lots of money on human data (RLHF / RLAIF with human oversight). Moonshot is accused of circumventing these costs. Frontier labs add research costs into their inference pricing. If the market doesn't permit them to sustain sufficient pricing to have a positive cash flow, then their business prospects become weaker and they risk insolvency. Furthermore, other leveraged companies are at risk.

The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".

Legality or lack thereof is one of many data points about whether a thing is noteworthy.

Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.

As consumers this is probably a positive development.

spaceman_2020 3 days ago | parent | next [-]

My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow

And OpenAI scraped and distilled that answer and gave me nothing

linkregister 3 days ago | parent | next [-]

What does your story have to do with Moonshot AI? Do you think they didn't also use the same corpus? Bizarre

voidnullvalue 3 days ago | parent | prev [-]

And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.

overgard 3 days ago | parent | next [-]

You could have gained that stuff prior to LLMs. The leg up you're describing is free information on the internet, not AI. AI just makes it a little easier to find, while also crushing the original sources in the process. (Even if it had a broken culture, is stack overflow even going to exist in a year? Where are they going to train on going forward?)

spaceman_2020 3 days ago | parent | prev | next [-]

Democratization of information, but Sam Altman gets a $100B net worth and I'm still broke :)

I would prefer some sort of democratiziation of the money made from the democratization of information as well

voidnullvalue 3 days ago | parent [-]

Agreed, i wish that it would have done more than change who gets rich off rent-seeking behavior surrounding the knowledge that others created, instead it just consolidated that from many gatekeepers to a few.

I can at least take some measure of pleasure in the fact that it has generally lessened the roadblocks in gathering information. I am still displeased that there are any gatekeepers of humanity's combined knowledge

spaceman_2020 2 days ago | parent [-]

It’s even worse than before

Businesses like these used to public at reasonable valuations. You could ride with them to trillion dollar valuations and grow your own fortune too. Everyone has a story of buying Apple or Google or Amazon stock and making millions

Now they’re going live at trillion dollar valuations and by the time you get in, all the upside has already gone (see Spacex IPO)

Not only did they steal all human data, they also made sure that the upside was only limited to themselves and their cronies

Balooga 3 days ago | parent | prev [-]

Not to be an arse, but didn't you have access to Stack Overflow with all questions/answers prior to LLMs?

voidnullvalue 3 days ago | parent [-]

Yes, but time is finite

noja 3 days ago | parent | prev | next [-]

Isn’t that the same argument they are making for replacing human labour?

Circumventing costs.

SubiculumCode 3 days ago | parent | next [-]

There are many frames that one can place upon this issue. They do not contradict the other. There are moral framings (stole the internet so go eff yourselves, is one), but so is national security, and so is the doomer recursive self improvement risk, and then there is the framing purely on what this implies for future AI training.

I mainly focus on the last.

It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:

A. Cease spending massive amounts of money and compute improving those models.

B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.

C. making the best models available only to select partners and government.

In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.

linkregister 3 days ago | parent | prev [-]

Do you get mad at your computer for replacing clerical workers? What does this nonsense comment have to do with the issue at hand?

watwut 3 days ago | parent [-]

Those were told "find another job" and in fact they were able to find different jobs.

AI companies are gleefully bragging and "making humans obsolete", "permanent underclass" and 40% unemployment rates they plan to create.

They pushed to replace people years BEFORE their technology even can produce that work.

So, you know, it is not the same. But also in fact, clerks did disliked when occasionally arrogant claimed to replace them while pushing unfinished software that dont quite work yet.

andyfilms1 3 days ago | parent | prev | next [-]

Oh, so mass theft is okay as long as American companies are doing it

ffsm8 3 days ago | parent | next [-]

copyright infringement is not theft, even if right holders often claim it is.

part of the definition of theft is that the original owner is deprived of it, which does not apply to copyright infringement.

You can only argue with damages from the perspective of potential profits, still not theft though.

https://en.wikipedia.org/wiki/Theft

hungryhobbit 3 days ago | parent | next [-]

So having tons of AIs quoting various literary works and reproducing knock-offs of them has a positive effect on those books' sales?

I think you're wrong: there is absolutely damage to the authors and publishers from what the AI companies have done.

ffsm8 3 days ago | parent [-]

? I literally said that, how am I wrong?

> You can only argue with damages from the perspective of potential profits, still not theft though.

Damages are not deprival of ownership. They're conceptually related but orthogonal

Also there was no moral judgement from my end, I just pointed out that an incorrect word is being applied. It's just not theft - by definition. But language is a fluid concept and definitions change over time. As people keep misusing it, it will eventually lose its original meaning. Which may have already happened for you, but this change hasn't been settled yet as can be seen from looking at the official definitions of the term, which as of today still mention the criteria

evanelias 3 days ago | parent | prev | next [-]

If you steal an unpopular product from a store, the damage is also only to "potential profits", so how does that differ? It's entirely possible no one would have purchased the product and it would have eventually been discarded/destroyed.

Or with services, if a barber cuts your hair and then you run away without paying them, do you not consider that theft, even though there's no change in ownership occurring?

BlackFingolfin 3 days ago | parent | prev [-]

This is almost funny to me, because in many jurisdictions, software companies sure invested a lot of effort into painting people copying software as thieves. In Germany, they (the software producer lobby, and later politicians influence by the former) even coined and spread the term "Raubkopie", which you could roughly translate as "robbed copy", i.e., that's one step worse than "theft", as a robbery in Germany legally means " theft accomplished by force or intimidation". So, yeah: like putting a knife to the throat of someone while you copy the software.

So, after literally decades of investing into advertising campaigns, lobbying to politicians to pass harsher and harsher laws against software "thieves and robbers", now that big tech are doing it, suddenly we are supposed to consider it with more nuance?

Ahhh... no thank you sir. I really enjoy them drinking their own kool-aid.

SubiculumCode 3 days ago | parent | prev | next [-]

Moreover, reading a copyrighted book and learning from it is not theft.

bigfishrunning 3 days ago | parent | next [-]

Generating a set of weights is not learning.

jayGlow 3 days ago | parent | next [-]

would you say that airplanes don't fly because they don't flap their wings? it's possible to achieve the same things with different approaches.

bigfishrunning 2 days ago | parent [-]

I would say airplanes fly, but I wouldn't say that submarines swim. Things have a bit more nuance, and the field of "learning" isn't as well understood as the ML proponents claim it is.

SubiculumCode 3 days ago | parent | prev [-]

That is a strong statement. I guess you are telling Machine Learning to go fuck itself.

bigfishrunning 3 days ago | parent [-]

No, Machine Learning is an unfortunate name for a well documented process for creating black-box classifiers. The process is good, the name is not.

SubiculumCode 3 days ago | parent [-]

And what, to your mind, would classify something as learning? I assume that your position is not the hard "only humans/living creatures can learn"

bigfishrunning 2 days ago | parent [-]

Honestly, I'm not sure. But I do know that there is an entire field of cognitive science dedicated to understanding learning, and quite frankly it's in its infancy. Evidence of this is that every elementary school introduces new teaching techniques from time to time, and very rarely do they result in any benefit to the people who are doing the learning (more often they benefit consultants...).

However, the current process of "Machine Learning" (which is a semi-random parameter descent/evolutionary replacement process) is unlikely to be equivalent to the way people learn, because we aren't copying/competing/replacing our brain constantly. People are actually very good at learning, but our brain material replaces itself partially and relatively slowly (when compared to how a neural network is trained).

trollbridge 3 days ago | parent | prev | next [-]

Great! Neither is distillation then.

SubiculumCode 3 days ago | parent [-]

Never said it was. Still, understanding to what extent the ability of Chinese labs to keep up to western models with much less compute needs to be understood.

Espressosaurus 3 days ago | parent | prev [-]

Machines are not humans.

linkregister 3 days ago | parent | prev [-]

Reread my comment and look for a value judgement on my part. The final sentence is probably a good clue as to my opinion.

robotpepi 3 days ago | parent | prev [-]

Chatgpt routinely cites and uses papers I don't have access to because they're behind a paywall. I don't think OpenAI is paying for all that copyright. That's in my opinion way more serious.

fc417fc802 3 days ago | parent | next [-]

Yes the fact that the scientific literature - created largely on the back of the tax payer - isn't open to all free of charge by force of law is a travesty. A cartel should not get to charge for access to the bulk of human knowledge. That is indeed a far more important issue than whether or not Moonshot violated the Anthropic ToS, possibly committing mass fraud in the course of doing so.

I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.

Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.

linkregister 3 days ago | parent | prev | next [-]

Your statement is orthogonal to my comment. Why reiterate the schadenfreude / fairness comment already stated several dozen times in this thread?

nylonstrung 3 days ago | parent | prev [-]

Who do you think is paying $100K+ for "Enterprise" access to Anna's Archive?