Remix.run Logo
▲ zzzeek 7 hours ago

I follow anti-LLM discourse quite a lot, and across the main bulletpoints: energy/carbon emissions, content worker harm, job displacement, deskilling, mental health effects, and copyright/plagiarism, the plagiarism one seems to have the most attention, and it's also the most solvable, if there were only more serious effort on ethically sourced models that can actually do the real science / math / code work that is what LLMs are best at. The whole world of LLMs to create videos/books/art/literature is where most of the offense is (the video/imagery side of it is where most of the energy/carbon emissions problems are too. and content worker harm).

I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.

▲UqWBcuFx6NV4r 7 hours ago | parent | next [-]

Yep. I’d probably be a lot less chastised in some circles for using Claude Code at work if it wasn’t misconstrued as being in support of, I don’t know, encroaching on the hypothetical commissions of a chronically online instagram furry artist or something.

It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.

I am skeptical of there being sufficient data to build “ethical” training datasets, and I’m confident that much of the same contingent will (somewhat rightfully) argue that ‘second-generation’ copyrighted AI material has already irreversibly made its way into every modern dataset.

▲__MatrixMan__ 5 hours ago | parent | next [-]

A reasonable next move, if the US was interested in acting like a democracy, would be legislation forcing this decision:

- prove that you had the rights for all of your training data

- open source the model

Give the labs a 3 month grace period in which to comply, so competition can persist even with dubiously sourced data, but the people can't be locked away from derivatives of their contributions for any significant amount of time.

▲latexr 7 hours ago | parent | prev | next [-]

> It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.

That’s not a justification. If a company were poisoning the water to your home as a byproduct, would you be satisfied if they told you “we don’t necessarily disagree with you about polluting the water, but in our field—which you do not understand, and in which the underlying build process is often not the water pollution—what we’re doing presents very real productivity gains”?

> I am skeptical of there being sufficient data to build “ethical” training datasets

Then you don’t build any. What fucked up world we live in where people think it’s OK to be unethical because they want something and can’t think of any other way to do it. What monumentally selfish rotten babies.

▲zzzeek 4 hours ago | parent [-]

is there nothing else that we take for granted as a convenience to humanity that also has serious negative externalities? just AI ?

▲ 9 minutes ago | parent | next [-]
[deleted]
▲ 12 minutes ago | parent | prev [-]
[deleted]
▲sublinear 7 hours ago | parent | prev | next [-]

> has already irreversibly made its way into every modern dataset

The "gray goo" scenario finally happens... for AI. That's actually the good ending for humanity. I love it! Poetic and believable. Data doesn't "heal" like nature. :D

▲zzzeek 7 hours ago | parent | prev | next [-]

I think you can train on math /science using synthetic generation to a significant extent. Training for coding requires more of the "scraping github / stackoverflow" angle but IMO that's a shallower hill to climb than scraping copyrighted art and literature.

There are actual models trained on ethical datasets but they are obviously not very high powered. If companies with the resources of an anthropic or openai were doing it (ha) it would be more feasible

▲johnnyanmac 7 hours ago | parent | prev | next [-]

>“I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains”

Being able to prove such gains in better products would be a start. And an emphasis on how it assists existing engineers/mathmaticians/researchers, not that any accomplishment made with AI assistance is "AI solves problem".

I don't know whatever happened to "words are cheap". I guess it literally made money to say words, so that adage is false for the time being.

>I am skeptical of there being sufficient data to build “ethical” training datasets

Well if all those scam job ads paying 100/hr to create AI training content was not a scam and instead the approach from the start, there may have been a chance to bridge that gap ethically. The industry chose to break things and is trying to act mad that people are mad at all the broken stuff.

These results are entirely a consequences of the actions chosen. And I don't believe there was ever an honest consideration of there being ethical training datasets. They just thought they could brute force society with fearmongering and bribes. The BOTD was already low in the beginning but completely gone now.

▲ 7 hours ago | parent | prev | next [-]
[deleted]
▲forthegains 7 hours ago | parent | prev [-]

Yeah and it only took stealing all the intellectual output of everyone ever made.

But sure, there's um, an ethical way of doing that?

▲johnnyanmac 7 hours ago | parent [-]

In theory, sure.

1. Only use open source/CC compliant assets.

2. Acquire rights/licenses to any datasets that do not fit #1. e.g. the Google deal with Reddit for 60m/yr.

3. Offer programs to have creatives willingly submit their data, with some sort of residual output based on the number of times their assets are sampled.

4. If all that is still not enough, hire creatives to create assets for you. This is something Spotify did recently with "ghost artists"[0]. The intentions here are suspect, but a non-consumer facing artist providing work for an LLM wouldn't have the same ethical dilemmas

5. Lastly, if all that still isn't enough: governmental programs to either provide grants, subsidies, or more outreach to get the ball rolling.

Would this cost tens, hundreds of billions of dollars? Yes. But clearly, that was not a barrier to entry for the industry anyway. So we can chalk this down to the personality of leadership or the wider culture of modern big tech

[0]: https://harpers.org/archive/2025/01/the-ghosts-in-the-machin...

▲fwip 5 hours ago | parent [-]

I agree with most of your post, but I don't think I agree that method 2 (Reddit deal) is necessarily ethical. I know it's too high of a bar for our nation to ever clear, but I think explicit author opt-in is the only ethical source of AI training data.

Like, legally, I'm sure Reddit had the right to sell it, but probably over half their content was written before ChatGPT was ever announced. The TOS allowing reddit to make "derivative works" was largely understood to mean things like cropping photos, using your viral post in an ad, or maybe auto-translating your comment.

▲altermetax 7 hours ago | parent | prev | next [-]

I don't really see the difference, code is protected by copyright (or copyleft) as much as art is, and yet the LLM scrapers use it without scruples. Same goes for math and science publications.

▲zahlman 5 hours ago | parent [-]

Artists take pride in the intentionality of every brushstroke, and for them any given work is much more likely to reach a point where it's considered "finished".

While coders may care about the craft (and I do), it's not as if the value of my code is in the exact variable names I chose.

▲CapsAdmin 7 hours ago | parent | prev | next [-]

In my experience, the science/math/code crowd don't care about copyright/plagiarism as much as the video/art/literature crowd, so the first crowd turns a blind eye to most of the latter crowd talks about.

Code can be art, and copyright/plagiarism is real. It sort of boils down to how much it bothers us.

▲harimau777 6 hours ago | parent | prev | next [-]

I don't think it is likely that they could get enough data without stealing. It would be incredibly costly to have to pay artists to church out art just to train an AI.

▲hardbass 6 hours ago | parent | prev | next [-]

I hope thats possible but I am not sure if proper knowledge of science can be had without also learning other literature (and vice versa).

▲asa123 7 hours ago | parent | prev | next [-]

i think math is close to art (just to be a contrarian, but kind of really)

its somewhat funny that math people are in a conundrum as to support or not support but this might partially be because some wish to believe that math itself is and can be useful and therefore accelerating is good

but the art people have no such delusions so they’re just strictly against

imo proof writing is more akin to art than coding/tech but…

▲vouaobrasil 7 hours ago | parent | prev | next [-]

> I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.

I disagree that they can be separated. Practically, I think they can't. Because the mere invention of new tools inspires even more AI advancement and that in turn will cause the other side (artistic side) to degenerate even more.

I'm anti-LLM all the way, 100%, no exceptions. Zero tolerance.

▲karahime 6 hours ago | parent | next [-]

So you recognize that discovery cannot be cut off from other discovery, that it's never just one or the other, and your solution is to say shut down all discovery?

▲zzzeek 4 hours ago | parent | prev [-]

you can distinguish between the LLMs you have zero tolerance for and a system like Google Translate? You have a sharp line you can draw for when something becomes "an LLM"?

▲PunchyHamster 7 hours ago | parent | prev | next [-]

They can't be ethically sourced and good at the same time.

The current models intelligence depends on massive training dataset of essentially stolen data

▲mehrzad 7 hours ago | parent | next [-]

While that is true, theoretically a regulation could be enacted that output tokens must focus on STEM research and other practical tasks and the LLM must refuse tasks outside of those areas, just as Claude disallowed cybersecurity tasks. Obviously this would never happen, but the theft of the training data wouldn’t matter as much if the usecases were less sinister.

▲zzzeek 7 hours ago | parent | prev [-]

openai and anthropic trained on actually stolen data since it was pirated datasets.

google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.

▲singpolyma3 7 hours ago | parent | prev [-]

Models aren't in school so plagiarism doesn't really apply

▲Loughla 7 hours ago | parent [-]

I can't tell if you're serious or not.