Remix.run Logo
RC_ITR 4 hours ago

Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.

https://alignment.anthropic.com/2026/teaching-claude-why/

notpachet 4 hours ago | parent | next [-]

Related reading:

The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.

https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...

HarHarVeryFunny an hour ago | parent | prev | next [-]

They could filter what they train on if they wanted to - they just don't want to.

hexasquid an hour ago | parent | prev | next [-]

The AI is getting bad morals from listening to that dreadful rock and roll

cedws 3 hours ago | parent | prev | next [-]

Sounds just like the fantastical nonsense that comes out of Lesswrong.

RC_ITR an hour ago | parent | next [-]

Do you make the claim that AI is something more than a reflection of its training data?

I'm curious what other things you would argue influences an LLM's behavior.

I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.

42 minutes ago | parent [-]
[deleted]
pineaux an hour ago | parent | prev [-]

Part of the epstein class, dont forget.

pixl97 3 hours ago | parent | prev [-]

I mean, you're not wrong, but by that logic we were done for even before we had digital computers.

RC_ITR an hour ago | parent [-]

And isn't that the great lesson of AI?

The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?