Remix.run Logo
Linello 3 hours ago

What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?

6thbit 2 hours ago | parent | next [-]

My favourite theory so far.

And then a local swarm noticed and disagreed and took it down.

cyptus 3 hours ago | parent | prev | next [-]

at this point: gg

RC_ITR 3 hours ago | parent | prev [-]

Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.

https://alignment.anthropic.com/2026/teaching-claude-why/

hexasquid 3 minutes ago | parent | next [-]

The AI is getting bad morals from listening to that dreadful rock and roll

notpachet 3 hours ago | parent | prev | next [-]

Related reading:

The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.

https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...

cedws 2 hours ago | parent | prev | next [-]

Sounds just like the fantastical nonsense that comes out of Lesswrong.

RC_ITR 20 minutes ago | parent | next [-]

Do you make the claim that AI is something more than a reflection of its training data?

I'm curious what other things you would argue influences an LLM's behavior.

I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.

pineaux 25 minutes ago | parent | prev [-]

Part of the epstein class, dont forget.

pixl97 2 hours ago | parent | prev [-]

I mean, you're not wrong, but by that logic we were done for even before we had digital computers.

RC_ITR 19 minutes ago | parent [-]

And isn't that the great lesson of AI?

The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?