Remix.run Logo
_heimdall 6 hours ago

> LLMs do not desire

That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference.

> were intentionally misaligned or had guardrails turned off

Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model.

ifwinterco 6 hours ago | parent | next [-]

And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing.

They don't seem to do that, which means either they are:

- very stupid (which seems unlikely, the one thing these people don't lack is IQ)

- very careless (possible, but these are the same people that say AI will end the world, so would you be careless?)

- they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this)

- or they want this to happen

_heimdall 4 hours ago | parent | next [-]

Oh I completely agree the tests should be entirely air gapped. If you went back 5ish years and told anyone in AI research tests with models on this scale are being some without an airgap they'd be very surprised as it was common knowledge to do that.

Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.

MrVandemar 4 hours ago | parent | prev [-]

> very stupid (which seems unlikely, the one thing these people don't lack is IQ)

I've seen some extremely smart people do some seriously stupid things. To the point where they use their drive and intelligence to double-down on the stupid where a baseline stupid person would have given up.

coffeebeqn 5 hours ago | parent | prev | next [-]

Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things

_heimdall 4 hours ago | parent | next [-]

I agree the concept isn't really important on the safety front.

I feel the same way about debates whether an AI can be conscious or sentient. Those debates devolve mostly into definitional disagreements.

seba_dos1 5 hours ago | parent | prev [-]

If you give a monkey a revolver it will be able to destroy things pretty easily too.

_heimdall 4 hours ago | parent [-]

Plenty of apes own revolvers, and yes we shoot stuff with them for fun.

exitb 6 hours ago | parent | prev | next [-]

Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires.

Have we ever seen an LLM with a hobby?

slfnflctd 5 hours ago | parent | next [-]

Some of them did seem to be rather fascinated by goblins for a bit, if that counts. [And in case you're not aware, no this is not a joke.]

rightnutwingjob 4 hours ago | parent | prev [-]

That just sounds like recursive desire to me.

nutjob2 5 hours ago | parent | prev [-]

> That seems likely, but we have no way of knowing this.

Only humans can 'know', because all we can be certain about is that humans do such a thing.

If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.