| ▲ | _heimdall 6 hours ago | ||||||||||||||||||||||
> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference. > were intentionally misaligned or had guardrails turned off Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model. | |||||||||||||||||||||||
| ▲ | ifwinterco 6 hours ago | parent | next [-] | ||||||||||||||||||||||
And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing. They don't seem to do that, which means either they are: - very stupid (which seems unlikely, the one thing these people don't lack is IQ) - very careless (possible, but these are the same people that say AI will end the world, so would you be careless?) - they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this) - or they want this to happen | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | coffeebeqn 5 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | exitb 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires. Have we ever seen an LLM with a hobby? | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | nutjob2 5 hours ago | parent | prev [-] | ||||||||||||||||||||||
> That seems likely, but we have no way of knowing this. Only humans can 'know', because all we can be certain about is that humans do such a thing. If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text. | |||||||||||||||||||||||