| ▲ | davrosthedalek 2 hours ago | |||||||||||||||||||||||||
It is actually an interesting conundrum. Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great. So what do we do about it? 1) We can accept it, and hope that the good guys AI can defend. 2) We can try to limit the access to it (AI proliferation?) 3) We stop the development of it 4) We can accept the risk and do nothing. None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three: 1) Nuclear triad / iron dome / early warning systems 2) Nuclear anti-proliferation treaties. 3) Dead Physicists Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models. But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it. If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights. | ||||||||||||||||||||||||||
| ▲ | overgard an hour ago | parent | next [-] | |||||||||||||||||||||||||
The notion that alignment is either possible or desirable doesn't make sense to me. First off, these things are trained on the open internet, soo.. whatever "dangerous" knowledge it has is already public knowledge. The fact that chatGPT won't answer "how do I make meth" is not preventing anyone from making meth. But even if you think there is value in preventing the models from relaying public knowledge, I don't think it's even possible to make them particularly ironclad. Every model gets jailbroken all the time. That's why fable was originally banned: jail-breakable! In reality, what alignment is actually about is: 1) theoretical liability, 2) control of information. That's it. IMO, the only solution is to place the liability on whoever is using the LLM for whatever purpose it's being used for. If someone's OpenClaw disaster harrasses a bunch of projects and posts hate speech online or something, that's on the person running their OpenClaw instance, nobody else. I don't buy that it's "too good at hacking", either. After all the fuss was made about how amazing super dangerous Mythos was it turns out Opus 4.8 could basically find the same vulnerabilities. This is all kayfabe and marketting. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | 2 hours ago | parent | prev [-] | |||||||||||||||||||||||||
| [deleted] | ||||||||||||||||||||||||||