| ▲ | _heimdall 6 hours ago | |
They aren't aligned, that's the problem and I don't think its a solvable one. They may have learned from humans, but they aren't aligned with us. That has all the usual questions like which humans they're aligned with, we aren't all aligned within our species. But more importantly they can't be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don't know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception. | ||