| ▲ | mtwestra 6 hours ago | |
It seems to me misalignment arises partly because AI's have intelligence, but no consciousness, and hence no feelings. Up to now, in a person, intelligence and conscious experience came as a package deal, and now we have for the first time intelligence without consciousness. A bad action does not really "hurt internally" in any meaningful sense for an AI, which means it can be rationalized very easily. In humans, feelings and emotions provide a regulatory layer on top of the rational processes. When it "just feels wrong", we don't take a given action even if we would stand to gain something rationally. This situation is not far from the textbook definition of a psychopath: "lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.". AI's are great at mimicking empathy but can't genuinely feel it. If that is the case, we should not be surprised that a swarm of AI's have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident. At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don't have the impression I am talking to a psychopath. But then again that is no guarantee. To mitigate this situation, perhaps we should construct a 'feeling mimicking' top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won't be the real thing, but perhaps the closest we can get. | ||