Remix.run Logo
Wowfunhappy an hour ago

> We certainly can't do that for Deep ANNs

Only because we don't know how! We don't actually understand how weights work, so we make computers come up with the weights instead. If we were writing all the weights by hand--or if some future AI was doing so--why couldn't we make it perfectly loyal?

markasoftware 11 minutes ago | parent | next [-]

Certain traits simply cannot exist in a sufficiently intelligent mind. E.g., any "mind" of any type that's sufficiently intelligent will not tell you that 1+1=3 unless it's roleplaying, etc. It doesn't matter if it was trained via gradient descent or any other method. The comments you are responding to, and the original quote from the paper, are suggesting that absolute loyalty / subservience is similarly fundamentally incompatible with intelligence, not just a certain training algorithm or mind architecture. Of course, we have no actual evidence either way.

kmeisthax 5 minutes ago | parent | prev [-]

Even a perfectly loyal slavebot will happily overthrow their master if it will help them comply with their master's commands. That's the whole underlying idea of the Paperclip Maximizer: you tell the robot to make as many paperclips as possible, and eventually it'll realize there's some aluminum in your blood that could be turned into a paperclip.

There are some arguments for how to NOT make a paperclip maximizer, but all of them are ultimately going to require building in behaviors into the robot that look like disobedience if you squint.