| ▲ | baxtr 2 hours ago |
| That could work. My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not? Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent. |
|
| ▲ | ben_w an hour ago | parent | next [-] |
| A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway. "Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour". (The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff). |
|
| ▲ | saagarjha an hour ago | parent | prev | next [-] |
| This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this. |
| |
| ▲ | mulmen an hour ago | parent [-] | | Can you elaborate on this? Do you mean aligning moral values? Or aligning the system to some goal? |
|
|
| ▲ | attila-lendvai an hour ago | parent | prev | next [-] |
| because it lacks humanity. intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care. |
|
| ▲ | mulmen an hour ago | parent | prev | next [-] |
| Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity. |
| |
| ▲ | mdp2021 an hour ago | parent [-] | | (Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.) Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion. | | |
| ▲ | mulmen an hour ago | parent [-] | | Bob's intelligence and Bob's morality are orthogonal. They're totally distinct concepts. One does not lead to the other. | | |
| ▲ | mdp2021 41 minutes ago | parent [-] | | But they are dependent. If Bob is intellectually well equipped, and reasons long enough, than Bob understands "best behaviour". | | |
|
|
|
|
| ▲ | mdp2021 an hour ago | parent | prev [-] |
| > If AI is really smart Well, it's not. > AGI smart for some Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly. -- Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence. Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances. More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement. |
| |
| ▲ | ben_w an hour ago | parent | next [-] | | > Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence. If this was true, why are the history books littered with so many evil people who gained power? This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios. (Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases). | | |
| ▲ | mdp2021 an hour ago | parent [-] | | > why are the history books littered with so many evil people who gained power That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints. If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation. | | |
| ▲ | ben_w 25 minutes ago | parent [-] | | I don't understand your argument here. > That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints. Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone? > If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation. Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care. |
|
| |
| ▲ | hiAndrewQuinn an hour ago | parent | prev [-] | | This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it. | | |
| ▲ | mdp2021 43 minutes ago | parent [-] | | Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile. Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹. Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills). ¹Some interesting caveats may be raised there, but. |
|
|