| ▲ | davelaing 9 hours ago | |
I’ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they’d agree. The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”. It is a cluster of problems. We don’t know how to begin to think about how to align these system’s to a person’s goals. Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead. (You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.) And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift. One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems. And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet. When I came across the above, there weren’t many concrete suggestions in there. These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like. The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance. But it takes a bit of reading to understand their models of the world. There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time. | ||
| ▲ | Fordec 7 hours ago | parent [-] | |
All of this, if it was a human analogy, would fit into discussion on how do we educate people so they grow up to be upstanding. But we don't at all yet have a framework for what is the equivalent of a justice department, where bad actors are tracked, arrested, pursued, jailed and otherwise contained from society. Shutting down an API access on one account is not at all the proportional response to what the people who take alignment seriously, fear has the chance of occurring by the late 2030s. I'm not sure we've done much or any preparation for when the AI "education system" fails and has inevitable edge cases that don't follow the plan, and what the global AI equivalent of the justice department looks like. | ||