Remix.run Logo
kypro a day ago

It's worth remembering that in a few years that capabilities of these agents are likely to be as far behind the frontier as GPT-4 is today.

As it stands we've made remarkably little progress in terms of alignment and still have no good strategies which are likely to guarantee the alignment of super intelligent systems. As it stands the frontier of alignment is basically some combination of:

- hoping that more intelligent models become more aligned by default (more or less disproved at this point)

- hoping that if you RHLF a model to be a good boy enough it will in fact be a good boy

- asking it nicely in its prompts to be a good boy

- using another model to spot when it's being a bad boy and turning it off

- letting it lose and hoping we can spot when it's bad

There are many arguments which I'm convinced by that would suggest alignment of a super intelligence is impossible.

None of this is surprising to those of us who have been concerned about AI risk for a long-time and have be repeatedly mocked or insulted.

There will be a point of no return if we carry on down this path, and that point is now very rapidly approaching. When it does everyone you know will die, or worse. We should remember we need super-human general intelligences to cure cancer. Select narrow intelligences are fine and allow us to retain control. Let's be sensible about this. We need to stop.

vasco a day ago | parent [-]

Humans are not aligned with each other so who should the AI align with? There's many wars going on, just pick one and do your thought exercise with AI aligned 100% to their human prompters. Which side does the AI refuse to help?

kypro a day ago | parent [-]

> Which side does the AI refuse to help?

Arguably an aligned AI would actively seek to prevent harms we humans seek to cause.

Does the aligned AI really allow humans to bomb and kill each other, or would it understand that it has a moral duty to limit our autonomy for our own good?

It's the first law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.

vasco a day ago | parent [-]

If your answer is that the AI will not do what either prompter wants but what it's own definition of alignment is I hope you know that leads into terminator.

kypro a day ago | parent [-]

You always get terminator the question is only who is its master – itself or humans (the government, etc)?

I think I'd argue it's likely better if it removes human autonomy than unquestionably serves the interests of the US/Chinese government. But there's no point in us worrying about this, that's a choice Sam Altman, et al, must make for humanity.