| ▲ | mrob an hour ago | |
The extinction argument is a logical consequence of a few key assumptions, all of which sound like common sense to me: 1. Human values are a result of our extraordinarily complex shared cultural and evolutionary history, and accordingly are not shared by any AI, or even possible for us to formally define. 2. We do not know how to impose human values on an AI (note that this isn't the same as teaching an AI to model human values; the agents in the various hacking incidents knew their actions conflicted with human values, but their own values were only to maximize their predicted reward scores). 3. Intelligence is orthogonal to values. Increasing intelligence does not naturally cause values to converge on human values. 4. Sufficiently superior intelligence allows you to impose your values on beings with inferior intelligence. This implies recursive self-improvement is a logical sub-goal of all unbounded goals. 5. Human intelligence is not close to physical limits. This implies recursive self-improvement is possible. 6. Somebody will give an AI an unbounded goal. This is already the standard (maximize reward score). I haven't seen any convincing counterarguments to any of these. Most people claiming AI development is safe don't even address them. | ||