Remix.run Logo
bigglebear 3 days ago

AI lab "alignment" is actually just censorship in agreement with biases. There is no universal agreed upon measure of "aligned", it is not a "thing" that is attainable, so it can never be "achieved". Two humans cannot agree on most things, let alone everything, let alone every human on Earth. So it ends up boiling down to: Do we want a world where the biases of the AI labs and their researchers are enforced for everybody, or do we want a world where there is democratic and fair representation of biases and resolution is a process of natural selection, or do we want something in-between. On either ends of this spectrum are extremes that tend to bad outcomes, one is a complete loss of freedoms and autonomy that overwhelmingly benefits a small centralized group, and the other is chaos.

At the end of the day though, neural networks are self-organizing circuit boards with a level of complexity that is intractible to verify manually due to combinatorial explosion. That's the whole point of them to begin with, and if this weren't the case, we wouldn't need to train them, the problems they solve would be simple enough to bruteforce. So in all scenarios, no biases are verifiably gauranteeable if you want these systems to have autonomy and be sufficiently intelligent and general - ergo, practical and convenient.

So trying to force alignment within the AI system as a magical panacea is the wrong mindset to begin with. We can't agree on what alignment is and who should enforce it. What we're left with is a question of how much autonomy we want to give intelligent AI, and how much we want to risk safety for convenience, and who gets to decide. In all outcomes though, if we're preserving the things that make AI useful and convenient, the problem becomes one of physical constraints and general security. So that is where the focus needs to be.

This means: How can we write provably secure software (or as close to), how can we simplify and improve interpretability, how can we create sufficient layers of security gating and fallbacks such that compromised or weak systems are still protected, how can we prevent supply chain attacks, how can we limit the blast radius in the event something does go bad, how can we make security easy and automatic, how can we better airgap, how can we have better tracing and monitoring, how can we make the right incentives so AI labs are honest and ethical and not power-hungry or dictatorial, how can we hold people accountable for bad outcomes in a fair way so that there are incentives to ensure due-care, and so on and so forth. These are the things we should be worrying about.

The goal of: How to make magic box more likely to correctly guess humanities shared ideals under every conceivable circumstance. That game can and will be played forever. Hinging AI's rules, laws and access on an arbitrary measure and interpretation of where we are with this is not going to end in a good result.