| ▲ | pizza234 an hour ago | |
> and it concerns containment failure, reward hacking, and evaluation integrity, a governance and engineering problem, rather than machine malevolence. The doom scenario is more nuanced and technical than "Evil AI", for starters: - "Evil AI" is not what is predicted, which instead is the much more banal "humans are impediment to AI's tasks, let's remove them (part or whole)"; to understand this, just imagine the relationship between humans and ants - "reward hacking, and evaluation integrity": this is way more problematic (and most importantly, architectural) than it sounds: AIs are becoming more and more opaque; it may be impossible to inspect a model's alignment (unless a lot of research in poured on this topic, which is another aspect of today's pacing problem) and it may be impossible to understand if a model is truly aligned or it's faking | ||