| ▲ | ozgung 21 hours ago | |
TL;DR: Researchers assumed it would be easier to align these models when they become smarter. In reality, agents are getting increasingly smarter but aligning them also becomes increasingly harder, or even impossible. They are losing their control over the agents. They don’t write their code anymore. They don’t know agents limits nor how to limit them. They are like lab scientists in a Hollywood thriller, watching a mutant organism rapidly evolving in front of their eyes. | ||