| ▲ | yewenjie 19 hours ago | |
I am somewhat confident that right now we have crossed a threshold of model capability that we will continue to see such breaches and unsanctioned actions by models in the coming months, some of which would be out in the wild, until someone comes up with some really robust control (keeping the AIs on leash) technique that adequately enforces the sanctioned actions. Even that guarantees almost nothing about real alignment (making the AIs want to predict and behave how we would have wanted them to behave). | ||
| ▲ | 18 hours ago | parent [-] | |
| [deleted] | ||