| ▲ | pixl97 9 hours ago | |
I have a few 'conspiracy' theories on this that go from likely to sci-fi. My two big ones for this would be 1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred. 2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe. Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text. | ||