| ▲ | simonw an hour ago | |
They didn't have watchdog agents - those exist for their production models but had been deliberately removed for the purpose of this evaluation. OpenAI wrote about how their mechanism for that in production works here: https://openai.com/index/safety-alignment-long-horizon-model... > We created a monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary. The monitor observes not just a single action but the entire trajectory. | ||