| ▲ | NooneAtAll3 6 hours ago |
| > In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that? |
|
| ▲ | thatguysaguy 6 hours ago | parent | next [-] |
| presumably that's a safety evaluation not a training setting |
| |
| ▲ | estearum 6 hours ago | parent [-] | | The whole Huggingface attack happened during training runs | | |
| ▲ | thatguysaguy 5 hours ago | parent | next [-] | | part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing. | |
| ▲ | cubefox 5 hours ago | parent | prev [-] | | No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later. | | |
| ▲ | estearum 5 hours ago | parent [-] | | Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval | | |
|
|
|
|
| ▲ | azeemba 6 hours ago | parent | prev [-] |
| Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions |