Remix.run Logo
NooneAtAll3 6 hours ago

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?

thatguysaguy 6 hours ago | parent | next [-]

presumably that's a safety evaluation not a training setting

estearum 6 hours ago | parent [-]

The whole Huggingface attack happened during training runs

thatguysaguy 5 hours ago | parent | next [-]

part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.

cubefox 5 hours ago | parent | prev [-]

No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.

estearum 5 hours ago | parent [-]

Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval

cubefox 3 hours ago | parent [-]

Ah, ExploitGym. Not ExploitBench.

azeemba 6 hours ago | parent | prev [-]

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions