Remix.run Logo
▲ ben_w an hour ago

A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".

(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).