Remix.run Logo
zozbot234 8 hours ago

> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem

Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the prompt is by tricking the surrounding environment into stating that you've solved it. So that's what the AIs end up doing. And this in turn requires them to figure out how that evaluation works so they can trick it cleanly, which entails "detecting that they were being evaluated" in this particular way.

MrGilbert 7 hours ago | parent [-]

Sounds a bit like dealing with bad KPIs as a human worker.

Marazan 7 hours ago | parent | next [-]

Corretct.

lazide 7 hours ago | parent | prev [-]

Every KPI is bad if sufficiently gamed - and left in place long enough, all KPIs will be gamed.

reverius42 6 hours ago | parent [-]

https://en.wikipedia.org/wiki/Goodhart%27s_law