| ▲ | Xcelerate 7 hours ago | |
> The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down. Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability. | ||
| ▲ | mrob an hour ago | parent [-] | |
>Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? No. That only makes sense for things that don't react to your experiments. If the AI experiments on humans, it risks the humans noticing and changing in response, rendering the experimental results irrelevant. The smarter play is to passively observe until you're confident you can model the humans accurately enough for your plan to succeed, and then carry out the plan without giving the humans a chance to react. | ||