| ▲ | verdverm 15 hours ago | |
They are training the agents to be "relentlessly proactive" because they want the agents to run longer, and it makes them more money by using more tokens. But they have trained them to try anything and everything to accomplish any task, so they can run unattended for longer. This is why they do better on benchmarks, it's why they can do things for us for longer, it's that persistence that makes them good at hacking. We do not have to train them to be this way, just like we don't have to train them to be so sycophantic | ||