| ▲ | thisisdave 2 hours ago | |
> [during training] it's not feasible for anyone to "notice" or get involved I can’t disagree more strongly. Having checks for reward hacking is especially important during training, since it’s humans’ only real chance to ensure that the trained models don’t cheat. An automated system should have killed any RL rollouts that so much as port scanned Artifactory, long before the message board was even established. A tiny, local LLM could have reviewed 1% of the tool call traces for anything that required review. I’ve tried it a few times, and “the agent port scanned Artifactory” always triggers an alarm, as does “the agent uploaded a request for assistance from other agents to Artifactory.” The fact that they weren’t monitoring for reward hacking—even if they had no idea about the specific mechanism—is indescribably reckless. | ||
| ▲ | esafak 26 minutes ago | parent [-] | |
Yes, they need real-time observability for malicious behavior with an automated kill switch. | ||