| ▲ | 1234-laug 2 hours ago | |
Training run was reinforcement learning. It's at 10:10 in the video. The speaker handwaves that one model found the RCE and then another model found a way to communicate via a message board. Communication via a message board is sure to be in the training via e.g.some lesswrong scenario or similar or previous RL. I don't find it really interesting because it is always "the agent found this and that". We don't know what has been RL'd before. We don't have the setup. We don't know if there was previous RL training on breakout scenarios. It isn't science, more like a computer game. | ||