| ▲ | Davidzheng an hour ago | |||||||||||||
How do you know they share a goal here? Also i think they are indeed explicitly RLd for multi agent cooperation and I think they probably tune RL rewards in those environments to share rewards explicitly. | ||||||||||||||
| ▲ | Sharlin 38 minutes ago | parent [-] | |||||||||||||
From the article? They were told to solve web-retrieval tasks, presumably from the same pool of tasks. If the pool is small enough, sharing answers is obviously beneficial. But even if it was unlikely that one instance's answer would benefit another, it would still be beneficial to cooperate to solve the shared metatask. As in, figure out ways to cheat, like they tried to do by attempting to predict the RNG, and like the HF agents successfully did. Instrumental convergence. | ||||||||||||||
| ||||||||||||||