| ▲ | xscott 4 hours ago | |
> [...] we evaluated the behavior of various Claude models in a setting with contradictory objectives. > We consistently saw a multiagent turf war... In fact, they sabotaged others with increasingly aggressive, self-replicating malware. Seems like Anthropic should withdraw their models until they can be taught to behave and cooperate as well their competitors (both open and closed) do. /s I hate fearmongering, and I don't trust Dario's intentions for doing it. | ||