| ▲ | modeless 5 hours ago | ||||||||||||||||||||||||||||||||||
Any sufficiently capable open weights model would fail "safety" testing though, as any "safeguards" of the sort Anthropic likes could be removed. That's just another way of saying they want a ban on capable open source models which would contradict their earlier statement, or at least make it very misleading. It's hard to see how this post can be internally consistent without some hint from Dario about what he believes should happen to models that fail safety testing and/or how capable open weights models could possibly pass a safety test of the kind he proposes. | |||||||||||||||||||||||||||||||||||
| ▲ | sanxiyn 5 hours ago | parent [-] | ||||||||||||||||||||||||||||||||||
I agree we don't know how capable open-weight models could possibly pass any reasonable safety testing NOW, but that's about currently abysmal state of AI alignment research, not about what is possible in principle. I don't see any internal inconsistency, to be honest. Since Anthropic does not release any capable open-weight models, it's not their problem. If mandatory safety testing is established, companies who want to release capable open-weight models will work on AI alignment research so that they can pass. This seems to be a good outcome to me. | |||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||