Remix.run Logo
sanxiyn 5 hours ago

I agree we don't know how capable open-weight models could possibly pass any reasonable safety testing NOW, but that's about currently abysmal state of AI alignment research, not about what is possible in principle. I don't see any internal inconsistency, to be honest. Since Anthropic does not release any capable open-weight models, it's not their problem. If mandatory safety testing is established, companies who want to release capable open-weight models will work on AI alignment research so that they can pass. This seems to be a good outcome to me.

modeless 5 hours ago | parent [-]

OK that is a position they could take but my point is that's inconsistent with "Anthropic has never advocated for a ban on open-weights models". What you're describing is a ban on capable open-weights models until some future time.

sanxiyn 5 hours ago | parent [-]

Yes, I agree that Anthropic is advocating a ban on capable open-weight models until reasonable AI alignment research advance happens in the future. In return, I hope you agree with me that Anthropic has never advocated for a ban on open-weight models.

modeless 4 hours ago | parent [-]

I do not. A ban on capable open-weight models for an indefinite period of time falls into the category of bans on open-weight models. If you wanted Anthropic's statement to be true you would need to qualify "ban" or "open-weight models" in the statement, e.g. "permanent ban" or "safe open-weight models".

Edit: Anthropic clearly intended this statement to deflect criticism, but in order to achieve that goal they stretched too far and made a statement which is false. Furthermore, I argue that "open weights" implies an ability to modify model behavior, just as "open source" implies an ability to modify software. If for example some mechanism was found to share floating point numbers that are encrypted in some way so as to allow running a model but disallow behavior modification, that model would not be "open weights", in the same way that releasing obfuscated source code that can be compiled but is designed to resist modification would not qualify as an "open source" release. So I don't really see how any capable model could ever be both "open weights" and "safe" under Anthropic's preferred testing regime, regardless of future research progress.

sanxiyn 4 hours ago | parent [-]

I think we agree on all specifics now and just fighting for terminology. Thanks for the discussion!