Remix.run Logo
sanxiyn 5 hours ago

If a model fails the test, it should be banned. He is not advocating a ban of open-weight models. He is advocating a ban of models that fail mandatory safety testing. Seems reasonable and straightforward.

EmbarrassedHelp 2 hours ago | parent | next [-]

> He is not advocating a ban of open-weight models

He is though. He wants open weight models banned that do not pass some set of tests.

And what does "safety" mean here? We constantly see these companies treating NSFW content as "unsafe", despite the fact that its not. Is being able to produce adult content going to result in a model being declared "unsafe"?

modeless 5 hours ago | parent | prev | next [-]

Any sufficiently capable open weights model would fail "safety" testing though, as any "safeguards" of the sort Anthropic likes could be removed. That's just another way of saying they want a ban on capable open source models which would contradict their earlier statement, or at least make it very misleading. It's hard to see how this post can be internally consistent without some hint from Dario about what he believes should happen to models that fail safety testing and/or how capable open weights models could possibly pass a safety test of the kind he proposes.

sanxiyn 5 hours ago | parent [-]

I agree we don't know how capable open-weight models could possibly pass any reasonable safety testing NOW, but that's about currently abysmal state of AI alignment research, not about what is possible in principle. I don't see any internal inconsistency, to be honest. Since Anthropic does not release any capable open-weight models, it's not their problem. If mandatory safety testing is established, companies who want to release capable open-weight models will work on AI alignment research so that they can pass. This seems to be a good outcome to me.

modeless 5 hours ago | parent [-]

OK that is a position they could take but my point is that's inconsistent with "Anthropic has never advocated for a ban on open-weights models". What you're describing is a ban on capable open-weights models until some future time.

sanxiyn 5 hours ago | parent [-]

Yes, I agree that Anthropic is advocating a ban on capable open-weight models until reasonable AI alignment research advance happens in the future. In return, I hope you agree with me that Anthropic has never advocated for a ban on open-weight models.

modeless 4 hours ago | parent [-]

I do not. A ban on capable open-weight models for an indefinite period of time falls into the category of bans on open-weight models. If you wanted Anthropic's statement to be true you would need to qualify "ban" or "open-weight models" in the statement, e.g. "permanent ban" or "safe open-weight models".

Edit: Anthropic clearly intended this statement to deflect criticism, but in order to achieve that goal they stretched too far and made a statement which is false. Furthermore, I argue that "open weights" implies an ability to modify model behavior, just as "open source" implies an ability to modify software. If for example some mechanism was found to share floating point numbers that are encrypted in some way so as to allow running a model but disallow behavior modification, that model would not be "open weights", in the same way that releasing obfuscated source code that can be compiled but is designed to resist modification would not qualify as an "open source" release. So I don't really see how any capable model could ever be both "open weights" and "safe" under Anthropic's preferred testing regime, regardless of future research progress.

sanxiyn 4 hours ago | parent [-]

I think we agree on all specifics now and just fighting for terminology. Thanks for the discussion!

makeitdouble 5 hours ago | parent | prev | next [-]

Can you think of anything open-source that has to go through mandatory tests to be distributed and still survived ?

There is a reason to it, that's as good as any angle to find why IMHO.

sanxiyn 5 hours ago | parent [-]

I think Gemma will be fine. Most open-weight models are not capable enough to be dangerous. Yes, I can't think of any capable open-weight model that would survive reasonable safety testing.

dnw 5 hours ago | parent | prev | next [-]

He is not advocating for banning models per se but the proposal makes a business model (i.e. serving open weight models) that is starting to work more expensive.

sanxiyn 4 hours ago | parent [-]

Agreed, and that serves Anthropic. It seems unproblematic to me. Dario probably sincerely believes in mandatory safety testing for capable models (open and closed), and likes the fact that it aligns with Anthropic's interest.

verdverm 5 hours ago | parent | prev [-]

abliteration and fine tuning makes it not so straightforward

https://huggingface.co/blog/mlabonne/abliteration

sanxiyn 5 hours ago | parent [-]

I know, but "we don't know how to make it not dangerous, so it should be allowed to release dangerous things" is... not convincing?

modeless 5 hours ago | parent | next [-]

My point is that advocating a de facto ban on capable open source models is inconsistent with Dario's statement here that "Anthropic has never advocated for a ban on open-weights models." Call a spade a spade.

sanxiyn 5 hours ago | parent [-]

De facto ban on capable open-weight models doesn't seem inconsistent with Dario's statement to me. One, it is de facto, not de jure, and it can and will change as AI alignment research advances. Two, it is only capable open-weight models, not open-weight models. In fact, Dario says non-dangerous (which for now is mostly non-capable) open-weight models are a public good, and I agree.

verdverm 4 hours ago | parent [-]

How do we define "capable"?

Is Kimi K3 capable? It's already out and being run by US companies on US hardware in US data centers.

https://huggingface.co/moonshotai/Kimi-K3

sanxiyn 4 hours ago | parent [-]

That is a difficult question I am not qualified to answer, but Mythos 5 was export controlled for a brief time due to its cybersecurity capability and implications to national security, so for cybersecurity "as capable as Mythos 5" seems to be a good baseline. I wouldn't know for biosecurity though.

UK AISI preliminary evaluation suggests Kimi K3 is not capable enough for cybersecurity in this sense.

https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...

verdverm 4 hours ago | parent [-]

I would be hesitant to extrapolate from this analysis

https://exploitbench.ai/#honest-limits

verdverm 5 hours ago | parent | prev [-]

There is no movement on global policy or enforcement. One country banning their people access to the best models hinders their people.

I am unconvinced that "this can be used dangerously, therefore we must ban it" argument. The OpenAI/Huggingface, needing to turn to Chinese open weight to defend themselves seems to support the case that we need open access and freedom to compute as we see fit.