Remix.run Logo
verdverm 5 hours ago

abliteration and fine tuning makes it not so straightforward

https://huggingface.co/blog/mlabonne/abliteration

sanxiyn 5 hours ago | parent [-]

I know, but "we don't know how to make it not dangerous, so it should be allowed to release dangerous things" is... not convincing?

modeless 5 hours ago | parent | next [-]

My point is that advocating a de facto ban on capable open source models is inconsistent with Dario's statement here that "Anthropic has never advocated for a ban on open-weights models." Call a spade a spade.

sanxiyn 5 hours ago | parent [-]

De facto ban on capable open-weight models doesn't seem inconsistent with Dario's statement to me. One, it is de facto, not de jure, and it can and will change as AI alignment research advances. Two, it is only capable open-weight models, not open-weight models. In fact, Dario says non-dangerous (which for now is mostly non-capable) open-weight models are a public good, and I agree.

verdverm 5 hours ago | parent [-]

How do we define "capable"?

Is Kimi K3 capable? It's already out and being run by US companies on US hardware in US data centers.

https://huggingface.co/moonshotai/Kimi-K3

sanxiyn 4 hours ago | parent [-]

That is a difficult question I am not qualified to answer, but Mythos 5 was export controlled for a brief time due to its cybersecurity capability and implications to national security, so for cybersecurity "as capable as Mythos 5" seems to be a good baseline. I wouldn't know for biosecurity though.

UK AISI preliminary evaluation suggests Kimi K3 is not capable enough for cybersecurity in this sense.

https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...

verdverm 4 hours ago | parent [-]

I would be hesitant to extrapolate from this analysis

https://exploitbench.ai/#honest-limits

verdverm 5 hours ago | parent | prev [-]

There is no movement on global policy or enforcement. One country banning their people access to the best models hinders their people.

I am unconvinced that "this can be used dangerously, therefore we must ban it" argument. The OpenAI/Huggingface, needing to turn to Chinese open weight to defend themselves seems to support the case that we need open access and freedom to compute as we see fit.