Remix.run Logo
thierrydamiba 6 hours ago

“Anthropic has never advocated for a ban on open-weights models.

Open-weights models that don’t have dangerous capabilities are a public good…”

A bit confused on this part, what model doesn’t have dangerous capabilities?

reasonableklout 6 hours ago | parent | next [-]

It seems possible to deliberately not train on some offensive capabilities and still have a very useful model. For example, Opus 5 deliberately did not train on exploiting vulnerabilities, and so performed less well on exploits than Mythos, yet was equally proficient at finding such vulnerabilities, according to the Opus 5 system card in their "OSS-Fuzz" eval [1].

[1]: https://www.securityweek.com/anthropics-opus-5-nears-mythos-...

philipkglass 6 hours ago | parent | next [-]

That's a good example, but I'm unsettled by Anthropic's growing refusals in the areas of chemistry and biology. If they think that scientific assistant models should be as unhelpful as Fable, because applied scientific knowledge is inherently dangerous, I don't want Anthropic or like-minded thinkers setting the standards for model safety.

jbstack 6 hours ago | parent | prev [-]

It's hard to imagine a frontier model being proficient at finding vulnerabilities, but not so proficient at exploiting them.

Surely finding is the hard part, and any LLM should be able to easily exploit a vulnerability it already knows about?

sanxiyn 5 hours ago | parent [-]

No, this is not the case for human security researchers and I don't see why it should be true for LLMs.

zer00eyz 6 hours ago | parent | prev [-]

Because Dario is still thinking in the past. He's having a Ben Carsons "the pyramids were to store grain" moment and no one is stopping him.

FTA > "My secondary concern is the risk that powerful AI models may be misused to carry out cyberattacks or biological attacks"

If this is the sort of attack he thinks is to be worried about then I dont know what to tell him. We already opened pandoras box on this. Look at what the Ukraine has done with open source drones (hunting people autonomously)

It takes minimal funding to build enough drones to destroy enough power infrastructure to shut down a large chunk of our grid. It takes even fewer talented resources to put that together with the help of already available AI.

The question I would ask Dario is this: what would some one like Ted Kazniski come up with given the resources of AI. It sure as shit would not be hacking or bioweapons or bombs in the mail.

IF they really gave a shit about safety, the would be funding (in conjunction with other AI companies) actual anonymous red teams (Ala wall facers) with some degree of independent over sight to put in the work that they arent. We're talking about a company that could not even keep its own harness code secure.