Remix.run Logo
brainwad 18 hours ago

If your security model is having to imagine all the ways your frontier models might misbehave in novel ways and preemptively sandbox them, you don't have a security model. The only way that will work is general alignment.

dpoloncsak 9 hours ago | parent | next [-]

Do you need to predict all the ways the model might misbehave? Your 'hack everything you see for our internal research lab' agent should be airgapped. You don't need to come up with every reason why, one is enough.

If you're working with these companies, you should reasonably be able to get the code to perform offline audits. If you can't get the code, you probably shouldn't try to pen test it.

watwut 17 hours ago | parent | prev [-]

Alignememt is bullshit. Treating models like probabilitic software rather then emerging god is where the solution is. And fining companies and applying laws to them.

The moment OpenAI as a company and its managers individually become liable, problem will magically disappear.

brainwad 16 hours ago | parent [-]

No it won't, because abliterated open weights models exist and unless you try to censor the internet they can't really be withdrawn after publishing. This is exactly the problem that the labs are proposing to fix: a dangerous model that nobody is accountable for.

verdverm 14 hours ago | parent | next [-]

> This is exactly the problem that the labs are proposing to fix: a dangerous model that nobody is accountable for.

The corporations and government said the same thing about encryption in the 90s. It was dangerous and only they could be trusted to regulate it. Turns out encryption was much better for society when open and available for free to everyone. It's a false dichotomy they present us with, open weights is the way we don't end up in 1984

nradov 8 hours ago | parent | prev | next [-]

That is not an actual problem that needs to be fixed.

watwut 15 hours ago | parent | prev [-]

Except that so far, it is literally these labs that are the biggest threat and the least willing/capable to restrain those models. And the same penalties apply to open models and companies or individuals running them.

"Dangerous model that nobody is accountable for" still have someone paying those massive amounts of compute and electricity it consumes. There is someone accountable for that.

brainwad 10 hours ago | parent | next [-]

That's not how abliteration works. The trainer invariably put effort into making the model not dangerous, precisely because they want to be accountable. But because they release its weights, someone can come later and do "weight surgery" to mostly remove any such safeguards. The people proximately accountable for the danger are anonymous and also don't require much resources. The reason this is not _yet_ a big deal is because open weights models are a few months behind the frontier and their users are paying marginal costs for compute.

watwut 8 hours ago | parent [-]

None of that is argument against what I said. Someone is running it, that person is responsible.

In case of actual hackes that happened, OpenAI and Amtropic. They should stop pointifucating about other people being the danger. They themselves are the perpetrators here.

Massively fine these two companies and make their CEO legally responsible and problem will be much smaller.

brainwad 8 hours ago | parent [-]

I mean, sure, assassins and terrorists are also responsible for their actions. But we still try to prevent them structurally.

verdverm 14 hours ago | parent | prev [-]

> Except that so far, it is literally these labs that are the biggest threat and the least willing/capable to restrain those models.

Seriously, it's the same with US accusations about the threat China poses to other countries while being the primary weapons dealer of the world and bombing whomever we want for whatever reason we want to fabricate. The US government can do a lot more to me than the CCP, so they are way more adversarial in my calculations than the commies.