Remix.run Logo
usef- 13 hours ago

What should they do instead? As far as I can tell from their actions they (Anth) are true believers about safety

pkulak 13 hours ago | parent | next [-]

They are true believers in preventing distillation.

usef- 13 hours ago | parent [-]

Is that anti safety? Most of the safety pilled people seem to want SOTA models closed.

There were so many security issues to patch in recent months that some open source projects stopped accepting them. I'm not sure how people still think Anthropic was lying.

nozzlegear 10 hours ago | parent | prev | next [-]

My contention is that, while they've converted and employ a legion of true believers (Dario may even be one himself), the veneration and flagellation at the altar of safety and alignment is conveniently self-serving. It demands constant scrutiny, because it only serves one single goal in the end: profit via regulatory capture.

revolvingthrow 13 hours ago | parent | prev [-]

So far the first and only sophisticated cyberattack was carried out by openai rather than by one the subversive commie models that hate us for our freedoms, so forgive me for taking their safety and alignment spiel somewhat less seriously than they’d like.

john_strinlai 12 hours ago | parent | next [-]

>So far the first and only sophisticated cyberattack

the only one that was publicly admitted to and which you happened to have read about.

even prior to ai, most companies and organizations would never admit to being breached unless irrefutable proof of the breach came to light. in other cases it is just immediately classified and the public finds out 20 years later (or, depending on the country, simply covered up).

usef- 13 hours ago | parent | prev [-]

That cyberattack is a sign of lack of safety, not of safety. It happened by accident.

"A model being tested broke out of its sandboxing and hacked a production system of a different organisation using a previously-unknown vulnerability, in order to steal test results, and that proves anthropic were worried for no reason months ago about future models being security issues" ...?

Anthropic have released a lot of security patches in recent months, so many that many open source maintainers are facing burnout (google it). It's public knowledge that they weren't lying, you can look at the code. Their ridiculous guardrails were so that projects could patch and prepare for more models like this one coming out.