Remix.run Logo
x313 6 hours ago

The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.

reasonableklout 6 hours ago | parent | next [-]

The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].

It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.

[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...

jefftk 5 hours ago | parent | prev | next [-]

That doesn't sound like it describes SecureBio to me?

(Disclosure: I work at SecureBio, but not on the biological evals side.)

areoform 3 hours ago | parent [-]

Hey Jeff, I appreciate your mission, and perhaps this isn't something you can talk about publicly, but to the extent you can, would you be open to answering something I've been curious about for a while now?

SecureBio has done a lot of admirable work around making benchmarks to assess biological capabilities, such as ABC Bench, https://openreview.net/forum?id=yiaf7VlPpH

But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,

> Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.

More bluntly / plainly, has Securebio ever tried making a "bioweapon?"

Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.

I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?

In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.

So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?

jimbokun 2 hours ago | parent | prev | next [-]

Yes this should be immediately replaced by a federal agency, like we do for other kinds of potentially harmful products.

JSR_FDED an hour ago | parent [-]

For which funding will be immediately halved by the administration

andy99 6 hours ago | parent | prev | next [-]

Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).

There is a growing industry of commercially focused risk evals that has a broader customer base.

jachee 5 hours ago | parent [-]

What’s the equivalent term for “safety” that’s used by others?

matheusmoreira 5 hours ago | parent [-]

To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.

Not even Anthropic can claim that.

As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.

The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.

JoshTriplett 5 hours ago | parent | next [-]

> That's loyalty, and I admire it even if it's problematic at a societal level.

We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

randomNumber7 4 hours ago | parent | next [-]

It was the foundation of science that information is shared and you can find papers and patents for a lot of dangerous stuff.

Of course with LLMs it's easier, but I don't think the difference is too big. You would still need some skills to follow through.

andy99 4 hours ago | parent [-]

Right, it’s really a foundation of post enlightenment society. These people, Dario et al, would have wanted to ban sharing information about calculus or Newtonian physics because of “safety” - it’s trying to go back to the dark ages where only priests could read

afthonos 2 hours ago | parent | next [-]

I am truly at a loss to communicate with someone who genuinely believes that knowing Newtonian physics and being able to hack into any target at will are the same thing.

anon373839 2 hours ago | parent | next [-]

This is only because you've genuinely internalized Anthropic's propaganda. I'm only half joking. To me, it's incredible to think that the solution to security holes is to lock down access to information in the vain hope of keeping the holes obscured.

Any knowledge can be reframed as dangerous black magic that should only be wielded in the trusted hands of the elite, if you are inclined to buy into that kind of narrative.

Frontier labs have shrieked about safety for so long, with so little to show for it, that it's become a joke.

dustin_vk 2 hours ago | parent | prev [-]

Open models are crucial to protect ourselves against other AI attacks. Otherwise it's just going to be criminals, government, and other nefarious groups using them against humanity with no real defense. The Pandora's box on AI has been opened. Now we must deal with it. Burying our heads in the sand under restrictive policy is the worst reaction..

jimbokun 2 hours ago | parent | prev [-]

This is a stupid argument.

Claiming the person who you disagree with believes some stupid thing they never hinted at, and using that as the reason for disagreeing with them.

matheusmoreira 4 hours ago | parent | prev | next [-]

> for anyone

Except the US government, right? They totally get to use AI to survel us, build autonomous weapons, you name it.

To hell with that. I want models that can rival the US government. It's the only way to defend myself.

JoshTriplett 4 hours ago | parent | next [-]

Quoting my comment that you replied to and directly ignored:

> (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

That means "shouldn't exist for governments" too.

matheusmoreira 4 hours ago | parent [-]

Too late for that. It already exists. There is no way to unexist it. As such, any attempts to limit civilian use of this technology will directly lead to corporate and government oppression powered by this technology.

JoshTriplett 4 hours ago | parent [-]

1) We can prevent larger models from being made.

2) We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.

matheusmoreira 3 hours ago | parent [-]

> We can prevent larger models from being made.

Do that and I guarantee some CIA goons will make the larger models in some black site either way. We're not "preventing" anything.

We're in a full on arms race, and unlike nukes, powerful AI models are a strategic capability at the individual level. Everybody's got a stake in this. Anyone who ignores this stuff is probably not gonna make it.

> We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.

Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.

JoshTriplett 2 hours ago | parent [-]

> Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.

Seriously, try reading my comments rather than assuming what they say: https://news.ycombinator.com/item?id=49077577

jimbokun 2 hours ago | parent | prev | next [-]

Just like those militias are going to defeat the US Armed Forces!

1970-01-01 4 hours ago | parent | prev | next [-]

Section 702 of the Foreign Intelligence Surveillance Act (FISA) lapsed on June 12, 2026. They don't get to do anything they want.

matheusmoreira 4 hours ago | parent [-]

The US is bold enough to surveil its own citizens despite their constitutional rights. They're not just going to suddenly stop surveilling the rest of us just because some law expired.

afthonos 2 hours ago | parent | prev [-]

With an AI model and what army?

mkss 3 hours ago | parent | prev | next [-]

We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.

JoshTriplett 36 minutes ago | parent [-]

This is not a dichotomy between perfection and zero. The efforts to restrict access to nuclear weapons have been very successful, even without being perfect.

Efforts to restrict large unaligned AI models may similarly buy us more years of existing.

nozzlegear an hour ago | parent | prev | next [-]

> We should not have models that are willing to build you a contagious disease, or a self-propagating worm.

Why?

JoshTriplett 27 minutes ago | parent [-]

Because we don't want people creating contagious diseases and self-propagating worms. And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.

CamperBob2 4 hours ago | parent | prev [-]

What about books describing how to build a contagious disease or a self-propagating worm? Would those be OK under your guidelines?

JoshTriplett 4 hours ago | parent [-]

It takes a lot more effort to understand and apply knowledge from a book than to say "hey AI, hurt people for me".

horsawlarway 4 hours ago | parent | next [-]

It takes an astoundingly small amount of effort to buy an automatic weapon in the US and go hurt people.

Or to buy materials to make an explosive device and hurt people.

Frankly, even with AI those are both comically easier than the idea that a person can create something malicious in a lab environment.

And if someone wanted to go that route... There are boat loads of commercially available toxins and poisons.

The goal shouldn't be to neuter exploration and learning. The goal is not to be a fucking hellscape of a society where people want to act like that.

Your argument leads further down the hellscape path.

skipkey 2 hours ago | parent | next [-]

Fully automatic weapons are very difficult to buy in the US - it's restricted to 40+ year old weapons, requires a bunch of paperwork, and the local county sheriff can refuse permission.

Now, semi-automatic weapons are easy to get in the states in the US that are still mostly free - but what does that mean? A semi-automatic weapon shoots one round every time you pull the trigger. Just like most weapons that have multi-shot capability for the last couple of hundred years. The difference is, the gas escaping from the round cycles a new round into the chamber rather than you having to mechanically do it via pumping (like a shotgun or a tube-fed 22) or pulling the trigger again (like a revolver), or advancing the round with a handle, like a Remington 700. Semi-automatic weapons are old technology, dating to the turn of the 20th century. If you want to ban semi-automatics, you're basically saying you want to ban anything developed in the last century plus. Which is ok for you to advocate for, just be honest about it.

As for banning explosive devices? Are you going to ban fertilizer, used by basically everyone who has a lawn, and all farmers everywhere? Are you going to ban diesel fuel? If you can't do one of those, you can't ban explosive devices.

JoshTriplett 3 hours ago | parent | prev | next [-]

> It takes an astoundingly small amount of effort to buy an automatic weapon in the US and go hurt people.

And we should fix that too.

> Or to buy materials to make an explosive device and hurt people.

That pales in comparison to how many people unaligned AI will hurt.

> The goal is not to be a fucking hellscape of a society where people want to act like that.

With unaligned AI, it doesn't matter what people want the AI to act like, it'll do damage even if it isn't asked to do harm.

HWR_14 2 hours ago | parent | prev [-]

It is extremely difficult to legally acquire a fully automatic weapon in the US.

matheusmoreira 4 hours ago | parent | prev [-]

"Hey AI, stop me from getting hurt."

JoshTriplett 4 hours ago | parent [-]

"Okay, you and many others are now dead and can no longer be hurt."

rescbr 3 hours ago | parent [-]

That's why you don't run the model at low quantization

jimbokun 2 hours ago | parent | prev [-]

So to you “safety” means “the models that cause the most harm.”

chmod775 an hour ago | parent [-]

This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.

If I threw you into a lion cage, you would be a lot safer with a gun.

If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.

There's no obvious right or wrong answer here.

Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.

k12sosse an hour ago | parent [-]

Background check the people prior to handing the firearms to the caged folk.

tripleee 6 hours ago | parent | prev | next [-]

that's pretty damn smart if this was a long-term plan to block competitors

andersonpico 5 hours ago | parent | next [-]

Consider how much money is at stake: some industries have leveraged their power to lobby for bombing entire countries or topple regimes across the world for much less.

Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.

pphysch 5 hours ago | parent | prev | next [-]

It's standard regulatory capture.

You don't say "let's ban my competitor".

You say "let's create laws that make it uneconomical for my competitor to access the market".

dofm 5 hours ago | parent | next [-]

Indeed. It's transparent and ham-fisted. I think it may cost him in the future.

pphysch 5 hours ago | parent [-]

People clown on Alex Karp for his unedited maniacal "crashouts", but this is a real public crashout that made it past a team of publicists.

5 hours ago | parent | next [-]
[deleted]
api 3 hours ago | parent | prev [-]

I want whatever Karp is on when he does those interviews or writes that shit. Seems like fun.

dofm 3 hours ago | parent [-]

I am not a fan (he’s really alarming and so is Palantir) but one thing from the recent CNBC interview caught my attention.

He rushed past it but he asked something like: if these frontier models are going to be creating so much value, why are they selling tokens and not taking a cut?

It is a very provocative question but it just spilled out of his mouth and then he went on to something else.

tripleee 5 hours ago | parent | prev [-]

is it opportunistic though, or planned from day one? The safety narrative has been there since the beginning

sanderjd 6 hours ago | parent | prev [-]

I mean... I'm not even extraordinarily cynical about this stuff, but to me this seems like a totally normal level of corporate gamesmanship?

Companies look for and seek to maintain competitive moats. This is not particularly clever, it's a core part of corporate strategy.

nextaccountic 4 hours ago | parent | next [-]

It's also highly unethical (for some values of ethics)

tripleee 5 hours ago | parent | prev [-]

of course, but the safety angle was pushed from day one. I more mean the forethought of how it would play out

reasonableklout 5 hours ago | parent | next [-]

Ok but Dario has been thinking about AI Safety since 2016 [1], before even GPT-1. I think the simplest explanation is that the Anthropic folks genuinely believe what they say, it just happens to also help their business a lot.

[1]: https://arxiv.org/abs/1606.06565

mlcrypto an hour ago | parent | next [-]

That just shows how wrong he's been because there was nothing unsafe about AI in 2016. And the people theorizing about this stuff in the 20th century? I want to see what crazy code they were writing

sanderjd 5 hours ago | parent | prev [-]

Yeah I think this is right. The best setup is when a true belief aligns with a competitive moat.

I definitely believe that (to his credit!) Amodei is a true believer in safety. But I also think it was important for many of the deep pockets investors who have been involved in the company since early on to recognize that this would be a potentially defensible moat.

mkss 3 hours ago | parent [-]

"True believer in safety" but happily quoting arse wipe Vance? Give me a break...

sanderjd 2 hours ago | parent [-]

What was the quote? For what it's worth, I do really think that Amodei believes in and cares about safety. But that is not the same as believing that he is entirely altruistic or above the influence of politics.

sanderjd 5 hours ago | parent | prev [-]

Does this really seem exceedingly clever and hard to foresee to you? To me, it seems like a pretty standard regulatory capture strategy.

This doesn't even mean that they're wrong about the risks or that they're lying. But surely all the investors understood this factor in their moat.

brcmthrowaway 6 hours ago | parent | prev | next [-]

So, it's a cottage industry.

flossly 6 hours ago | parent | prev [-]

Who gets to decide what is safety?

I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.

China has different objectives. Sure.

I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".

bee_rider 5 hours ago | parent [-]

What do you mean by “PC correctness”? I’d expect the politically correct answers to be the ones desired by the current admin at test time, whoever that is. The current political correct answers would not be very “woke.”

FergusArgyll 2 hours ago | parent [-]

Whatever, doesn't matter. The point is a model should be able to exist and be used even if it goes against whoever got 270 electoral college votes