Remix.run Logo
netinstructions an hour ago

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this:

Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.

Chance-Device an hour ago | parent | next [-]

What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.

I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.

overgard 27 minutes ago | parent | next [-]

I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as "apolitical", he just wants to buy the politicians, he doesn't believe in their causes. I hate these people.)

JumpCrisscross 23 minutes ago | parent [-]

> all that regulation will do at this point is help the incumbents who are failing

This really depends. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

XorNot an hour ago | parent | prev | next [-]

This is marketing.

Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?

Chance-Device an hour ago | parent [-]

It’s marketing the same way shitting your pants in public is marketing. People notice you.

krick 9 minutes ago | parent [-]

Apparently this is totally legit marketing strategy now. It truly is, especially if there are enough people who think that shitting your pants is cool, and the people that form the "market" nowadays may have a very different idea from yours about what is cool. Their ideas about coolness are very different from mine, that's for sure.

urams 44 minutes ago | parent | prev [-]

> What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?

Chance-Device 41 minutes ago | parent | next [-]

Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.

JumpCrisscross 32 minutes ago | parent [-]

> It’s got nothing to do with safety

Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.

Avicebron 28 minutes ago | parent [-]

Gatekeeping the public's access to models is "good policy" now? I suppose you think you'll get a dispensation to use Fable and Mythos?

JumpCrisscross 24 minutes ago | parent [-]

> Gatekeeping the public's access to models is "good policy" now?

Sorry, I was unclear. I mean that politicians being self serving doesn't tell you whether a policy is good or not.

matheusmoreira 17 minutes ago | parent | prev [-]

> Why do you think there is no policy appetite?

Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

Chance-Device 10 minutes ago | parent [-]

How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop?

During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.

JumpCrisscross an hour ago | parent | prev | next [-]

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because we continue to have zero evidence that aligment is an actual risk.

joe_the_user 19 minutes ago | parent | next [-]

I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?

echelon an hour ago | parent | prev | next [-]

Thank you.

We have wasted so much time and energy building up what has effectively become a marketing stunt.

Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.

JumpCrisscross 28 minutes ago | parent [-]

> We have wasted so much time and energy building up what has effectively become a marketing stunt

Genuine question: have we? AI is effectively unregulated in America.

ai_fry_ur_brain 27 minutes ago | parent | prev [-]

[dead]

justinnk an hour ago | parent | prev | next [-]

Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.

rubyfan an hour ago | parent | prev | next [-]

This is marketing+. They will look for policy action here to try to capture tax payer dollars.

andruc a minute ago | parent | next [-]

What incentive does HF have here?

cayley_graph an hour ago | parent | prev | next [-]

The timing after the release of GLM 5.2 and Kimi K3 is quite convenient, too, as an angle for regulatory quashing of open-weights models just as they're entering the mainstream conversation around usurping the American frontier labs. I accept my thinking here is conspiratorial, but there's also a hell of a lot of money on the line to encourage the unscrupulous.

ofjcihen an hour ago | parent | prev [-]

I don’t know if the initial “incident” was purposeful but I can tell that if I were in this position that would be my pivot.

mkagenius 33 minutes ago | parent | prev | next [-]

It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.

In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...

cududa 14 minutes ago | parent [-]

Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????

I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif

Nition 14 minutes ago | parent | prev | next [-]

In a way the intelligence of the AI itself allows them to offload responsibility to the AI. As you say, if one was simply writing software that did all this due to some insane programming decisions you'd be in big trouble.

Wowfunhappy 32 minutes ago | parent | prev | next [-]

Why was this test even connected to the public internet?

Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

JumpCrisscross 27 minutes ago | parent [-]

> why aren't they saying their next test will be air gapped in light of what happened?

Because they want to talk about how clever this model is for figuring out how to break out, hoping nobody asks why a company pitching its every-inflating agents as a replacement for software engineers can't ship a decent Mac client nor code a sandbox.

If they airgap it, they not only lose that PR angle, they also risk someone taking them seriously and requiring models be airgapped in general. That, in turn, trashes their sales pitch.

bbor 6 minutes ago | parent | prev | next [-]

I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.

Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.

The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”

karmasimida an hour ago | parent | prev | next [-]

Because the model capability is beyond their expectation.

This is brilliant marketing but I think it is real.

user43928 an hour ago | parent [-]

Interestingly OpenAI benchmarking 'an even more capable pre-release model' lines up with rumors of GPT-6 releasing in early August.

I hope that with the existing safety guardrails in place, they can roll it out to all users.

overgard 34 minutes ago | parent | prev | next [-]

I don't trust these people, this reads 100% like PR BS.

micromacrofoot an hour ago | parent | prev | next [-]

because "money" with a little "who's going to stop us"

31 minutes ago | parent | prev | next [-]
[deleted]
ofjcihen 40 minutes ago | parent | prev | next [-]

I’m honestly impressed that they managed to screw this up somehow.

Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee.

This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.

arisAlexis an hour ago | parent | prev | next [-]

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

pizzafeelsright 7 minutes ago | parent | next [-]

I really like this question because here is my situation and why my mind may have changed.

I do not think it is marketing directly but strategic release of info is plausible.

I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.

"I can't get access to the ~/.ssh so I will write a script to copy the file"

I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.

cayley_graph an hour ago | parent | prev | next [-]

They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.

nozzlegear 27 minutes ago | parent | next [-]

Precisely. "Aw jeez, we finally built the T-1000, but all it wants to do is kill John Connor – just like we warned! Why did I give it live ammunition and unsupervised time machine access?"

cwnyth an hour ago | parent | prev | next [-]

He wouldn't be the first reckless CEO...

mplappert an hour ago | parent | prev | next [-]

“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)

overgard 11 minutes ago | parent | next [-]

I'm fairly certain they're both malicious and stupid.

rubyfan an hour ago | parent | prev | next [-]

I would attribute it to profit motive instead of either stupidity or malice.

cryptoz an hour ago | parent | prev [-]

FWIW, I used to love this phrase but over recent years have come to understand it is quite damaging. We live in a society where evil frequently hides behind a ‘stupid’ label, and people bring this quote up to defend or soften actions that are indeed done out of specific malicious intent.

arisAlexis an hour ago | parent | prev [-]

They said: AI is becoming dangerously autonomous and capable. Proof of today's breach. Crowd "hey why didn't you say so, c'mon it's marketing". Them "we said so".

JumpCrisscross 33 minutes ago | parent | prev | next [-]

> What would change your mind on this?

Evidence of an AI doing one of the alignment things. At this point, Sam and Dario have lost credibility on this question.

w4yai an hour ago | parent | prev | next [-]

Oh... if Sam and Dario say so, then it must be true.

arisAlexis an hour ago | parent [-]

About their creation? Yes as most of inventors about their invention usually

iamnothere 3 minutes ago | parent | next [-]

Yes, just like Elizabeth Holmes. Or all the “free energy” crackpots. Or the people promoting radium baths for random ailments. Or Tesla’s late-in-life claims about wireless energy, death rays, and cosmic energy. Or purveyors of “snake oil” and all manner of “tonics”. The list goes on and on.

overgard 9 minutes ago | parent | prev | next [-]

These guys are not creators or inventors. They're hype men.

23 minutes ago | parent | prev [-]
[deleted]
fidotron an hour ago | parent | prev | next [-]

Demonstration of personal responsibility and accountability?

Or is that too much?

Terr_ an hour ago | parent | prev | next [-]

I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex:

1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!"

2. "So, uh, nobody can consistently steer our car properly, it keeps veering sideways sometimes, especially at high speeds, and people are finding sneaky ways of tricking it into slamming into barriers and turning pedestrians into pink fog..."

SpicyLemonZest 44 minutes ago | parent [-]

They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.

Avicebron 2 minutes ago | parent [-]

That's not why critics make fun of them. It's because their answer to "oh no we're accidentally creating the godhead. Someone please, give us power, your money, and praise, it's the only thing we can do."

It's vile hypocrisy. If they want to be priests, strip them of everything and they can live and work out of a concrete box in a mid-western cornfield. Why the material distraction if they are so religiously pure.

I know these people and I can tell you they aren't close to as smart as they think they are. Do you remember Yudowsky's "math petss"?

joe_the_user an hour ago | parent | prev | next [-]

I think you're making a false dictomy. The these models can be actually dangerous - in reality and the people in charge of their development can believe this is true (on various levels) but still not take it super seriously and instead mostly use the fact as marketing rather than being super cautious once they see the danger in action. This is behavior that's characteristic of extreme arrogance, which we know is rife in these circles.

29 minutes ago | parent | prev | next [-]
[deleted]
throwuxiytayq an hour ago | parent | prev [-]

I used to think people would wake the fuck up when AI starts killing people, these days I'm not so sure. Maybe if it caused an Instagram outage? Almost worked in Russia.

paxys an hour ago | parent | prev [-]

Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.

gowld 34 minutes ago | parent [-]

What's happening in Iran, if not world government?

paxys 22 minutes ago | parent | next [-]

How is whatever is happening in Iran related to a world government?

romanhounds 29 minutes ago | parent | prev [-]

Are you calling Israel the world government? What's happening in Iran is on them.