Remix.run Logo
▲ OpenAI fires three safety researchers for "mishandling research information"(techcrunch.com)
82 points by trakkstar 4 hours ago | 43 comments
▲csbrooks an hour ago | parent | next [-]

Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?

▲chinathrow an hour ago | parent | next [-]

At this point in our shared timeline, I do believe that wouldn't be crazy, no.

▲loveparade 42 minutes ago | parent | prev | next [-]

Done by an internal model that is too dangerous to release.

▲righthand 28 minutes ago | parent | prev | next [-]

Figured out a way? These employees are most likely at-will.

▲lapcat 31 minutes ago | parent | prev [-]

> LLMs figured out a way to get these safety researchers fired

This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

▲ryeats 19 minutes ago | parent [-]

A Subliminal controlled human did the firing obviously.

▲staticman2 13 minutes ago | parent [-]

No Sam just does whatever ChatGPT 4 tells him to do. It was too dangerous to release but those fools did it anyway.

There's no deception it's very straightforward per this 2023 post:

"I mean, what if most of this is just ChatGPT [4 era] running the company..."

https://news.ycombinator.com/item?id=35281863

▲AndrewDucker 3 hours ago | parent | prev | next [-]

Fired OpenAI researchers say they were let go for 'prioritising safety'

https://www.bbc.co.uk/news/articles/cvlydn8d3lkjo

▲fn-mote 2 hours ago | parent | next [-]

The BBC link is fine, but the TechCrunch article contains all of that information and more.

▲testfrequency an hour ago | parent [-]

Except, the BBC headline is more respectful and clear as to what happened..TC is a bit vague

▲s0ss 15 minutes ago | parent [-]

Why do we care about the composition of the headlines?

▲s_dev 2 hours ago | parent | prev | next [-]

Leopold Aschenbrenner said the exact same thing after he was fired from OpenAI. It's a great excuse to explain a sudden loss of employment to others so you're still employable.

It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.

▲ericb 29 minutes ago | parent [-]

I think when it is one at a time, that's reasonable to suspect.

But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?

▲embedding-shape 2 hours ago | parent | prev | next [-]

Seems they're well aware they got fired for sharing private company information with 3rd parties, the submission article contains their admission of this:

> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them

I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.

▲thatsabadlook 2 hours ago | parent | next [-]

To be fair that is nearly exactly how openai makes its money. What's good for the gander is good for the goose.

▲irthomasthomas 2 hours ago | parent | prev [-]

Where is the admission?

▲watwut an hour ago | parent | prev [-]

Ok, but considering how weird definitions of "safety" are floating out of these companies, it does not mean much.

▲fredoliveira 30 minutes ago | parent [-]

I guess I'll ask what weird definitions of safety you've been seeing.

▲realo 7 minutes ago | parent | prev | next [-]

In related news...

"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."

▲staticman2 an hour ago | parent | prev | next [-]

> continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.

How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?

▲peri-cl 3 hours ago | parent | prev | next [-]

> "[...]OpenAI told her she’d been fired because she accessed an executive’s email. “OpenAI delegated that access to me for recruiting,”"

How exactly does this work? Struggling to comprehend the scenario.

▲rcr-anti an hour ago | parent | next [-]

There's a feature in most enterprise email, say Outlook, where you can delegate access to an inbox/address without sharing creds. Very common and normal use case, either for assistants/secretaries, common/shared inboxes, that kind of thing.

▲Ozzie_osman 3 hours ago | parent | prev | next [-]

Sometimes a recruiter or hiring manager wants to do outreach as if it's coming from a more senior person, with the assumption that the candidates are more likely to respond.

Assuming this is what was intended, there are far more secure ways of doing this.

▲TeMPOraL 2 hours ago | parent [-]

This IMO shouldn't be done more securely. It should be considered fraud.

▲pasquinelli 2 hours ago | parent [-]

not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.

▲disgruntledphd2 2 hours ago | parent [-]

> not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.

it was most likely for her team, which would explain why she was doing it.

▲pasquinelli an hour ago | parent | next [-]

if she's hiring for her own team why does she need to use someone else's email?

▲creativeSlumber an hour ago | parent | prev [-]

even if it's for her team, should have been a recruiters job.

▲nradov an hour ago | parent [-]

This may shock you but in agile, growing organizations employees sometimes have multiple job responsibilities. I've done a bit of recruiting even though I'm not a recruiter or hiring manager.

▲pasquinelli an hour ago | parent [-]

and did you use someone else's email for that?

▲chatmasta 3 hours ago | parent | prev | next [-]

This is pretty common practice for EAs.

▲comboy 3 hours ago | parent | prev | next [-]

Your have powerful agents at your disposal, so hey how to best optimize for increasing my payroll? On it. But since the agent was lunched by her, well there's consequences to ones actions, right?

(just to be clear, this is made up)

▲QuadmasterXLII 3 hours ago | parent | prev [-]

“Hi chatgpt! Please set up alice to get emails sent to me from bob so she can coordinate his inferviews. here is my gmail username and password”

Chain Of Thought: I dont have bob’s email. I don’t have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to alice…”

▲varjag 2 hours ago | parent | prev | next [-]

The purges will continue until reported AI safety improves.

▲htrp 34 minutes ago | parent | prev | next [-]

are these the employees that invited the METR team to do a debrief on huggingface?

▲diamondDrill 12 minutes ago | parent | prev | next [-]

'safety researchers' lol

▲loopglitch26 3 hours ago | parent | prev | next [-]

at this rate open ai will be "something" without it's people

▲oh_ok_lol 24 minutes ago | parent | prev | next [-]

Oh, ok. lol.

▲himata4113 3 hours ago | parent | prev [-]

You can coax openai models into hacking critical infrastructure* so I am not surprised that these people were sounding alarms at a time where openai appears to be struggling as they're failing to compete with anthropic and this months chinese models (should) be around the corner, notably a new revision of kimi should be coming out really soon.

* It's not easy, but it's possible. Although the techniques are more basic than one would expect because at the end of the day words dictate the line between what is criminal and what is not.

▲gnfargbl 3 hours ago | parent [-]

You could hack critical infrastructure before AI. Any and all of the bulk internet scanners have had lists of exposed critical infrastructure for quite a while now. At first it was shocking that nothing ever got done about it, then it became routine.

All that AI has done is to lower the bar of entry for criminal activity. Which is a concern, but it's not the primary concern. The primary concern remains that so much critical infrastructure is poorly secured.

▲abm53 2 hours ago | parent | next [-]

There’s obviously a strong interaction between the hackability of the target and the economic value of hacking the target.

Perhaps the latest models change that relationship in a meaningful way.

▲himata4113 2 hours ago | parent | prev [-]

The bigger problem here is that you can hack everything, all at once, for very cheap.

Don't get me wrong I have general disgust towards these companies that are trying to get regulatory capture on AI when they can't even secure their own systems. I believe if people know that a random AI agent can hack their systems they will put in a lot more effort into making sure it doesn't happen. This is a personal example, but I didn't really care about securing few systems as I knew no human would be ever interested in finding a vulnerability in proprietary software, however, AI has no concept of that and would hack a random rpi server running a completely undocumented unknown API just because it can't distinguish value and it costs nothing.

▲pixl97 22 minutes ago | parent [-]

Ya, quantity is a quality in of itself. In the past hackers may have used something unimportant to get a foothold but almost always tried to get to worthwhile machines. An AI will compromise everything in the network it can quickly simply because it can (assuming the attacker has a large budget, but I'll assume they stole the tokens).

It's like a new form of spam. Only far more dangerous.