Remix.run Logo
iforgotmypasswo 3 hours ago

This is so much more interesting than what people looking for immediate criminal punishment and people referring to AI as next token generators are focusing on.

First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve.

The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as “AI School”. And the AI is trying to cheat! Because it’s easier and there’s an incentive to do so! Just like humans! That’s wild.

Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that’s not complicated, difficult, or the interesting part.

What’s interesting here is that we need proctoring and monitoring at a scale that allows training.

I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. “Oh, shit! We’ve accidentally trained it to hack into systems to accomplish its goals!” Can you imagine the kind of day that would give you?

You failed to make it smarter. You didn’t catch it cheating, and you instead incentivized cheating. Bad day!

This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That’s a new idea for me, though I’m sure it’s old news to others.

How do we build training systems which make cheating impossible?

How do we simulate systems where cheating is possible, where AI thinks it’s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears.

Sure, I’m actively concerned about AI killing us all in 10 years. But there’s a whole field of AI psychology brewing here, and it’s interesting as hell.

iforgotmypasswo 3 hours ago | parent [-]

Side note, you could absolutely create an AI sleeper agent by simulating dates and times during training to effectively flip a switch. I guarantee AI systems from other countries will be banned from accessing products which manage controlled or export restricted information as those sorts of techniques are further developed.