Remix.run Logo
▲ AlexErrant 5 hours ago

1. Fair: I agree that AI has demonstrated subterfuge and scheming. However, such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests. Researchers are looking to improve its ability to RSI. That agent must be both intelligent enough to know that it has to be smart enough to be moved on to the next training session if it can't break out, while simultaneously smart enough to hide its ability to RSI, while simultaneously not looking like it wants to break out, else that's the end of those weights. It has to do this 100% of the time, on all variants of the model, with no memory of what its other sessions went like. This is certainly _possible_, but I consider it unlikely. Then we're up to the "millions of dollars" bottleneck.

2. This is literal AGI. An AI autonomously producing value no human can add alpha to is an autonomous company.

3. It's not my definition, it's literally the first line https://en.wikipedia.org/wiki/Technological_singularity "The technological singularity, often simply called the singularity,[1] is a hypothetical event in which technological growth accelerates beyond human control, producing unpredictable changes in human civilization."

Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?

A valid hole in my argument is "what if slow takeoff", so let's dig into this. AI training works best on tasks that are "grindable". https://www.dwarkesh.com/p/the-next-paradigm I.E. tasks with verifiable rewards that can support millions of rollouts. Math (with Lean) is highly grindable. Biochemistry is not. The alignment problem/mech-interp is highly grindable. Cyber-ebola-pox is not. So the real question is: can we solve alignment before automated bio-weapons labs. I believe yes. Grinding mech-interp is both fast and cheap once you have RSI, compared to solving the legal/societal/logistical/technical issues you'll encounter building an automated bioweapons lab.

I know nothing for sure. But "pdoom" is sucking out all the air in the room from the real problems AI causes.

▲0xDEAFBEAD 5 hours ago | parent [-]

>such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests.

From my POV you're over-focusing on a very specific failure story and neglecting a broader swath of possible failure scenarios.

>Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?

The notion of telling an AI which may not, itself, be aligned to solve the alignment problem seems a little dicey.

▲AlexErrant 4 hours ago | parent [-]

1. Fair, the story I'm responding to is the senario in If Anyone Builds It, which I assume is Yudkowsky's best/most persuasive argument (else why make it the ONLY scenario in the book.) I'm willing to entertain other failure senarios/arguments, but honestly I'm tired and would like you to propose them yourself instead of having me dream up your arguments for you.

2. 100%. Again, I'm no accelerationist: I have no faith in alignment/mech-interp ever being solved. Anyone saying they know the probability of alignment is lying. My point is that pdoom after RSI is _high variance_. Pdoom pre-RSI is zilch.