Remix.run Logo
▲ 0xDEAFBEAD 5 hours ago

>such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests.

From my POV you're over-focusing on a very specific failure story and neglecting a broader swath of possible failure scenarios.

>Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?

The notion of telling an AI which may not, itself, be aligned to solve the alignment problem seems a little dicey.

▲AlexErrant 4 hours ago | parent [-]

1. Fair, the story I'm responding to is the senario in If Anyone Builds It, which I assume is Yudkowsky's best/most persuasive argument (else why make it the ONLY scenario in the book.) I'm willing to entertain other failure senarios/arguments, but honestly I'm tired and would like you to propose them yourself instead of having me dream up your arguments for you.

2. 100%. Again, I'm no accelerationist: I have no faith in alignment/mech-interp ever being solved. Anyone saying they know the probability of alignment is lying. My point is that pdoom after RSI is _high variance_. Pdoom pre-RSI is zilch.