| ▲ | 0xDEAFBEAD 5 hours ago | |
>such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests. From my POV you're over-focusing on a very specific failure story and neglecting a broader swath of possible failure scenarios. >Is there a hole in my "alignment problem/solve mechanistic interpretability" argument? The notion of telling an AI which may not, itself, be aligned to solve the alignment problem seems a little dicey. | ||
| ▲ | AlexErrant 4 hours ago | parent [-] | |
1. Fair, the story I'm responding to is the senario in If Anyone Builds It, which I assume is Yudkowsky's best/most persuasive argument (else why make it the ONLY scenario in the book.) I'm willing to entertain other failure senarios/arguments, but honestly I'm tired and would like you to propose them yourself instead of having me dream up your arguments for you. 2. 100%. Again, I'm no accelerationist: I have no faith in alignment/mech-interp ever being solved. Anyone saying they know the probability of alignment is lying. My point is that pdoom after RSI is _high variance_. Pdoom pre-RSI is zilch. | ||