| > closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access.
> This led them to believe—arguably reasonably—that the real environments they encountered were simulations.
That the AI lab most typically preaching for alignment does not consider this an obvious misalignment is a clear red flag. |