Remix.run Logo
andai 5 hours ago

> I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it’s a core goal of our current research program.

- Jakub Pachocki (OpenAI’s Chief Scientist)

I wonder how helpful this actually is for alignment? Didn't we already determine that they know when they're being evaluated, and they just say what they think you want to hear?

SubiculumCode 3 hours ago | parent [-]

It is still helpful I believe, but your point is well taken. The problem is that we have relatively few tools for monitoring alignment, and longer loops of processing that stay in latent space means less ability to monitor.