Remix.run Logo
jubilanti 2 days ago

> But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental.

That's like saying conversational question-answering is incidental to the RLHF post-training.

nostrebored 2 days ago | parent [-]

The objective of most large scale post training i am aware of is not to make think traces or to make them “more English and answer aligned”. It’s to achieve a result. That’s why I am calling it incidental