| ▲ | vessenes 2 days ago | |
Anthropic’s Mechinterp did some very fine work on this. TLDR - you can; you train a decoder on neuralese to english and then add a loss function for a roundtrip of english -> neuralese -> english (or possibly n -> e -> n? I don’t recall), giving a pretty strong indication that you have a good ‘translation’. They published open weights versions of these interpreters for a number of open models sometime in the last year. Very cool idea. By the way, they concluded CoT often lied, based on the neuralese interpretation. EDIT: a comment below linked to https://www.anthropic.com/research/natural-language-autoenco..., which is what I was referring to. | ||