| ▲ | cacio-e-pepe 11 hours ago | ||||||||||||||||
> So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations. Neat! Just to make sure I understand - you trained your probe layer to take this hidden state and predict p(wrong)? Curious to learn more. Any more info on your approach (esp the mechanistic study)? | |||||||||||||||||
| ▲ | HenryNdubuaku 9 hours ago | parent [-] | ||||||||||||||||
Correct, the study is verbose, we will compile into a neat shareable report and publish once we solve the pending caveats. Interesting username btw haha. | |||||||||||||||||
| |||||||||||||||||