| ▲ | notpachet 4 hours ago | |
Related reading: The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property. https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig... | ||