| ▲ | hodgehog11 7 hours ago | |
A group already (sort-of) has for the broader theory: https://arxiv.org/abs/2604.21691 It's a good list, but "alignment" is much more targetted. I attended a workshop with researchers from OpenAI and Anthropic also present, where the objective was to hash out what the problems in alignment even are. This, it turned out, was extraordinarily difficult. The problems with alignment are to do with how we define and search for this vague notion when it is mathematically ill-posed at present. | ||