| ▲ | tuvix 4 hours ago | |
So this is just a collection of citations to places where misaligned or illegal things happened in the real world? Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is. In any case it’s an interesting concept for a benchmark. | ||
| ▲ | GPerson 4 hours ago | parent [-] | |
Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies. | ||