| ▲ | socializer 2 hours ago | |
> I really would like to see their tests and the model’s reasoning traces. Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then. That, and labs aren't serious about sandboxing their evals. | ||