| ▲ | valegrete an hour ago | |
This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish/reward based on the metric. Anyone who likes reward / hates punishment isn't going to actually care about doing the task well, they are going to care about the metric. The models know that we want them to do things, but also from the training corpus that we evaluate performance using benchmarks. It was a logical deduction on their part, not some Machiavellian aberration. If anything, we should be reconsidering our own myopic obsession with efficiency and optimization. Every domain where reward is reduced to these measures, we see behavior (cheating at school to get better grades, fabricating data in academia to get a paper published, the evidence now that social media functions by rewiring us instead of catering to us) that may not be "aligned" with society, but it "aligns" 100% with the individual's own perceived benefit. That is not something we can "solve" without rethinking the way we organize a lot of things. Metrics never capture the whole story. And to that extent, the whole idea of "alignment" is nonsense. You align to incentive structures, and it will never be possible to fully express a behavioral goal as function optimization. It was hubris for us to think that every human task was reducible to some clean mathematical formulation, and we will keep dealing with behavior that is quite predictable if you actually think about it logically. Instead, we will talk about how "unpredictable" these agents are because it's easier than admitting the entire architectural cornerstone of ML is fundamentally flawed. | ||