| ▲ | MichaelGlass 4 hours ago | |
I see a lot of claims in this article without ... any proof? Both can be true: - It's useful to anthropomorphize agents when predicting behavior and - we have to use specific language to specify what we mean. What does the author mean by "confuse the models" ? Are they talking about not picking right information? Picking the wrong information? Losing their previous context / task? Part of setting up a proper eval is also deciding what we actually mean ourself. What are we actually optimizing for? It's not, e.g. % confusion, %rubbish, etc. The article does point to it: retrieval latency, accuracy, etc. | ||