| ▲ | bunderbunder 5 hours ago | |
It’s hard to say. But supposedly the counter example wasn’t found by an agent running in full auto; it came out of a bunch of back and forth with a human operator. Without, in addition to the aforementioned access to currently non-public information about these models, a detailed transcript of the chat sessions leading up to the discovery, it’s hard to ascribe the reasoning steps involved to any source in particular. Part of my concern here is that simply pointing out that LLMs appear to be performing tasks that can be done through reasoning, and using that in and of itself as evidence of reasoning, is affirming the consequent. | ||