| ▲ | smallmancontrov 3 hours ago | |
Sibling posts are correct -- the chain-of-thought is doing hidden computation, it has been shown in the linked papers. If you want to see it yourself: load up Qwen 3.8 in LM Studio and watch the CoT stumble around like a drunken sailor before miraculously jumping to the correct result. If you want an example of subversion, Anthropic has some good ones: https://transformer-circuits.pub/2025/attribution-graphs/bio... https://transformer-circuits.pub/2025/attribution-graphs/bio... | ||