| ▲ | lucisferre 2 hours ago | |
I wonder if some of this explains why people have been finding with Opus 5 that running it with lower settings than "High" is producing better or at least just as good results. > Not so fast. A 2025 paper(opens a new tab) from Northeastern University and the University of California, Berkeley on frontier open-source LRMs showed that between 30% and 60% of their “thinking steps” had “minimal causal impact” on the answers the models produced to benchmark math questions. Chop half of them out, and a model’s performance barely suffers. “We want to be careful when we review these chain-of-thought prompts because they may not be linked to the final output,” said Weiyan Shi(opens a new tab), one of the study’s authors | ||