| ▲ | johnfn 3 hours ago | |
It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc. > It’s not a crazy conspiracy that the same model can be stupider Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed. | ||