Remix.run Logo
▲ pazimzadeh 2 hours ago

Can someone explain to me why on these benchmarks like these a higher effort level often has a lower score?

For example, GPT-6.1 Sol High gets 75.2% on DeepSWE and XHigh gets 71.9% and is more expensive

https://openai.com/index/introducing-gpt-6-1-sol/#deepswe

Also, how many times did they test each condition - just once or a few times? are they showing an average of multiple attempts, etc..

▲zamadatix 2 hours ago | parent [-]

More thought can cause the important info to leave the context or hallucinated info to be enshrined in the context and later acted upon, especially in long horizon benchmarks like DeepSWE.

With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.