| ▲ | 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS(swe-rebench.com) | |||||||
| 17 points by ibragim_bad 2 hours ago | 4 comments | ||||||||
| ▲ | sathish316 27 minutes ago | parent | next [-] | |||||||
What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd? | ||||||||
| ▲ | spullara an hour ago | parent | prev | next [-] | |||||||
They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models. | ||||||||
| ▲ | dia80 an hour ago | parent | prev [-] | |||||||
Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels. | ||||||||
| ||||||||