Remix.run Logo
13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS(swe-rebench.com)
17 points by ibragim_bad 2 hours ago | 4 comments
sathish316 27 minutes ago | parent | next [-]

What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?

spullara an hour ago | parent | prev | next [-]

They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.

dia80 an hour ago | parent | prev [-]

Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.

cbg0 an hour ago | parent [-]

I think you just answered your own question.

Edit: In DeepSWE Sol High scores the same as Fable High for ~1/3 of the cost.