| ▲ | Agents on Rails: Best model solves 35% of feature benchmark runs(rubyonrails.org) | |||||||
| 20 points by chalmovsky a day ago | 5 comments | ||||||||
| ▲ | andhuman a day ago | parent | next [-] | |||||||
This benchmark expects models to not only do the happy path, but also edge cases, even though the ticket they give to the model doesn’t specify if. This to align more to real world tickets. Usually with benchmarks it’s the other way around: only implement what’s asked. So I welcome this type of benchmark because this is how I use the models. | ||||||||
| ▲ | klooney 20 hours ago | parent | prev | next [-] | |||||||
Is this a proprietary harness? I feel like harnesses have a huge influence on how models behave | ||||||||
| ||||||||
| ▲ | chalmovsky a day ago | parent | prev | next [-] | |||||||
this seems really low? | ||||||||
| ▲ | entaroadun123 10 hours ago | parent | prev [-] | |||||||
[flagged] | ||||||||