| ▲ | hnd9q09qk4 2 hours ago | |
Exact match graders are the real variable here, we had F1 or a judge model swing passage QA scores by ten points on identical answers. | ||
| ▲ | Theory42 2 hours ago | parent [-] | |
A good point. I initially had a model for a judge, but it seemed to give very lenient scores. I'm open to learning about how best to benchmark the phenomenon, though. | ||