| ▲ | theChris-in 9 hours ago | |||||||
We can do a mix of general use (as in user stories) plus a few academic benchmarks. So you have any specific ideas? You can hmu at iam@thechris.in | ||||||||
| ▲ | GodelNumbering 9 hours ago | parent [-] | |||||||
Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_... > We can do a mix of general use (as in user stories) plus a few academic benchmarks. Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness. > So you have any specific ideas? Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability. | ||||||||
| ||||||||